Post Graduate Program in AI Agents and Generative AI for Business Applications — McCombs School of Business, UT Austin (via Great Learning). Program completion certificate, completed September 2026. Verify ↗ (opens in a new tab)
The plant floor taught the rule before the agents did.
My route into AI systems ran through business operations — from SAP production planning and manufacturing data at GSK to Haleon, where I led data engagement and delivery for the global SAP IBP transformation and now lead Quality & Supply Chain data strategy. The daily pain of manual workflow taught me what automation actually has to earn.
Employer work · GSK and Haleon
Described qualitatively.
2017SAP production-planning go-lives at manufacturing sites across several countries
2019Product owner for planning solutions used across manufacturing plants
2021Global information standards for a manufacturing network
2022Data foundations after the GSK–Haleon separation
2024Data engagement and delivery for the global SAP IBP transformation
2026Quality & Supply Chain data strategy
Coursework & competition builds
Academic and hackathon work — not employer or client deployments.
Jun–Aug 2026Three graded coursework proofs of concept, each with an evaluation gate
In the programmePractice builds: guardrails, tool use, policy checks
Sep 2026Two hackathon prototypes with a human stop rule
Sep 2026Programme completed — certificate from McCombs, UT Austin, via Great Learning
I completed the postgraduate program in AI agents in September 2026. The projects below are academic coursework and proofs of concept, not production client deployments.
03 / Selected work
Three builds. Each time, the evidence said no to the easy option.
Academic coursework and proofs of concept from the postgraduate program, in the order I built them — each one harder than the last. Every case leads with the decision, then shows the architecture, the measured result, what was rejected and what it changed.
The architecture diagrams animate the order data moves through each pipeline.
Case 01
Support Ticket Analysis
Academic coursework · proof of concept — not a production or client deployment
Submitted
2026-06-11
Focus
Single pipeline + LLM-as-judge
Stack
gpt-4o-mini · gpt-4o · LLM-as-judge · pandas
The decision
A low-scoring summary was flagged, not waved through.
The gpt-4o judge scores every summary. A score below 4 on any criterion flags that summary for human review; the flag is not evidence that a person has already reviewed it.
The problem
A POC to summarize unstructured support tickets, judge summary quality, generate empathetic customer responses, judge them, and consolidate each output for downstream use.
Architecture
A sequential two-tier pipeline: gpt-4o-mini generates summaries and responses, while gpt-4o evaluates both stages before consolidation.
Two generate-then-judge stages; a summary scoring below 4 on any criterion goes to a person.
Actual result
Summaries flagged for human review
6/30
Judge averages out of 5: summaries / responses
4.51 / 4.48
Placeholder leaks
0
The weakest result, kept visible
Ticket 3 is the weakest summary. It scored 2.25 out of 5 after the model over-committed to a concrete quality_defect category and medium urgency despite the ticket lacking information.
What it changed
A judge only enforces the failure modes its rubric explicitly defines.
All measured results (4) — Support Ticket Analysis
Submission results from the academic coursework run.
Measured output
Submission result
Summary overall score
4.51/5
Response overall score
4.48/5
Placeholder leakage
zero placeholder leaks
Summaries flagged for review
6/30
Case 02
Grounded Synthesis
Academic coursework · proof of concept — not a production or client deployment
The automatically tuned prompt lost on unseen questions. The hand-written prompt shipped.
GEPA, a tool that rewrites prompts automatically, produced a version that scored better on its own practice questions. On six questions neither prompt had seen, it scored lower — so the held-out test made the call.
The problem
GlobalEdge brokers needed grounded answers from financial news, price data, and SEC filings without the manual review burden or loss of evidence behind their responses.
Architecture
Chroma RAG with chunk and top-k sweeps, DeepEval GEval judging, held-out prompt comparison, and a locked winning configuration.
Retrieval settings were swept against a judge; the prompt was chosen on held-out questions, not by the optimizer.
Actual result
Held-out mean score: hand-written prompt vs GEPA-optimized prompt
0.708 vs 0.658
Final held-out questions passed
5/5
Final held-out groundedness / broker actionability
0.820 / 0.860
Rejected: the GEPA-optimized prompt
GEPA underperformed the hand-authored prompt on held-out data. The submission kept Baseline v1 rather than shipping the optimized prompt.
Honest limitation
The starting point was a 45.0% baseline benchmark pass rate, and the held-out sets are small — the prompt comparison is a gate, not a statistical claim.
What it changed
An optimizer that regresses trustworthiness is not a win.
All measured results (6) — Grounded Synthesis
Submission results from the academic coursework run.
Running both critics on gpt-4o-mini would have made the run meaningfully cheaper. It was tested on an isolated copy first, cost accuracy on a borderline escalation, and was turned down.
The problem
A POC for a mid-sized retailer to ingest noisy delivery logs, resolve actionable exceptions, escalate when policy requires, generate customer notifications, and retain auditable evaluation metrics.
Architecture
LangGraph pipeline with retrieved playbook context, deterministic routing, two critic stages, capped revision loops, and an auditable final result.
Deterministic code routes and caps the loop; gpt-4o-mini writes, gpt-4o critics judge. Noise exits before any model call.
Actual result
Escalation accuracy in the final run — 7 of 8 scored cases
88%
LLM calls across 10 cases; 2 of 10 shipments cost $0
46
Task completion · 6.07s mean end-to-end latency
10/10
Rejected: downgrading both critics to gpt-4o-mini
An earlier, separate run: escalation accuracy was 100% with gpt-4o critics and fell to 88% when both critics were downgraded — the downgraded critic missed SHP-008, a discretionary-escalation case. Downgrade rejected. That experiment is distinct from the final run’s 88% (7 of 8 scored cases).
Honest limitation
n=10 is small. Metrics moved between real runs; a production pilot needs a larger, more adversarial test set before any single run’s percentages are trusted.
What it changed
Aggregate metrics start the investigation; traces reveal the failure mode.
All measured results (8) — Last-Mile Delivery Exception Handling
Submission results from the academic coursework run.
Two competition prototypes built on synthetic data. In both, the system stops and a person makes the decision that matters.
hackIAthon 2026 · Challenge 2 · September 2026
JIDOKA
Stop rule
The agent never approves a payment.
Claims-invoice auditor. Deterministic rules check each repair-shop invoice against the claim file and the agreed rate card; every finding cites its evidence line by line. A person accepts, requests evidence or dismisses each finding, and the agent never approves a payment. An LLM drafts the report only from findings the rules already proved.
Decentralized AI Hackathon 2026 · Tether QVAC Psy track · 9–11 September 2026
Notare
Stop rule
Nothing is saved until the clinician approves the exact text.
Local clinical-documentation prototype. A printed note or consented voice note is captured on Android, transcribed locally with QVAC VisionPsy, reviewed and corrected by the clinician on a Windows PC, and saved as an encrypted, patient-scoped record only after the clinician approves the exact text.
Stack: QVAC VisionPsy · Local LLMs · Electron · React Native · Encrypted records · Synthetic data
Source code is private; available to reviewers on request. Clinical validation is not claimed.
Industry experience — 2017 to nowEmployer work · qualitative
Jan 2026 — Present
Global Data Strategy Lead — Quality & Supply Chain
Haleon · Panama City, Panama
Designs the Quality & Supply Chain data-governance model and the operating model for how data is defined, owned and used across functions, markets and plants.
Sep 2024 — Jan 2026
Global Data Engagement & Delivery Lead — SAP IBP Transformation
Haleon · Panama City, Panama
Led process, data, role and control changes for the global SAP IBP migration; end-to-end owner for the North America and Europe markets.
Apr 2022 — Sep 2024
End-to-End Data Engagement Lead — Quality & Supply Chain
Haleon · Panama City, Panama
Built data-governance and stewardship foundations after the GSK–Haleon separation. Acting head of the data-engagement team in 2024.
Jun 2025 — Present
Board Member
APROVIVA · Panama
Part-time. Strategy, project management and governance for a residential community.
Jan 2021 — Apr 2022
Global Data Management Subject Matter Expert
GSK · United Kingdom
Global owner of information standards and governance for the manufacturing network; chaired the monthly global data-management forum.
Feb 2019 — Feb 2020
SAP Production Planning Product Owner & Logistics Manager
GSK · United Kingdom
Product owner for planning solutions used across manufacturing plants; managed material schedulers, suppliers and inventory.
Apr 2017 — Jan 2021
Founder & Managing Director
GCO Solutions Ltd
Technology platform for barbering businesses and professionals. Side venture, concurrent with GSK employment.
Jan 2017 — Sep 2019
SAP Production Planning Functional Consultant
GSK · United Kingdom / international
Manufacturing SAP deployments, including international go-live assignments.
Enterprise delivery
ERP and planning-data delivery — GSK SAP production-planning deployments (2017–2019), then Haleon’s planning transformation, from data-stewardship foundations to programme-level execution.
Period
Delivery
Contribution
2024
Data-quality turn-around — European markets
Led cross-functional data-quality improvement and remediation across European markets.
2024
Large-scale master-data remediation
Coordinated master-data remediation with business owners and regional teams.
2025
Deployment-wave readiness — SAP IBP planning transformation
Led planning-data readiness, validation and remediation for regional deployment waves.
2024–25
Operating model and governance
Established cross-functional governance and readiness tracking, with documentation and training handed to business-as-usual teams.
2024
Acting head of the data team
Ran team operations while leading strategic initiatives; proposed new roles and mentored the team, which later ran European deployment waves autonomously.
Credential
Post Graduate Program in AI Agents and Generative AI for Business Applications
McCombs School of Business, The University of Texas at Austin — via Great Learning
Program completion certificate · completed September 2026