01 / Decide

Jeffery Gyamerah enterprise data & applied AI

People decide.Systems prove.

Enterprise data & supply-chain leader (GSK, Haleon) who builds AI agent systems that are measured, not promised.

Portrait of Jeffery Gyamerah
Jeffery Gyamerah · Panama

Post Graduate Program in AI Agents and Generative AI for Business Applications — McCombs School of Business, UT Austin (via Great Learning). Program completion certificate, completed September 2026. Verify (opens in a new tab)

LinkedIn (opens in a new tab)Email

02 / The through-line

The plant floor taught the rule before the agents did.

My route into AI systems ran through business operations — from SAP production planning and manufacturing data at GSK to Haleon, where I led data engagement and delivery for the global SAP IBP transformation and now lead Quality & Supply Chain data strategy. The daily pain of manual workflow taught me what automation actually has to earn.

Employer work · GSK and Haleon

Described qualitatively.

  1. 2017 SAP production-planning go-lives at manufacturing sites across several countries
  2. 2019 Product owner for planning solutions used across manufacturing plants
  3. 2021 Global information standards for a manufacturing network
  4. 2022 Data foundations after the GSK–Haleon separation
  5. 2024 Data engagement and delivery for the global SAP IBP transformation
  6. 2026 Quality & Supply Chain data strategy

Coursework & competition builds

Academic and hackathon work — not employer or client deployments.

  1. Jun–Aug 2026 Three graded coursework proofs of concept, each with an evaluation gate
  2. In the programme Practice builds: guardrails, tool use, policy checks
  3. Sep 2026 Two hackathon prototypes with a human stop rule
  4. Sep 2026 Programme completed — certificate from McCombs, UT Austin, via Great Learning

I completed the postgraduate program in AI agents in September 2026. The projects below are academic coursework and proofs of concept, not production client deployments.

03 / Selected work

Three builds. Each time, the evidence said no to the easy option.

Academic coursework and proofs of concept from the postgraduate program, in the order I built them — each one harder than the last. Every case leads with the decision, then shows the architecture, the measured result, what was rejected and what it changed.

The architecture diagrams animate the order data moves through each pipeline.

Case 01

Support Ticket Analysis

Academic coursework · proof of concept — not a production or client deployment

Submitted
2026-06-11
Focus
Single pipeline + LLM-as-judge
Stack
gpt-4o-mini · gpt-4o · LLM-as-judge · pandas

The decision

A low-scoring summary was flagged, not waved through.

The gpt-4o judge scores every summary. A score below 4 on any criterion flags that summary for human review; the flag is not evidence that a person has already reviewed it.

The problem

A POC to summarize unstructured support tickets, judge summary quality, generate empathetic customer responses, judge them, and consolidate each output for downstream use.

Architecture

A sequential two-tier pipeline: gpt-4o-mini generates summaries and responses, while gpt-4o evaluates both stages before consolidation.

Support ticket analysis: animated architectureSupporttickets30 unstructuredticketsLoad +validatepandasSummarizegpt-4o-mini ·category · urgencySummary judgegpt-4o · 4 criteriaDraft responsegpt-4o-miniResponse judgegpt-4o · 4 criteriaConsolidate +exportone row per ticketHuman review6 of 30 summariesscore < 4Data in / outDeterministic codeLLM stepLLM judge / gateHuman reviewResultSupport ticket analysis: animated architectureSupport tickets30 unstructuredticketsLoad + validatepandasSummarizegpt-4o-mini · category· urgencySummary judgegpt-4o · 4 criteriaDraft responsegpt-4o-miniResponse judgegpt-4o · 4 criteriaConsolidate +exportone row per ticketHuman review6 of 30 summariesscore < 4Data in / outDeterministic codeLLM stepLLM judge / gateHuman reviewResult
Two generate-then-judge stages; a summary scoring below 4 on any criterion goes to a person.

Actual result

Summaries flagged for human review
6/30
Judge averages out of 5: summaries / responses
4.51 / 4.48
Placeholder leaks
0

The weakest result, kept visible

Ticket 3 is the weakest summary. It scored 2.25 out of 5 after the model over-committed to a concrete quality_defect category and medium urgency despite the ticket lacking information.

What it changed

A judge only enforces the failure modes its rubric explicitly defines.

All measured results (4) — Support Ticket Analysis
Submission results from the academic coursework run.
Measured outputSubmission result
Summary overall score4.51/5
Response overall score4.48/5
Placeholder leakagezero placeholder leaks
Summaries flagged for review6/30

Case 02

Grounded Synthesis

Academic coursework · proof of concept — not a production or client deployment

Submitted
2026-07-19
Focus
RAG + evaluation discipline
Stack
Chroma · DeepEval GEval · GEPA · gpt-4o-mini · gpt-4o

The decision

The automatically tuned prompt lost on unseen questions. The hand-written prompt shipped.

GEPA, a tool that rewrites prompts automatically, produced a version that scored better on its own practice questions. On six questions neither prompt had seen, it scored lower — so the held-out test made the call.

The problem

GlobalEdge brokers needed grounded answers from financial news, price data, and SEC filings without the manual review burden or loss of evidence behind their responses.

Architecture

Chroma RAG with chunk and top-k sweeps, DeepEval GEval judging, held-out prompt comparison, and a locked winning configuration.

Grounded synthesis RAG: animated architectureSourcesnews · prices · SECfilingsChunk + indexChroma · 1,495documentsRetrieve top-kbest sweep: k = 8Generateanswergpt-4o-mini · sayswhat is missingGEval judgegpt-4o · grounded /actionableHeld-out gatehand prompt 0.708 ·GEPA 0.658Lock confighand-written promptkeptFinal held-out5 questions · 5/5passconfig sweepData in / outDeterministic codeLLM stepLLM judge / gateResultGrounded synthesis RAG: animated architectureSourcesnews · prices · SECfilingsChunk + indexChroma · 1,495documentsRetrieve top-kbest sweep: k = 8Generate answergpt-4o-mini · sayswhat is missingGEval judgegpt-4o · grounded /actionableHeld-out gatehand prompt 0.708 ·GEPA 0.658Lock confighand-written promptkeptFinal held-out5 questions · 5/5 passconfig sweepData in / outDeterministic codeLLM stepLLM judge / gateResult
Retrieval settings were swept against a judge; the prompt was chosen on held-out questions, not by the optimizer.

Actual result

Held-out mean score: hand-written prompt vs GEPA-optimized prompt
0.708 vs 0.658
Final held-out questions passed
5/5
Final held-out groundedness / broker actionability
0.820 / 0.860

Rejected: the GEPA-optimized prompt

GEPA underperformed the hand-authored prompt on held-out data. The submission kept Baseline v1 rather than shipping the optimized prompt.

Honest limitation

The starting point was a 45.0% baseline benchmark pass rate, and the held-out sets are small — the prompt comparison is a gate, not a statistical claim.

What it changed

An optimizer that regresses trustworthiness is not a win.

All measured results (6) — Grounded Synthesis
Submission results from the academic coursework run.
Measured outputSubmission result
Baseline benchmark pass rate45.0%
Final held-out Groundedness0.820
Final held-out Broker Actionability0.860
Final held-out result5/5 pass (100%)
Held-out prompt comparisonGEPA rejected: v1 0.708 vs v2 0.658
Indexed documents1,495
Read the write-up: The optimizer I didn’t ship

Case 03

Last-Mile Delivery Exception Handling

Academic coursework · proof of concept — not a production or client deployment

Submitted
2026-08-17
Focus
Multi-agent + cost / latency engineering
Stack
LangGraph · Chroma · gpt-4o-mini · gpt-4o · LangSmith

The decision

Cheaper critics were measured, then refused.

Running both critics on gpt-4o-mini would have made the run meaningfully cheaper. It was tested on an isolated copy first, cost accuracy on a borderline escalation, and was turned down.

The problem

A POC for a mid-sized retailer to ingest noisy delivery logs, resolve actionable exceptions, escalate when policy requires, generate customer notifications, and retain auditable evaluation metrics.

Architecture

LangGraph pipeline with retrieved playbook context, deterministic routing, two critic stages, capped revision loops, and an auditable final result.

Last-mile delivery exception handling: animated architectureDelivery logsnoisy CSV · 10shipmentsPreprocessordedupe · noise ·injection scanOrchestratorrule-based routerResolutionagentgpt-4o-mini +playbook RAGResolutioncriticgpt-4o · accept /revise / escalateCommunicationagentgpt-4o-mini · toneper customerCommunicationcriticgpt-4oFinalizenotify or escalate· audit trailEvaluation5 metrics per casenoise → $0REVISE · max 2ACCEPTData in / outDeterministic codeLLM stepLLM judge / gateResultLast-mile delivery exception handling: animated architectureDelivery logsnoisy CSV · 10shipmentsPreprocessordedupe · noise ·injection scanOrchestratorrule-based routerResolution agentgpt-4o-mini + playbookRAGResolutioncriticgpt-4o · accept /revise / escalateCommunicationagentgpt-4o-mini · tone percustomerCommunicationcriticgpt-4oFinalizenotify or escalate ·audit trailEvaluation5 metrics per casenoise → $0REVISE · max 2Data in / outDeterministic codeLLM stepLLM judge / gateResult
Deterministic code routes and caps the loop; gpt-4o-mini writes, gpt-4o critics judge. Noise exits before any model call.

Actual result

Escalation accuracy in the final run — 7 of 8 scored cases
88%
LLM calls across 10 cases; 2 of 10 shipments cost $0
46
Task completion · 6.07s mean end-to-end latency
10/10

Rejected: downgrading both critics to gpt-4o-mini

An earlier, separate run: escalation accuracy was 100% with gpt-4o critics and fell to 88% when both critics were downgraded — the downgraded critic missed SHP-008, a discretionary-escalation case. Downgrade rejected. That experiment is distinct from the final run’s 88% (7 of 8 scored cases).

Honest limitation

n=10 is small. Metrics moved between real runs; a production pilot needs a larger, more adversarial test set before any single run’s percentages are trusted.

What it changed

Aggregate metrics start the investigation; traces reveal the failure mode.

All measured results (8) — Last-Mile Delivery Exception Handling
Submission results from the academic coursework run.
Measured outputSubmission result
Task completion100% (10/10)
Tool-call accuracy100%
Escalation accuracy88% — 7 of 8 scored cases (SHP-003 over-escalated; 2 noise cases unscored)
Reasoning coherence4.4/5
Mean end-to-end latency6.07s
Run cost engineering46 LLM calls across 10 cases
Guardrail / noise short-circuit2/10 shipments cost $0
Critic downgrade experiment (earlier 100% run)100% → 88% — downgrade rejected
Read the write-up: 46 LLM calls, 10 shipments, 2 that cost $0

04 / Stop the line

Under a deadline, the stop rule stays on.

Two competition prototypes built on synthetic data. In both, the system stops and a person makes the decision that matters.

hackIAthon 2026 · Challenge 2 · September 2026

JIDOKA

Stop rule

The agent never approves a payment.

Claims-invoice auditor. Deterministic rules check each repair-shop invoice against the claim file and the agreed rate card; every finding cites its evidence line by line. A person accepts, requests evidence or dismisses each finding, and the agent never approves a payment. An LLM drafts the report only from findings the rules already proved.

Team: Jeffery Gyamerah and Gilberto (@cpu-16)

Stack: Next.js 16 · Prisma + SQLite · Rule engine · Human-in-the-loop · Synthetic data

Decentralized AI Hackathon 2026 · Tether QVAC Psy track · 9–11 September 2026

Notare

Stop rule

Nothing is saved until the clinician approves the exact text.

Local clinical-documentation prototype. A printed note or consented voice note is captured on Android, transcribed locally with QVAC VisionPsy, reviewed and corrected by the clinician on a Windows PC, and saved as an encrypted, patient-scoped record only after the clinician approves the exact text.

Stack: QVAC VisionPsy · Local LLMs · Electron · React Native · Encrypted records · Synthetic data

Source code is private; available to reviewers on request. Clinical validation is not claimed.

05 / The record

The full record, for the second read.

Industry experience — 2017 to nowEmployer work · qualitative

  1. Jan 2026 — Present

    Global Data Strategy Lead — Quality & Supply Chain

    Haleon · Panama City, Panama

    Designs the Quality & Supply Chain data-governance model and the operating model for how data is defined, owned and used across functions, markets and plants.

  2. Sep 2024 — Jan 2026

    Global Data Engagement & Delivery Lead — SAP IBP Transformation

    Haleon · Panama City, Panama

    Led process, data, role and control changes for the global SAP IBP migration; end-to-end owner for the North America and Europe markets.

  3. Apr 2022 — Sep 2024

    End-to-End Data Engagement Lead — Quality & Supply Chain

    Haleon · Panama City, Panama

    Built data-governance and stewardship foundations after the GSK–Haleon separation. Acting head of the data-engagement team in 2024.

  4. Jun 2025 — Present

    Board Member

    APROVIVA · Panama

    Part-time. Strategy, project management and governance for a residential community.

  5. Jan 2021 — Apr 2022

    Global Data Management Subject Matter Expert

    GSK · United Kingdom

    Global owner of information standards and governance for the manufacturing network; chaired the monthly global data-management forum.

  6. Feb 2019 — Feb 2020

    SAP Production Planning Product Owner & Logistics Manager

    GSK · United Kingdom

    Product owner for planning solutions used across manufacturing plants; managed material schedulers, suppliers and inventory.

  7. Apr 2017 — Jan 2021

    Founder & Managing Director

    GCO Solutions Ltd

    Technology platform for barbering businesses and professionals. Side venture, concurrent with GSK employment.

  8. Jan 2017 — Sep 2019

    SAP Production Planning Functional Consultant

    GSK · United Kingdom / international

    Manufacturing SAP deployments, including international go-live assignments.

Enterprise delivery

ERP and planning-data delivery — GSK SAP production-planning deployments (2017–2019), then Haleon’s planning transformation, from data-stewardship foundations to programme-level execution.

PeriodDeliveryContribution
2024Data-quality turn-around — European marketsLed cross-functional data-quality improvement and remediation across European markets.
2024Large-scale master-data remediationCoordinated master-data remediation with business owners and regional teams.
2025Deployment-wave readiness — SAP IBP planning transformationLed planning-data readiness, validation and remediation for regional deployment waves.
2024–25Operating model and governanceEstablished cross-functional governance and readiness tracking, with documentation and training handed to business-as-usual teams.
2024Acting head of the data teamRan team operations while leading strategic initiatives; proposed new roles and mentored the team, which later ran European deployment waves autonomously.

Credential

Post Graduate Program in AI Agents and Generative AI for Business Applications

McCombs School of Business, The University of Texas at Austin — via Great Learning

Program completion certificate · completed September 2026

A programme completion certificate, not a degree.

Verify the certificate (opens in a new tab)

Practice builds

Smaller academic builds from the programme, exercising one capability each.

  • Customer support agent

    Kartify

    Single-agent ReAct · SQL tool · Input and output guardrails · Evaluation retry threshold 0.75

  • Expense workflow automation

    Reimbursement

    FastMCP server · 9-tool single agent · Multimodal OCR · Validation and policy steps

  • Responsible AI claim auditing

    ClaimAudit

    Input guardrails · SELECT-only SQL safety · PII-aware output checks · Audit logging and bias tracking

Capability map

Agent orchestration
LangGraph · Critic loops · Revision caps
Retrieval
Chroma · Chunk + top-k sweeps · DeepEval GEval · GEPA
Evaluation
LLM-as-judge design · 3-layer QA
Guardrails and responsible AI
Input / output checks · PII-aware controls
Tool integration
MCP · SQL · OCR
No-code
n8n
Multimodal GenAI
Multimodal OCR · Vision-based document extraction

06 / Contact

Let’s discuss the evidence.

Based in Panama · data strategy, governance and applied AI