Every RAG project reaches the moment where someone — a teammate, a tutorial, a paper — says: don’t hand-tune the prompt, run an optimizer.
For my brokerage intelligence assistant (500 news documents, 500 price documents, 126 SEC filing pages in one Chroma store; gpt-4o-mini generating, gpt-4o judging), that someone was GEPA — genetic-pareto prompt optimization, built into DeepEval. Feed it a training split, let it mutate and select prompts, ship the winner. Why would my hand-written prompt survive an evolutionary search?
The baseline was honest about its limits: 45% pass on a 20-question golden set — groundedness 0.600, actionability 0.580. Plenty of room for an optimizer to “win.”
I let GEPA run. Then I did the part that matters: scored both prompts on six held-out questions neither had seen.
- Baseline v1 — groundedness / actionability
- 0.733 / 0.683 · mean 0.708 · 50% pass
- GEPA v2 — groundedness / actionability
- 0.667 / 0.650 · mean 0.658 · 33% pass
The optimizer had optimized — for its training data. On evidence it hadn’t seen, it was worse on every metric that mattered. I shipped neither instinct nor hype; I shipped the gate: BEST_PROMPT = Baseline v1.
Six held-out questions is small — I’ll say it first. But the discipline doesn’t change with n: the set the optimizer never saw is the only vote that counts. A bigger budget buys a bigger held-out set, not a different rule.
The compounding came after. With the prompt locked, I swept retrieval configs. One knob — widening top-k from 5 to 8 — lifted both metrics (groundedness 0.655→0.700), the opposite of my starting hypothesis that tighter chunks would win. The locked pipeline scored 0.820 groundedness and 0.860 actionability on five final held-out broker questions, 5/5 pass.
The result I’m proudest of isn’t the 0.820. It’s two of those five final answers: asked for revenue comparisons the evidence couldn’t support, the system said “the data is missing” instead of inventing figures. That behavior lived in the hand-written prompt GEPA wanted to replace. For a brokerage, a confident wrong answer is worse than a partial one — honesty is the trust product.
What I kept:
- Verify the optimizer; don’t assume it. An optimizer that regresses trustworthiness is not a win.
- Keep tuning data and reported data separate — reusing the tuning set as your “final” number inflates the grade.
- Prompt rules like “state what’s missing” are product features, not decoration.
- Aggregates start the investigation; per-question traces end it. The worst baseline failures were all SEC-filing questions — a section-aware chunking problem no global knob fixed.
The tools will keep changing — GEPA today, something else next quarter. The gate doesn’t: held-out evidence decides, or you’re just generating confidence.