Ask a seven-node agent pipeline what it costs and most people answer with a shrug and a token price. The better answer is a ledger: which node, which model, how many calls, per case. Here is the ledger for a last-mile delivery exception-handling POC — a LangGraph pipeline (preprocessor → orchestrator → resolution → critic → communication → critic → finalize) that ran 10 noisy shipments end to end.
- Total LLM calls, 10 cases
- 46 (36 pipeline + 10 judge)
- Shipments costing $0 (noise short-circuit)
- 2/10
- Shipments costing 4 calls each
- 7/10
- REVISE case (revision loop fired)
- 1/10 — 8 calls
Three design decisions made that ledger small. The preprocessor is deterministic — dedupe, noise detection, prompt-injection scan — so routine noise never touches a model; that’s where the two $0 shipments come from. The orchestrator is a deterministic router, not an agent: guardrail blocks, escalation triggers, and a hard revision cap (max_loops=2) mean the expensive loops are bounded by code, not by hope. And generation runs on gpt-4o-mini while the two critics run on gpt-4o — the cheap model does the volume, the strong model does the judging.
Then the obvious follow-up: if gpt-4o-mini is good enough to generate, why pay for gpt-4o critics? Downgrade both critics and the run gets meaningfully cheaper. So I tested it — full pipeline, isolated scratch copy, critics on gpt-4o-mini.
Escalation accuracy dropped from 100% to 88%. The downgraded critic missed SHP-008, a discretionary-escalation case — the exact class of borderline judgment call the build’s own trace debugging had already shown was fragile. The saving was real; so was the failure. Downgrade rejected.
Run note: the 100% → 88% comparison above is the earlier critic-downgrade experiment. It is a different run from the final submitted result, where escalation accuracy was 88% — 7 of 8 scored cases.
That experiment cost one scratch run and told me where the money actually lives: you can cheapen the talking, not the judging. The review also identified possible cost and latency levers: sample the coherence judge less frequently and process independent shipments concurrently. Those are proposed changes, not a measured speedup; a sampled judge call still has a cost.
What I kept:
- Every agent boundary is a cost decision. Deterministic code in front of the pipeline is the cheapest model you’ll ever run.
- Cap loops in the router, not in the prompt.
max_loops=2is a guarantee; “please be concise” is a suggestion. - Measure the downgrade before you ship it. Cost levers fail in the borderline cases, not the average ones — and averages won’t show you that.
- Keep a per-case call ledger. “46 calls across 10 cases” is an answer; “it depends” is not.
Cost engineering in agent systems isn’t an invoice review at the end of the month. It’s a set of architectural choices — where code replaces a model, where a cheap model replaces a strong one, and which of those trades you refuse after measuring them. The ledger is the design document.