Three Agents, Same Outcome Accuracy: What the Dashboard Hides
Three versions of a document-retrieval agent each score 92% on outcome accuracy against the same 150-task eval set. Trace review reveals: Version A averages 4 tool calls per task with tight correlation between tool calls made and task difficulty; Version B averages 11 tool calls per task, largely by re-running the same search with minor query rewording when the first few results look "not quite right"; Version C averages 4 tool calls per task but 15% of its trajectories include a tool call with an argument that doesn't match what the retrieved context actually supports (a citation to a document ID that was never actually retrieved in that trajectory).
- For each version, name the specific trajectory-level problem (if any) that outcome accuracy alone would hide.
- Which version would you be most worried about shipping, and why — rank them and justify the ranking using the metrics this subject defines.
- Propose the one additional metric or check that would have caught each version's problem before it reached a dashboard.
1. What outcome accuracy hides per version
Version A has no trajectory-level problem apparent from what's described — 4 tool calls correlated with actual task difficulty is the efficient, well-formed shape trajectory evaluation is looking for. Its outcome accuracy is a reasonably trustworthy summary of its quality.
Version B has a step-efficiency problem: nearly 3x the tool calls of Version A for the same outcome accuracy, driven by redundant re-search rather than genuine task complexity. Outcome accuracy is blind to this because the redundant calls don't change whether the final answer was right — they only change cost and latency, both invisible to an outcome-only check.
Version C has a tool-call accuracy problem specifically around argument correctness — a tool call citing a document ID that wasn't actually retrieved in that trajectory is a hallucinated citation masquerading as a correct-looking action. This is the most dangerous of the three because it can coexist with a "correct" final answer by coincidence (the answer happens to be right even though the citation supporting it is fabricated) — exactly the brittleness-to-future-edge-cases risk this subject names: nothing about the current eval set is stopping this from becoming a visibly wrong citation the moment the retrieved-context distribution shifts even slightly.
2. Ranking by risk
Version C is the most concerning to ship: a fabricated citation that happens to coincide with a correct answer is a correctness and trust problem waiting to surface in production the moment luck runs out, and it's the kind of error a user or downstream system can be actively misled by (trusting a citation that doesn't actually support the claim). Version B is a real but lower-severity problem — it's a cost and latency issue, not a correctness or trust issue; it costs money and time but the answers it eventually gives are (per the given accuracy) as reliable as Version A's. Version A is the one to actually ship as-is.
3. The catching check for each
Version B: track step efficiency against a small set of hand-verified reference trajectories per task, or at minimum trend median tool-calls-per-task over time — a 3x gap against a reference trajectory (or against Version A's own number, once compared side-by-side) would have flagged this before it reached a dashboard that only shows outcome accuracy.
Version C: tool-call accuracy checked specifically for argument correctness against what was actually retrieved in that trajectory — not just "is this a syntactically valid tool call" but "does this citation's document ID appear among the documents this specific trajectory actually retrieved." This is exactly the kind of mechanical, trace-level check that doesn't require a judge's opinion — the tool call and the retrieval log in the same trace are enough to verify it automatically, and it should be run on 100% of trajectories, not sampled, since it's cheap to compute mechanically.
Share this question