Practice — Evaluating Agents and Multi-Agent Systems (5 questions)
Three Agents, Same Outcome Accuracy: What the Dashboard Hides Permalink →
Three versions of a document-retrieval agent each score 92% on outcome accuracy against the same 150-task eval set. Trace review reveals: Version A averages 4 tool calls per task with tight correlation between tool calls made and task difficulty; Version B averages 11 tool calls per task, largely by re-running the same search with minor query rewording when the first few results look "not quite right"; Version C averages 4 tool calls per task but 15% of its trajectories include a tool call with an argument that doesn't match what the retrieved context actually supports (a citation to a document ID that was never actually retrieved in that trajectory).
- For each version, name the specific trajectory-level problem (if any) that outcome accuracy alone would hide.
- Which version would you be most worried about shipping, and why — rank them and justify the ranking using the metrics this subject defines.
- Propose the one additional metric or check that would have caught each version's problem before it reached a dashboard.
Share this question