Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — Evaluating Agents and Multi-Agent Systems (5 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Advanced Open Free

Three Agents, Same Outcome Accuracy: What the Dashboard Hides Permalink →

Three versions of a document-retrieval agent each score 92% on outcome accuracy against the same 150-task eval set. Trace review reveals: Version A averages 4 tool calls per task with tight correlation between tool calls made and task difficulty; Version B averages 11 tool calls per task, largely by re-running the same search with minor query rewording when the first few results look "not quite right"; Version C averages 4 tool calls per task but 15% of its trajectories include a tool call with an argument that doesn't match what the retrieved context actually supports (a citation to a document ID that was never actually retrieved in that trajectory).

  1. For each version, name the specific trajectory-level problem (if any) that outcome accuracy alone would hide.
  2. Which version would you be most worried about shipping, and why — rank them and justify the ranking using the metrics this subject defines.
  3. Propose the one additional metric or check that would have caught each version's problem before it reached a dashboard.

Share this question

Advanced Open Pro

Computing and Interpreting pass@k and pass^k for a Batch-Processing Agent

Unlock this question →
Advanced Open Pro

Diagnosing a Trajectory Judge That Rewards the Wrong Thing

Unlock this question →
Advanced Open Pro

Credit Assignment in a Three-Stage Research Pipeline

Unlock this question →
Advanced Open Pro

Comparing Two Proposals with Cost-Per-Successful-Task

Unlock this question →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.