Evaluating Agents and Multi-Agent Systems
Outcome vs. trajectory evaluation, tool-call accuracy, pass@k vs. pass^k, LLM-as-judge over trajectories, sandboxed benchmarks, cost-per-successful-task, and multi-agent-specific failure modes
Agent- and multi-agent-specific evaluation for AI engineering interviews: outcome vs. trajectory evaluation and why trajectory matters even when the answer is right; concrete agent metrics (tool-call accuracy, step efficiency, task completion rate); pass@k vs. pass^k explained and measured with a worked numeric example, framed as evaluation metrics rather than a design-time cost calculation; LLM-as-judge applied to trajectories specifically, with the biases unique to judging a process rather than an answer; sandboxed benchmark environments (τ-bench/SWE-bench-style) as the rigorous alternative; cost-per-successful-task as the unifying production metric; multi-agent-specific coordination failures, credit assignment, and handoff loss; agent-specific regression testing and canarying; and a worked 200-task eval set where pass@1 and pass^3 disagree on a ship decision.
Practice questions (5)
-
View →
Three Agents, Same Outcome Accuracy: What the Dashboard Hides
Advanced · Free -
View →
Computing and Interpreting pass@k and pass^k for a Batch-Processing Agent
Advanced -
View →
Diagnosing a Trajectory Judge That Rewards the Wrong Thing
Advanced -
View →
Credit Assignment in a Three-Stage Research Pipeline
Advanced -
View →
Comparing Two Proposals with Cost-Per-Successful-Task
Advanced