Evaluating Agents and Multi-Agent Systems
A support agent that resolves 95% of tickets correctly sounds like a success until you look at how: one version reaches the right answer in 3 tool calls every time, another reaches it in anywhere from 2 to 14 depending on what it happened to try first, and a third gets there by calling a refund tool twice — once correctly, once by mistake, luckily canceling itself out. All three can post the same 95% "accuracy" number on a dashboard, and that number is the least useful thing you could report about the difference between them. Evaluating an agent is not the same problem as evaluating a single LLM call, and the gap between the two is exactly what this subject is built around: an agent's output is the end of a trajectory — a sequence of decisions, tool calls, and intermediate states — and a trajectory can be wasteful, fragile, or one lucky roll away from failing even when the final answer it happened to produce this time was correct.
llm-observability-and-evaluation already owns the general machinery this subject leans on constantly and does not re-explain: what a request trace needs to capture, offline evals vs. online monitoring as two complementary loops, LLM-as-judge's core mechanics and calibration discipline, CI regression gates, and canarying. Read that subject first if you haven't — this one assumes it and adds only what's specific to evaluating a multi-step, tool-using system rather than a single generative call: how to score a trajectory and not just an answer, the metrics unique to tool use and step count, the precise difference between pass@k and pass^k, what changes when the judge is looking at a process instead of a final response, sandboxed environments as judge alternatives, and the coordination-specific failure modes that only exist once more than one agent is in the loop. agent-architectures-and-the-agentic-loop and multi-agent-orchestration cover the architectures these evaluation methods are applied to — the loop shapes, the supervisor/worker patterns, the compounding-error math, the handoff design — and are cross-linked throughout rather than repeated.