Advanced
Open
Pro
Diagnosing a Trajectory Judge That Rewards the Wrong Thing
A team introduces an LLM-as-judge to score agent trajectories for a code-review-assistant agent, rubric: "rate how thorough and careful this trajectory was, 1-10." After a month, they notice: judge scores have crept up 1.5 points on average, but a separately-tracked step- efficiency metric shows trajectories have gotten 40% longer over the same period, and a human spot-check of 20 recent high-judge-scoring trajectories finds several with irrelevant tool calls (re-reading files already read, re-running a search that returned nothing new) padded in.
- Name the specific bias at work and explain the mechanism by which it produced this exact pattern of evidence.
- Is this evidence of the agent actually improving, actually regressing, or something else? Justify using all three signals given (judge score, step efficiency, human spot-check).
- Rewrite the rubric instruction that caused this, and explain specifically what changed and why that fixes the mechanism from part 1.
Share this question