Advanced
Open
Pro
Calibrating an LLM-as-Judge and Defending Against Its Biases
Your team wants to use an LLM-as-judge to score 10,000 production conversations a day for correctness, because human review of that volume is impossible. Leadership is excited about the judge score trending up after a recent prompt change.
- Before trusting that upward trend, what would you check, and why?
- Name three specific, known biases of LLM judges and, for each, describe a concrete scenario where it would produce a misleading score for this system.
- Design a pairwise comparison setup (judge picks the better of two answers) that avoids one of those biases specifically.
- What ongoing process keeps the judge trustworthy over time, rather than calibrated once and then drifting silently?
Share this question