Advanced
Open
Pro
A Vendor Cohort's Quality Is Degrading — Find It Before It Ships
Part of the AI Engineer Interview path →
Part of the Reinforcement Learning & Long-term Optimization path →
Your annotation platform serves comparisons to ~220 vendor annotators and 15 internal experts. This cycle's inter-annotator agreement (IAA), aggregated across the whole pool, has dropped from a stable 0.71 (Cohen's kappa) over the last four cycles to 0.61 this cycle. The prompt mix and guidelines are unchanged from last cycle.
- Give two distinct hypotheses for the drop that are consistent with "guidelines and prompt mix unchanged," and explain what data you'd pull to distinguish between them.
- Suppose you find the drop is concentrated in one 40-person vendor cohort added three weeks ago, with all other annotators near their historical baseline. What's your immediate action on the training data collected so far this cycle, and why?
- What ongoing mechanism, if it had existed before this cycle, would have surfaced this problem within days instead of at the aggregate end-of-cycle IAA check?
Share this question