Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

A Vendor Cohort's Quality Is Degrading — Find It Before It Ships

Your annotation platform serves comparisons to ~220 vendor annotators and 15 internal experts. This cycle's inter-annotator agreement (IAA), aggregated across the whole pool, has dropped from a stable 0.71 (Cohen's kappa) over the last four cycles to 0.61 this cycle. The prompt mix and guidelines are unchanged from last cycle.

  1. Give two distinct hypotheses for the drop that are consistent with "guidelines and prompt mix unchanged," and explain what data you'd pull to distinguish between them.
  2. Suppose you find the drop is concentrated in one 40-person vendor cohort added three weeks ago, with all other annotators near their historical baseline. What's your immediate action on the training data collected so far this cycle, and why?
  3. What ongoing mechanism, if it had existed before this cycle, would have surfaced this problem within days instead of at the aggregate end-of-cycle IAA check?

Share this question

← Back to Case Study: Designing an RLHF / Preference-Tuning Platform practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.