Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Choosing a Comparison UI for a New Adversarial Category

The red-team lead wants annotators to evaluate the model's responses to a new adversarial category: prompts that embed a subtly false technical claim, where a good response should correct the claim without being preachy about it. She proposes collecting a 1-5 Likert "how well did the response handle this" score from annotators, because it's fast to build and gives one number per response to trend over time.

  1. What's the risk in using a Likert scale as the primary data source for this category, given what this case says about relative versus absolute human judgment?
  2. Propose an alternative collection design for this specific category, and explain what each part of your design is for.
  3. The red-team lead still wants a trendable "how are we doing on this category over time" number for a leadership dashboard. How do you give her that without making Likert the primary training-data mechanism?

Share this question

← Back to Case Study: Designing an RLHF / Preference-Tuning Platform practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.