Advanced
Open
Pro
Choosing a Comparison UI for a New Adversarial Category
Part of the AI Engineer Interview path →
Part of the Reinforcement Learning & Long-term Optimization path →
The red-team lead wants annotators to evaluate the model's responses to a new adversarial category: prompts that embed a subtly false technical claim, where a good response should correct the claim without being preachy about it. She proposes collecting a 1-5 Likert "how well did the response handle this" score from annotators, because it's fast to build and gives one number per response to trend over time.
- What's the risk in using a Likert scale as the primary data source for this category, given what this case says about relative versus absolute human judgment?
- Propose an alternative collection design for this specific category, and explain what each part of your design is for.
- The red-team lead still wants a trendable "how are we doing on this category over time" number for a leadership dashboard. How do you give her that without making Likert the primary training-data mechanism?
Share this question