Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — Prompt Evaluation and Versioning (5 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Intermediate Open Free

Building a Golden Set From Scratch for a Refund-Request Classifier Permalink →

You're taking over a prompt that classifies refund requests into approve, deny, or escalate_to_human. There is currently no golden set at all — the previous owner "tested" changes by pasting a handful of examples into a playground and eyeballing the output before shipping. You have access to six months of production logs (roughly 40,000 requests) with no labels, plus a support-ticket archive showing which requests were later overturned on appeal.

  1. Describe concretely how you would build the first version of the golden set from this raw material — what you'd sample, how much, and what you'd deliberately over-represent.
  2. What labeling process would you put in place to avoid the common pitfalls of inconsistent or self-referential labels?
  3. Six months from now, how do you know if the golden set has gone stale, and what would you do about it?

Share this question

Intermediate Open Pro

Diagnosing a Suspiciously Confident LLM-as-Judge

Unlock this question →
Intermediate Open Pro

Setting a CI Regression Threshold That Doesn't Cry Wolf

Unlock this question →
Intermediate Open Pro

A/B Test Looks Like a Win — Until You Check the Guardrail Metrics

Unlock this question →
Intermediate Open Pro

A Metric Regressed but the Prompt Diff Is Clean — Now What?

Unlock this question →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.