Practice — Prompt Evaluation and Versioning (5 questions)
Intermediate
Open
Free
Building a Golden Set From Scratch for a Refund-Request Classifier Permalink →
You're taking over a prompt that classifies refund requests into
approve, deny, or escalate_to_human. There is currently no golden
set at all — the previous owner "tested" changes by pasting a handful
of examples into a playground and eyeballing the output before shipping.
You have access to six months of production logs (roughly 40,000
requests) with no labels, plus a support-ticket archive showing which
requests were later overturned on appeal.
- Describe concretely how you would build the first version of the golden set from this raw material — what you'd sample, how much, and what you'd deliberately over-represent.
- What labeling process would you put in place to avoid the common pitfalls of inconsistent or self-referential labels?
- Six months from now, how do you know if the golden set has gone stale, and what would you do about it?
Share this question
Intermediate
Open
Pro
Setting a CI Regression Threshold That Doesn't Cry Wolf
Unlock this question →
Intermediate
Open
Pro
A/B Test Looks Like a Win — Until You Check the Guardrail Metrics
Unlock this question →
Intermediate
Open
Pro