Advanced
Open
Pro
Building Golden Sets Retroactively Under a Deadline
36 of your 40 production prompts have no golden set. You have roughly 55 days left after finishing the triage in Step 1. There is no hand-written "ideal answer" reference for most of these prompts — just two years of production traffic.
- Describe how you'd construct a golden set for a prompt that has never had one, using only production traces.
- For open-ended generation prompts with no verifiable ground truth, what do you use as the target to evaluate against, and how do you trust it?
- Given the time constraint, how do you size golden sets differently across the high-traffic/high-criticality, high-traffic/ low-criticality, and long-tail buckets from Step 1?
Share this question