Paths Subjects Questions Quizzes Pricing Search
Advanced Open Pro

Building Golden Sets Retroactively Under a Deadline

36 of your 40 production prompts have no golden set. You have roughly 55 days left after finishing the triage in Step 1. There is no hand-written "ideal answer" reference for most of these prompts — just two years of production traffic.

  1. Describe how you'd construct a golden set for a prompt that has never had one, using only production traces.
  2. For open-ended generation prompts with no verifiable ground truth, what do you use as the target to evaluate against, and how do you trust it?
  3. Given the time constraint, how do you size golden sets differently across the high-traffic/high-criticality, high-traffic/ low-criticality, and long-tail buckets from Step 1?

Share this question

← Back to Case Study: Migrate a Production Prompt Suite Across a Model Deprecation practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.