Paths Subjects Questions Quizzes Pricing Search
Advanced Open Pro

Designing a Gold Evaluation Set From Scratch

You're joining a team that has a RAG system in production with no formal evaluation set — quality has been judged so far by engineers eyeballing a handful of chat transcripts. You have access to three months of real conversation logs. Design the gold evaluation set: where the examples come from, how you decide what to include, how big it should be, and how it gets used day to day once built. Be specific about what categories of example a set built only from "typical" logged questions would be missing, and why that gap matters.

Share this question

← Back to Evaluating RAG Systems practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.