Advanced
Open
Pro
Designing a Gold Evaluation Set From Scratch
You're joining a team that has a RAG system in production with no formal evaluation set — quality has been judged so far by engineers eyeballing a handful of chat transcripts. You have access to three months of real conversation logs. Design the gold evaluation set: where the examples come from, how you decide what to include, how big it should be, and how it gets used day to day once built. Be specific about what categories of example a set built only from "typical" logged questions would be missing, and why that gap matters.
Share this question