Advanced
Open
Pro
Designing an Evaluation Plan Before Launch
Your team is two weeks from launching an LLM assistant that drafts replies to customer emails using retrieved policy documents. Leadership asks: "How will we know it is good, and how will we know it has not got worse after a prompt or model change?"
- Design an offline evaluation set and the metrics you would compute on it, distinguishing retrieval from generation.
- Explain how you would use LLM-as-judge, and name three caveats and how you mitigate each.
- Describe the online plan and the guardrail metrics that would stop a rollout.
Share this question