Practice — LLM Observability and Evaluation (5 questions)
Advanced
Open
Free
Diagnosing a Two-Week Quality Regression from a Dashboard Permalink →
You are the on-call AI engineer for a RAG-based support assistant. This morning you notice, on the standing dashboard:
- The "no confident source" rate (fraction of requests where retrieval confidence fell below the answer threshold) has climbed steadily from 4% to 13% over the last two weeks.
- The daily LLM-as-judge faithfulness score has dropped from 0.94 to 0.86 over the same window.
- Both trends are gradual — no single sharp step change.
- Your deploy log shows no prompt, model, or retrieval-config change in that window.
- Explain why these two symptoms are likely connected rather than coincidental.
- List, in the order you would check them, at least three plausible root causes given that nothing in your deploy log changed.
- Describe exactly what you would pull from request traces to confirm (not just guess at) the cause.
- What would you change afterward so this class of regression is caught automatically next time, both before and after it reaches production?
Share this question
Advanced
Open
Pro
Designing a CI Regression Gate for Prompt and Model Changes
Unlock this question →
Advanced
Open
Pro