Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — LLM Observability and Evaluation (5 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Advanced Open Free

Diagnosing a Two-Week Quality Regression from a Dashboard Permalink →

You are the on-call AI engineer for a RAG-based support assistant. This morning you notice, on the standing dashboard:

  • The "no confident source" rate (fraction of requests where retrieval confidence fell below the answer threshold) has climbed steadily from 4% to 13% over the last two weeks.
  • The daily LLM-as-judge faithfulness score has dropped from 0.94 to 0.86 over the same window.
  • Both trends are gradual — no single sharp step change.
  • Your deploy log shows no prompt, model, or retrieval-config change in that window.
  1. Explain why these two symptoms are likely connected rather than coincidental.
  2. List, in the order you would check them, at least three plausible root causes given that nothing in your deploy log changed.
  3. Describe exactly what you would pull from request traces to confirm (not just guess at) the cause.
  4. What would you change afterward so this class of regression is caught automatically next time, both before and after it reaches production?

Share this question

Advanced Open Pro

Designing a CI Regression Gate for Prompt and Model Changes

Unlock this question →
Advanced Open Pro

Calibrating an LLM-as-Judge and Defending Against Its Biases

Unlock this question →
Advanced Open Pro

Designing a Trace Schema for a Tool-Using Agent

Unlock this question →
Advanced Open Pro

Canarying a Prompt Change with Rollback Criteria

Unlock this question →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.