Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — Evaluating RAG Systems (5 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Advanced Open Free

Diagnosing a RAG Metrics Dashboard Permalink →

Your team's weekly offline eval dashboard for a RAG-based support assistant shows:

  • recall@5 = 0.65
  • judge correctness (vs. gold answer) = 0.90
  • faithfulness = 0.80

Your manager sees the 0.90 correctness number and says "that's a great score, let's ship this to more traffic." Walk through how you would respond: what do these three numbers actually tell you together, what do you suspect is really happening, and what would you fix first before shipping wider? Be specific about what evidence would confirm or rule out your hypothesis.

Share this question

Advanced Open Pro

Computing Retrieval Metrics From Raw Results

Unlock this question →
Advanced Open Pro

Designing a Gold Evaluation Set From Scratch

Unlock this question →
Advanced Open Pro

Calibrating an LLM-as-Judge Pipeline

Unlock this question →
Advanced Open Pro

Building a Hallucination-Detection Pipeline for Production

Unlock this question →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.