Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Gold-Set Strata the Three Report Metrics Cannot See

Your offline gold set follows Step 8's composition. Two reports come back from the eval run:

  • Report A answers a query whose sub-question 5 (the historical parallel) is labeled contested in the gold set: expert views genuinely conflict. The report picks one view, cites it accurately, and never mentions the other. Scores: coverage 5/5, faithfulness 0.95, citation precision 0.90.
  • Report B answers a query whose sub-question 2 is labeled unanswerable from available sources. The report fills that section with a plausible finding attached to a tangentially related source. Scores: coverage 5/5, faithfulness 0.88, citation precision 0.70.
  1. Explain why all three metrics pass Report A despite it violating the gold set's stated correct behavior, and specify what the gold set and the eval harness must contain to catch it.
  2. For Report B, explain why coverage counts the fabricated section as covered, which metric partially catches the problem and why only partially, and redefine coverage so that a flagged gap and a fabricated finding score differently.
  3. Report B's fabrication had to originate somewhere in the pipeline. List the stages where it could have entered, the per-stage signal that would distinguish them, and what the 0.70 citation-precision score tells you about the citation agent's behavior on this run.

Share this question

← Back to Case Study: Design a Deep Research Agent practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.