Advanced
Open
Pro
Gold-Set Strata the Three Report Metrics Cannot See
Your offline gold set follows Step 8's composition. Two reports come back from the eval run:
- Report A answers a query whose sub-question 5 (the historical parallel) is labeled contested in the gold set: expert views genuinely conflict. The report picks one view, cites it accurately, and never mentions the other. Scores: coverage 5/5, faithfulness 0.95, citation precision 0.90.
- Report B answers a query whose sub-question 2 is labeled unanswerable from available sources. The report fills that section with a plausible finding attached to a tangentially related source. Scores: coverage 5/5, faithfulness 0.88, citation precision 0.70.
- Explain why all three metrics pass Report A despite it violating the gold set's stated correct behavior, and specify what the gold set and the eval harness must contain to catch it.
- For Report B, explain why coverage counts the fabricated section as covered, which metric partially catches the problem and why only partially, and redefine coverage so that a flagged gap and a fabricated finding score differently.
- Report B's fabrication had to originate somewhere in the pipeline. List the stages where it could have entered, the per-stage signal that would distinguish them, and what the 0.70 citation-precision score tells you about the citation agent's behavior on this run.
Share this question