Advanced
Open
Pro
Citation Existence vs. Citation Faithfulness
Your offline eval dashboard shows citation-existence at 0.995 and faithfulness (entailment) at 0.81, against gates of 0.99 and 0.90 respectively.
- Explain what each metric is actually measuring, and why a system can score near-perfectly on one while failing the other.
- Give a concrete example answer/citation pair that would pass the existence check but fail the faithfulness check.
- The faithfulness score is below gate. Name two plausible root causes in the pipeline (not "the model is bad") and how you'd distinguish between them using the trace.
Share this question