Why In-Distribution AUROC Is the Wrong Headline Metric
Your team reports a new detector's evaluation results: AUROC = 0.97 against a held-out test set built from the same five generator pipelines present in training data. Leadership wants to ship it as the new production detector based on this number.
- Explain concretely why 0.97 in-distribution AUROC alone does not tell you how this detector will perform in production, citing the relevant published finding.
- Propose the evaluation split you would insist on seeing before sign-off, and how you would construct it.
- Suppose the same detector scores AUROC = 0.79 on your proposed split. Should it still ship? What else would you want to know before deciding?
1. Why 0.97 alone is misleading: The DeepFake Detection Challenge (DFDC), run by Meta AI with a public leaderboard, published exactly this failure mode: models that scored strongly against generator methods present during development saw their performance fall sharply against a held-out black-box test set built from generator methods the competitors had never seen. A detector's AUROC against generators it was trained and evaluated on tells you it has learned to recognise those specific artefacts — frequency-domain signatures, blending patterns, whatever those particular pipelines leave behind — not that it has learned a generalisable notion of "synthetic." Since the whole point of this system is to keep working as generators improve, and a new generator family can appear within weeks, in-distribution AUROC measures the wrong thing: how well the detector has memorised the past, not how it will do against what's coming.
2. The split to insist on: An unseen-generator split: a held-out evaluation set built entirely from one or more generator families that were excluded from training entirely — ideally a generator added to the in-house generator zoo after the training data was frozen, so there is no possibility of leakage. This requires deliberately holding back at least one generator family from training specifically to serve as this test, and refreshing it as new families are added to the zoo, so the test itself doesn't go stale the same way a static public benchmark does.
3. Should it ship at 0.79 unseen-generator AUROC? 0.79 is meaningfully better than chance and still useful as one signal inside the cascade (feeding an ensemble, not standing alone as a verdict), but it should not be evaluated in isolation. Before deciding: check precision at the actual review-capacity operating point (not just aggregate AUROC), check calibration (is 0.79 AUROC paired with a trustworthy probability the policy engine can threshold on), check per-demographic error parity, and confirm the cascade's other layers (provenance, hashing, human review of the ambiguous middle) are absorbing the gap this number implies. A cascade is built assuming no single learned layer is perfect — the question isn't "is 0.79 good enough alone" but "does the full system's expected performance, with this detector as one layer, meet the bar," which is a different and more complete question than the leadership team's original framing.
Share this question