Practice — Case Study: Detecting Deepfakes and AI-Generated Media at Platform Scale (6 questions)
Advanced
Open
Free
Why In-Distribution AUROC Is the Wrong Headline Metric Permalink →
Your team reports a new detector's evaluation results: AUROC = 0.97 against a held-out test set built from the same five generator pipelines present in training data. Leadership wants to ship it as the new production detector based on this number.
- Explain concretely why 0.97 in-distribution AUROC alone does not tell you how this detector will perform in production, citing the relevant published finding.
- Propose the evaluation split you would insist on seeing before sign-off, and how you would construct it.
- Suppose the same detector scores AUROC = 0.79 on your proposed split. Should it still ship? What else would you want to know before deciding?
Share this question
Advanced
Open
Free
Provenance and Detection Are Not the Same Layer Permalink →
A product manager proposes: "Now that SynthID-class watermarking exists, can we just require all AI-generated content to be watermarked and skip building a learned detector entirely? It would be much cheaper."
- Explain why this proposal does not cover the platform's actual upload stream, using the provenance-vs-detection distinction.
- Separately, a security researcher points out that C2PA credentials can be stripped by re-saving an image. Does this mean C2PA is not worth implementing? Explain what it is still good for.
- Design the minimal combination of provenance and detection that would be defensible for a platform launching this system today.
Share this question
Advanced
Open
Pro