Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Free

Provenance and Detection Are Not the Same Layer

A product manager proposes: "Now that SynthID-class watermarking exists, can we just require all AI-generated content to be watermarked and skip building a learned detector entirely? It would be much cheaper."

  1. Explain why this proposal does not cover the platform's actual upload stream, using the provenance-vs-detection distinction.
  2. Separately, a security researcher points out that C2PA credentials can be stripped by re-saving an image. Does this mean C2PA is not worth implementing? Explain what it is still good for.
  3. Design the minimal combination of provenance and detection that would be defensible for a platform launching this system today.
Solution

1. Why watermarking alone doesn't cover the stream: SynthID-class watermarking only works if the generator that produced the content chose to embed it — a compliant, cooperative generator. It says nothing about content from a non-compliant generator (an open-source model run without any watermarking step), content that was manipulated rather than fully generated (a face-swap onto a real background frame, which a generation-time watermark on the swapped region wouldn't necessarily carry through), or a bad actor who deliberately picks tools that don't watermark specifically to evade this check. Today's actual upload stream is overwhelmingly content with no verifiable origin claim either way — no credential, no watermark — precisely because watermark-and-credential-emitting tools are not yet universal. A learned detector exists for exactly that unlabelled majority; skipping it means the system only ever catches the minority of fakes made carelessly enough to use a compliant, watermarking generator.

2. C2PA's value despite being strippable: A stripped credential doesn't mean C2PA has no value; it means C2PA answers "is there a verifiable claim this asset is real/generated, unaltered since capture" rather than "is this asset fake." An asset with an intact, valid C2PA credential is a strong, cheap, deterministic signal it can be trusted without running any learned model. An asset without one (whether never signed, or stripped) simply falls through to the next layer — it is not treated as suspicious purely for lacking a credential, since most content today has none for entirely benign reasons (older capture devices, tools that don't yet implement the standard). The strippability point argues for defence-in-depth, not for skipping C2PA: it is one deterministic, high-precision layer among several, not a standalone proof of authenticity.

3. A minimal defensible combination: At minimum: (a) verify and honour C2PA credentials where present, to cheaply clear or label the subset of traffic that carries one; (b) extract SynthID-class or equivalent watermarks where present, for the subset of generated content that carries one even without a full credential chain; (c) a perceptual-hash lookup against known-bad assets, to catch repeat-uploads of previously confirmed fakes for near-free; (d) a learned detector — even a single frequency-aware image-level model to start — covering everything that survives (a)–(c), since that is most of the stream on day one; (e) a human review path for the detector's ambiguous middle. Launching with only provenance, or only a learned detector, each leaves a large, predictable gap; the minimal defensible system needs at least one deterministic layer and at least one learned layer, because each covers population the other structurally cannot.

Share this question

← Back to Case Study: Detecting Deepfakes and AI-Generated Media at Platform Scale practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.