Provenance and Detection Are Not the Same Layer
A product manager proposes: "Now that SynthID-class watermarking exists, can we just require all AI-generated content to be watermarked and skip building a learned detector entirely? It would be much cheaper."
- Explain why this proposal does not cover the platform's actual upload stream, using the provenance-vs-detection distinction.
- Separately, a security researcher points out that C2PA credentials can be stripped by re-saving an image. Does this mean C2PA is not worth implementing? Explain what it is still good for.
- Design the minimal combination of provenance and detection that would be defensible for a platform launching this system today.
1. Why watermarking alone doesn't cover the stream: SynthID-class watermarking only works if the generator that produced the content chose to embed it — a compliant, cooperative generator. It says nothing about content from a non-compliant generator (an open-source model run without any watermarking step), content that was manipulated rather than fully generated (a face-swap onto a real background frame, which a generation-time watermark on the swapped region wouldn't necessarily carry through), or a bad actor who deliberately picks tools that don't watermark specifically to evade this check. Today's actual upload stream is overwhelmingly content with no verifiable origin claim either way — no credential, no watermark — precisely because watermark-and-credential-emitting tools are not yet universal. A learned detector exists for exactly that unlabelled majority; skipping it means the system only ever catches the minority of fakes made carelessly enough to use a compliant, watermarking generator.
2. C2PA's value despite being strippable: A stripped credential doesn't mean C2PA has no value; it means C2PA answers "is there a verifiable claim this asset is real/generated, unaltered since capture" rather than "is this asset fake." An asset with an intact, valid C2PA credential is a strong, cheap, deterministic signal it can be trusted without running any learned model. An asset without one (whether never signed, or stripped) simply falls through to the next layer — it is not treated as suspicious purely for lacking a credential, since most content today has none for entirely benign reasons (older capture devices, tools that don't yet implement the standard). The strippability point argues for defence-in-depth, not for skipping C2PA: it is one deterministic, high-precision layer among several, not a standalone proof of authenticity.
3. A minimal defensible combination: At minimum: (a) verify and honour C2PA credentials where present, to cheaply clear or label the subset of traffic that carries one; (b) extract SynthID-class or equivalent watermarks where present, for the subset of generated content that carries one even without a full credential chain; (c) a perceptual-hash lookup against known-bad assets, to catch repeat-uploads of previously confirmed fakes for near-free; (d) a learned detector — even a single frequency-aware image-level model to start — covering everything that survives (a)–(c), since that is most of the stream on day one; (e) a human review path for the detector's ambiguous middle. Launching with only provenance, or only a learned detector, each leaves a large, predictable gap; the minimal defensible system needs at least one deterministic layer and at least one learned layer, because each covers population the other structurally cannot.
Share this question