Designing the Real-Data-Anchor Filter Against Model Collapse
Your synthetic-data engine has been running for a year across many slices. A retrospective analysis finds that a specific slice's synthetic pool is now composed as follows: cycle 1 was generated from real images (100% real-anchored); each subsequent quarterly cycle generated new synthetic examples using, in part, the previous cycle's synthetic output as reference/style examples to speed up generation, because doing so produced batches that passed the validation gate faster. By cycle 5, roughly 60% of the slice's active training pool traces back through at least one prior synthetic generation rather than directly to a real image.
- Explain why "passed the validation gate faster" is itself a warning sign here, not evidence the pipeline is healthy.
- Describe concretely how you would detect that this slice is at risk of model collapse, using signals available in this pipeline.
- Propose a specific, mechanical policy change (not just "be more careful") that would have prevented this drift from cycle 1 to cycle 5, and explain the trade-off it introduces.
Share this question