Advanced
Open
Pro
Why Synthetic Defects Can Validate but Must Not Train
Part of the ML System Design Interview path →
Part of the MLOps & Production ML path →
Part of the Generative Vision & Image AI System Design path →
A junior teammate proposes: "We have almost no real defect images. Let's generate thousands of synthetic defective images with cut-paste and generative edits, label them 'defective,' and train a supervised binary classifier (normal vs. defective) directly on real normal images plus this synthetic defective set — it solves the data scarcity problem and lets us skip the one-class anomaly detection machinery entirely."
- What goes wrong with this proposal, concretely — what will the resulting classifier actually learn to detect?
- Contrast this with how this case actually uses synthetic defects, and explain why that use doesn't have the same failure mode.
- This case's synthetic-defect vocabulary is cross-referenced from
case-study-synthetic-data-generation-for-computer-vision. Which concept from that case applies almost directly here, and how?
Share this question