Case Study: Synthetic Training Data for Computer Vision
Design the internal data engine that fixes a detector's rare-case failures — targeted generation triggered by measured production failures, and why the realism of the synthetic images is never the metric that matters
Model interview answer for synthetic training-data generation (autonomous driving, retail shelf detection, medical imaging-class systems): why this case is the track's internal flywheel pattern — generation as a data engine feeding a detector's own retraining loop, not a user-facing product, with recall lift per dollar of compute as the ROI story instead of engagement or conversion. The approach ladder from classical augmentation through simulation/rendering (Omniverse Replicator / Unity Perception-class domain randomization), CycleGAN-style unpaired sim-to-real translation, conditional diffusion augmentation, and fully layout-conditioned generation, argued against the GAN-vs-diffusion comparison table from this track's GAN foundation subject. The failure-slice mining loop, label preservation under editing, dedup/leakage checks against the test set, and the real-data-anchor filter that guards against model collapse from training on your own outputs. Train-on-synthetic-test-on-real as the only metric that matters, with FID demoted to a diagnostic; a telemetry-to-promotion pipeline as the system design; and a back-of-envelope ROI argument comparing recall lift per dollar of compute against the cost of collecting and labelling the same failure slice in the real world.
Practice questions (6)
-
View →
Matching the Approach Ladder to Three Failure Slices
Advanced · Free -
View →
When FID Improves but Recall Doesn't
Advanced · Free -
View →
Designing the Real-Data-Anchor Filter Against Model Collapse
Advanced -
View →
Choosing a Mixing Ratio and a Stopping Point
Advanced -
View →
Investigating a Recall Number That's Too Good
Advanced