Practice — Case Study: Synthetic Training Data for Computer Vision (6 questions)
Matching the Approach Ladder to Three Failure Slices Permalink →
You run the synthetic-data engine for three different production detectors, and each has just had a new failure slice mined from telemetry:
- An autonomous-driving perception stack misses pedestrians who are more than 50% occluded by a parked delivery van, at night, in light rain. A driving simulator with a street-scene asset library is already available to the team.
- A retail shelf-detection system fails to recognize a brand-new SKU that launched last week — there are a handful of real shelf photos containing it, and a 3D product-packaging model from the manufacturer.
- A medical-imaging classifier under-detects a rare pathology subtype; there are ~40 confirmed real scans showing it, and no anatomical simulator of any kind.
For each slice, pick the approach-ladder rung (or rungs) you would reach for first, and justify the choice against the specific characteristics of that slice — don't just name a rung, explain why the alternative rungs are a worse fit here.
Share this question
When FID Improves but Recall Doesn't Permalink →
Your team generates two candidate synthetic batches for the same "occluded pedestrian at night" slice using two different diffusion configurations. An automated report shows:
- Batch A: FID = 18.2 (against a real reference sample of the slice)
- Batch B: FID = 31.6
Batch A looks more photorealistic in every way a human reviewer checks. After retraining the detector on real data plus each batch separately and evaluating on a real, held-out sample of the slice:
- Detector retrained with Batch A: slice recall +1.5 points over baseline
- Detector retrained with Batch B: slice recall +9.0 points over baseline
- Explain, mechanistically, how this outcome is possible — what could Batch B be doing that Batch A isn't, despite scoring worse on FID and looking less realistic to a human?
- What decision do you make, and what do you do before trusting Batch B's number?
- A teammate proposes optimizing the generation pipeline's hyperparameters directly against FID going forward, to make this kind of discrepancy less likely. Is that a good idea? Why or why not?
Share this question