Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — Case Study: Synthetic Training Data for Computer Vision (6 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Advanced Open Free

Matching the Approach Ladder to Three Failure Slices Permalink →

You run the synthetic-data engine for three different production detectors, and each has just had a new failure slice mined from telemetry:

  1. An autonomous-driving perception stack misses pedestrians who are more than 50% occluded by a parked delivery van, at night, in light rain. A driving simulator with a street-scene asset library is already available to the team.
  2. A retail shelf-detection system fails to recognize a brand-new SKU that launched last week — there are a handful of real shelf photos containing it, and a 3D product-packaging model from the manufacturer.
  3. A medical-imaging classifier under-detects a rare pathology subtype; there are ~40 confirmed real scans showing it, and no anatomical simulator of any kind.

For each slice, pick the approach-ladder rung (or rungs) you would reach for first, and justify the choice against the specific characteristics of that slice — don't just name a rung, explain why the alternative rungs are a worse fit here.

Share this question

Advanced Open Free

When FID Improves but Recall Doesn't Permalink →

Your team generates two candidate synthetic batches for the same "occluded pedestrian at night" slice using two different diffusion configurations. An automated report shows:

  • Batch A: FID = 18.2 (against a real reference sample of the slice)
  • Batch B: FID = 31.6

Batch A looks more photorealistic in every way a human reviewer checks. After retraining the detector on real data plus each batch separately and evaluating on a real, held-out sample of the slice:

  • Detector retrained with Batch A: slice recall +1.5 points over baseline
  • Detector retrained with Batch B: slice recall +9.0 points over baseline
  1. Explain, mechanistically, how this outcome is possible — what could Batch B be doing that Batch A isn't, despite scoring worse on FID and looking less realistic to a human?
  2. What decision do you make, and what do you do before trusting Batch B's number?
  3. A teammate proposes optimizing the generation pipeline's hyperparameters directly against FID going forward, to make this kind of discrepancy less likely. Is that a good idea? Why or why not?

Share this question

Advanced Open Pro

Designing the Real-Data-Anchor Filter Against Model Collapse

Unlock this question →
Advanced Open Pro

Choosing a Mixing Ratio and a Stopping Point

Unlock this question →
Advanced Open Pro

Investigating a Recall Number That's Too Good

Unlock this question →
Advanced Open Pro

Recall Lift per Dollar Versus Real-World Collection

Unlock this question →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.