Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Free

Matching the Approach Ladder to Three Failure Slices

You run the synthetic-data engine for three different production detectors, and each has just had a new failure slice mined from telemetry:

  1. An autonomous-driving perception stack misses pedestrians who are more than 50% occluded by a parked delivery van, at night, in light rain. A driving simulator with a street-scene asset library is already available to the team.
  2. A retail shelf-detection system fails to recognize a brand-new SKU that launched last week — there are a handful of real shelf photos containing it, and a 3D product-packaging model from the manufacturer.
  3. A medical-imaging classifier under-detects a rare pathology subtype; there are ~40 confirmed real scans showing it, and no anatomical simulator of any kind.

For each slice, pick the approach-ladder rung (or rungs) you would reach for first, and justify the choice against the specific characteristics of that slice — don't just name a rung, explain why the alternative rungs are a worse fit here.

Solution

1. Night-rain occluded pedestrian — rung 4 (conditional diffusion augmentation), with rung 2/3 as a secondary source. The core problem here is a scene-condition change (daytime/clear → night/rain) plus an occlusion geometry change, applied to a subject (a pedestrian) the real world already has abundant well-captured examples of in daytime, unoccluded form. Starting from a real captured daytime image and editing it — darkening and adding rain via mask/layout-conditioned diffusion, and compositing a synthetic occluder over part of the pedestrian's bounding box with the label updated to reflect the new visible fraction — keeps everything except the edited attributes anchored to a real photo, which is exactly what minimizes domain gap for this kind of slice. Because a simulator is also available, rung 2/3 (render occluded night-rain scenes directly, optionally translated toward the real sensor's look) is a reasonable secondary source to increase volume and geometric diversity beyond what editing real images alone can provide — but it shouldn't be the primary rung, because it discards the "start from something already real" advantage that makes rung 4 the cheaper, lower-domain-gap choice when a real base image exists. Rung 5 (fully generated from a layout) is unnecessary here — nothing about this slice requires generating a pedestrian or a street scene from nothing; the object and the base scene both already exist in reality.

2. Brand-new SKU on the shelf — rung 2 (simulation/rendering) as primary, rung 5 as a fallback. This slice is fundamentally different: the object itself — the new SKU's packaging — does not exist in any prior real photo in enough volume to edit from, so rung 4's "start from a real base image" advantage isn't available for the object itself (only for the shelf background). But a 3D packaging model is available, which is exactly the input rung 2 needs: render the new product, using domain randomization over shelf position, lighting, camera angle and neighboring products, to get perfect boxes/masks as a byproduct of rendering with no extra labelling step. If the rendered product's appearance (label print, packaging material) needs to look more like real photos of that packaging under store lighting, a lightweight sim-to-real translation (rung 3) trained on the retailer's existing shelf-photo domain narrows that gap further. Rung 5 (layout-conditioned generation with no 3D asset) is the fallback only if no packaging model exists — here it's strictly worse than rung 2 because rung 2's 3D asset gives exact, controllable geometry a purely generative layout model would have to approximate.

3. Rare pathology subtype — rung 4 (conditional diffusion editing of real scans), used conservatively. With no anatomical simulator available, rungs 2/3/5 are effectively off the table or would require building simulation capability from scratch — a much larger investment than this single slice likely justifies. The ~40 confirmed real scans are the valuable asset: use them, and a larger pool of real scans without the pathology, as the basis for mask-conditioned edits that introduce or exaggerate the pathology's visual characteristics onto real anatomy, keeping everything outside the lesion region anchored to a genuine scan. This should be treated as the most conservative rung choice of the three — medical imaging has the smallest acceptable "hallucination budget," so every edited scan needs heavier human clinical review before entering training than either of the other two slices would need, and the real-data-anchor filter should be tuned tighter (a lower tolerance for "looks almost real") than for the driving or retail slices, precisely because an anatomically implausible synthetic lesion is a more direct risk than a slightly-off rain texture or shelf lighting.

Share this question

← Back to Case Study: Synthetic Training Data for Computer Vision practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.