Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Recall Lift per Dollar Versus Real-World Collection

A retail detector's "new SKU on shelf" slice currently misses the SKU 44% of the time (56% recall). You have two options to close the gap:

Option A — synthetic generation. Render 8,000 images of the new SKU using the manufacturer's 3D packaging model with domain randomization (rung 2), at an estimated $0.15/image generation cost, plus an estimated $1,400 of incremental retrain compute. Based on a pilot mixing-ratio ablation, this is expected to lift slice recall to approximately 85%.

Option B — real-world collection. Dispatch a data-collection team to photograph the SKU on shelves across a sample of stores over three weeks, then send the images to a labelling vendor at $3.50/image for bounding boxes. The team estimates they can collect and label 1,200 real images in that window, at a fully-loaded field-collection cost (staff time, travel) of approximately $9,000 on top of labelling, and based on similar past efforts expects this to lift slice recall to approximately 78% (real images are higher-fidelity per example, but far fewer of them and slower to arrive).

  1. Compute the total cost and the recall-lift-per-dollar for each option. State any assumptions you make explicitly.
  2. Beyond the recall-lift-per-dollar number, name two other factors that should influence the decision, and explain which option they favor.
  3. Is choosing purely on recall-lift-per-dollar always correct for this kind of decision? Give a scenario where it wouldn't be.

Share this question

← Back to Case Study: Synthetic Training Data for Computer Vision practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.