Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Choosing a Mixing Ratio and a Stopping Point

You run the mixing-ratio ablation for the "new SKU on shelf" slice at four real:synthetic ratios, each retrain using the same real data plus an increasing amount of synthetic data for the slice, evaluated on a real held-out sample:

Ratio (real:synthetic) Slice recall Overall-distribution recall
100:0 (baseline) 41% 92.0%
90:10 63% 91.8%
75:25 74% 91.5%
50:50 79% 90.3%

Generating and validating each additional 15-percentage-point increment of synthetic data for this slice costs approximately $900 in generation and incremental retrain compute (treat this as a stated estimate, not a precise figure).

  1. Given a hard floor of "overall-distribution recall must not drop by more than 1.0 point from baseline," which ratio would you ship, and why?
  2. Compute the cost per point of slice recall gained at each step from 90:10 onward, and use it to argue for or against pushing to 50:50 even though it technically clears the floor.
  3. The slice owner argues for 50:50 because "79% recall is just better than 74%." What's the strongest argument against defaulting to the ratio with the single highest slice-recall number?

Share this question

← Back to Case Study: Synthetic Training Data for Computer Vision practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.