Advanced
Open
Pro
Investigating a Recall Number That's Too Good
A generation job for the "occluded pedestrian at night" slice produces a batch using rung 4 (conditional diffusion editing of real daytime images). After retraining, slice recall jumps from 41% to 88% — a much larger gain than any previous cycle for any slice in the program's history. The team is excited and wants to promote the candidate immediately.
- Before promoting, what specific checks would you run, and in what order, and why does the order matter?
- Describe two distinct, plausible root causes for a suspiciously large recall jump like this one — one relating to data leakage and one relating to the evaluation set itself — and how you'd tell them apart.
- Suppose the leakage check comes back clean, but you discover the real held-out evaluation sample for this slice has only 18 real examples. Does that change your confidence in the 88% number, and what would you recommend?
Share this question