Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Free

When FID Improves but Recall Doesn't

Your team generates two candidate synthetic batches for the same "occluded pedestrian at night" slice using two different diffusion configurations. An automated report shows:

  • Batch A: FID = 18.2 (against a real reference sample of the slice)
  • Batch B: FID = 31.6

Batch A looks more photorealistic in every way a human reviewer checks. After retraining the detector on real data plus each batch separately and evaluating on a real, held-out sample of the slice:

  • Detector retrained with Batch A: slice recall +1.5 points over baseline
  • Detector retrained with Batch B: slice recall +9.0 points over baseline
  1. Explain, mechanistically, how this outcome is possible — what could Batch B be doing that Batch A isn't, despite scoring worse on FID and looking less realistic to a human?
  2. What decision do you make, and what do you do before trusting Batch B's number?
  3. A teammate proposes optimizing the generation pipeline's hyperparameters directly against FID going forward, to make this kind of discrepancy less likely. Is that a good idea? Why or why not?
Solution

1. How this is mechanistically possible: FID measures how close the overall feature distribution of a generated batch is to a real reference sample — it says nothing about whether the specific pixels that matter to the detector (occlusion boundaries, the precise silhouette of a partially visible pedestrian against a dark, wet background) are well represented. Batch A could be photorealistic in a generic sense — good lighting, good texture, low artifact rate overall — while systematically under-representing or softening exactly the occlusion boundary detail the detector needs to learn from (e.g. a diffusion configuration that produces smoother, more "plausible-looking" but less geometrically precise edges around the occluder). Batch B could look visibly rougher or more artifact-prone in general (explaining its worse FID and worse human-realism rating) while happening to preserve sharper, more varied occlusion-boundary geometry that is the actual signal the detector is missing. FID averages over the whole image and the whole batch distribution; it has no mechanism to weight the specific sub-region and specific geometric property that determines whether this detector's recall on this exact failure mode improves.

2. The decision, and what to check first: Ship Batch B's result as the one that matters, because slice recall on real held-out data is the metric defined as the success criterion — FID was always a diagnostic, not the grading criterion, and this is a direct, concrete illustration of exactly why. Before trusting the +9.0 number, though: (a) run the leakage/dedup check between Batch B's source images and the real held-out test set specifically — a surprisingly large recall jump on a small slice is exactly the signature a leaked near-duplicate source image would produce, and it needs to be ruled out before celebrating; (b) check the overall-distribution recall/precision for both retrains, not just the slice number, to confirm Batch B's larger, rougher-looking batch isn't buying its slice gain at the cost of a regression elsewhere; (c) run the real-data-anchor filter on Batch B specifically, since a batch that looks rougher to a human might also be further from the anchor distribution in a way that matters for model-collapse risk even if it helps this one retrain.

3. Optimizing generation hyperparameters directly against FID — not a good idea: This would optimize the pipeline against exactly the metric this case demonstrates is disconnected from the metric that actually matters, and could plausibly make results worse: a generation configuration tuned to minimize FID is being tuned to match the overall feature distribution of a generic real reference sample, which has no reason to correlate with — and, as this scenario shows, can actively trade off against — preserving the specific geometric detail a downstream detector needs on a specific slice. The correct optimization target is TSTR slice recall itself (via the mixing-ratio-style ablation this subject describes, run per generation configuration), with FID kept only as a cheap pre-flight sanity check for catching a badly broken generation run before spending retrain compute on it — not promoted into the thing hyperparameters are tuned against.

Share this question

← Back to Case Study: Synthetic Training Data for Computer Vision practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.