When FID Improves but Recall Doesn't
Your team generates two candidate synthetic batches for the same "occluded pedestrian at night" slice using two different diffusion configurations. An automated report shows:
- Batch A: FID = 18.2 (against a real reference sample of the slice)
- Batch B: FID = 31.6
Batch A looks more photorealistic in every way a human reviewer checks. After retraining the detector on real data plus each batch separately and evaluating on a real, held-out sample of the slice:
- Detector retrained with Batch A: slice recall +1.5 points over baseline
- Detector retrained with Batch B: slice recall +9.0 points over baseline
- Explain, mechanistically, how this outcome is possible — what could Batch B be doing that Batch A isn't, despite scoring worse on FID and looking less realistic to a human?
- What decision do you make, and what do you do before trusting Batch B's number?
- A teammate proposes optimizing the generation pipeline's hyperparameters directly against FID going forward, to make this kind of discrepancy less likely. Is that a good idea? Why or why not?
1. How this is mechanistically possible: FID measures how close the overall feature distribution of a generated batch is to a real reference sample — it says nothing about whether the specific pixels that matter to the detector (occlusion boundaries, the precise silhouette of a partially visible pedestrian against a dark, wet background) are well represented. Batch A could be photorealistic in a generic sense — good lighting, good texture, low artifact rate overall — while systematically under-representing or softening exactly the occlusion boundary detail the detector needs to learn from (e.g. a diffusion configuration that produces smoother, more "plausible-looking" but less geometrically precise edges around the occluder). Batch B could look visibly rougher or more artifact-prone in general (explaining its worse FID and worse human-realism rating) while happening to preserve sharper, more varied occlusion-boundary geometry that is the actual signal the detector is missing. FID averages over the whole image and the whole batch distribution; it has no mechanism to weight the specific sub-region and specific geometric property that determines whether this detector's recall on this exact failure mode improves.
2. The decision, and what to check first: Ship Batch B's result as the one that matters, because slice recall on real held-out data is the metric defined as the success criterion — FID was always a diagnostic, not the grading criterion, and this is a direct, concrete illustration of exactly why. Before trusting the +9.0 number, though: (a) run the leakage/dedup check between Batch B's source images and the real held-out test set specifically — a surprisingly large recall jump on a small slice is exactly the signature a leaked near-duplicate source image would produce, and it needs to be ruled out before celebrating; (b) check the overall-distribution recall/precision for both retrains, not just the slice number, to confirm Batch B's larger, rougher-looking batch isn't buying its slice gain at the cost of a regression elsewhere; (c) run the real-data-anchor filter on Batch B specifically, since a batch that looks rougher to a human might also be further from the anchor distribution in a way that matters for model-collapse risk even if it helps this one retrain.
3. Optimizing generation hyperparameters directly against FID — not a good idea: This would optimize the pipeline against exactly the metric this case demonstrates is disconnected from the metric that actually matters, and could plausibly make results worse: a generation configuration tuned to minimize FID is being tuned to match the overall feature distribution of a generic real reference sample, which has no reason to correlate with — and, as this scenario shows, can actively trade off against — preserving the specific geometric detail a downstream detector needs on a specific slice. The correct optimization target is TSTR slice recall itself (via the mixing-ratio-style ablation this subject describes, run per generation configuration), with FID kept only as a cheap pre-flight sanity check for catching a badly broken generation run before spending retrain compute on it — not promoted into the thing hyperparameters are tuned against.
Share this question