Advanced
Open
Pro
A Suspiciously Good FID Score
Part of the ML System Design Interview path →
Part of the Generative Vision & Image AI System Design path →
A new checkpoint of your face generator posts the best FID score the team has ever recorded, and the promotion review is ready to ship it as the new production model. Before signing off, you look at a grid of 50 freshly sampled faces and notice they look strikingly similar to each other — similar pose, similar lighting, similar general appearance — even though each individual face looks highly realistic.
- Explain how a generator in this state could still post an excellent FID score, referencing what FID actually measures.
- What additional offline metric would you compute before approving promotion, and what result would confirm your suspicion?
- Propose one additional evaluation to run given this specific suspicion, beyond what's already in the standard evaluation checklist, and justify it.
Share this question