Advanced
Open
Pro
Why a Great FID Score Didn't Predict the Complaint Spike
Part of the AI Engineer Interview path →
Part of the Generative Vision & Image AI System Design path →
Your team ships a new inpainting model version. Offline, it posts a better FID and LPIPS score (measured on the masked region only, against held-out ground truth) than the previous version. A week after rollout, user complaints spike: "the fill looks obviously fake" and "you can see where the edit is," even though users generally agree the content of the fill (what got generated) looks plausible in isolation.
- Explain how a model can improve on region-level FID/LPIPS while the product-level complaint (visible editing) gets worse. What is the evaluation gap?
- Propose a concrete offline metric that would have caught this before rollout, and describe how you'd compute it.
- Name one online metric from this case's evaluation suite that should also have moved in response to this regression, and explain the causal link between the two.
Share this question