Advanced
Open
Pro
A Great FID Score, a Warped Logo in Production
Part of the ML System Design Interview path →
Part of the Generative Vision & Image AI System Design path →
Your team ships a new try-on model version. Offline, it posts a better FID and LPIPS score than the previous version, measured on held-out paired test data. A week after rollout, customer support tickets spike: shoppers report that brand logos and printed slogans look "smeared" or "wrong" in try-on previews, even though they generally agree the overall photo looks realistic.
- Explain how a model can improve on FID/LPIPS while this specific complaint gets worse. What is the evaluation gap?
- Propose a concrete offline metric (or pair of metrics) that would have caught this before rollout, and describe exactly how you'd compute it.
- Name one online, non-support-ticket signal that should also have moved in response to this regression, and explain the causal chain from root cause to that signal.
Share this question