Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

A Great FID Score, a Warped Logo in Production

Your team ships a new try-on model version. Offline, it posts a better FID and LPIPS score than the previous version, measured on held-out paired test data. A week after rollout, customer support tickets spike: shoppers report that brand logos and printed slogans look "smeared" or "wrong" in try-on previews, even though they generally agree the overall photo looks realistic.

  1. Explain how a model can improve on FID/LPIPS while this specific complaint gets worse. What is the evaluation gap?
  2. Propose a concrete offline metric (or pair of metrics) that would have caught this before rollout, and describe exactly how you'd compute it.
  3. Name one online, non-support-ticket signal that should also have moved in response to this regression, and explain the causal chain from root cause to that signal.

Share this question

← Back to Case Study: Virtual Try-On for Fashion E-commerce practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.