Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Why a Great FID Score Didn't Predict the Complaint Spike

Your team ships a new inpainting model version. Offline, it posts a better FID and LPIPS score (measured on the masked region only, against held-out ground truth) than the previous version. A week after rollout, user complaints spike: "the fill looks obviously fake" and "you can see where the edit is," even though users generally agree the content of the fill (what got generated) looks plausible in isolation.

  1. Explain how a model can improve on region-level FID/LPIPS while the product-level complaint (visible editing) gets worse. What is the evaluation gap?
  2. Propose a concrete offline metric that would have caught this before rollout, and describe how you'd compute it.
  3. Name one online metric from this case's evaluation suite that should also have moved in response to this regression, and explain the causal link between the two.

Share this question

← Back to Case Study: Generative Fill — Inpainting, Outpainting and Object Removal practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.