Advanced
Open
Pro
A Restoration That Scores Well and Is Still Wrong
A cloud photo-restoration job on an old, damaged family photo scores well on every metric your team tracks: high LPIPS-based similarity to a reference restoration, a strong NIQE naturalness score, and a positive human MOS rating from a rater who was not shown the original damaged photo. The customer who uploaded the photo says the restored face "doesn't look like my grandmother" and requests a refund.
- Explain why every metric in the pipeline could look good here while the output is still a real failure.
- Design a concrete automatic check that should run on every job (not just in offline evaluation) to catch this class of failure before delivery, and explain what it measures that the existing metrics don't.
- A colleague proposes fixing this by simply lowering the GAN-loss weight globally for all cloud restoration jobs. Evaluate this fix, including any real cost.
Share this question