A Model That Improved PSNR and Made Users Angrier
Your team ships a new checkpoint of the on-device upscaling model. Offline, it improves PSNR by 0.8 dB and SSIM by 0.02 over the previous checkpoint — a clear win by the metrics the team has always tracked. After rollout, user complaints about "blurry" or "fake-looking" zoomed photos rise, and the regeneration/re-crop rate in the cloud restoration product (a related but separate model family) also ticks up over the same period.
- Explain, using the perception-distortion trade-off, how a model can improve PSNR/SSIM while making users less happy with the result.
- What evaluation should have caught this before rollout, and why didn't PSNR/SSIM catch it?
- Propose a decision rule for what "improved" should mean for this product going forward, and justify it.
1. How PSNR can improve while perceived quality drops
PSNR and SSIM measure closeness to one specific ground-truth image in pixel space. When the true high-resolution content is genuinely ambiguous given the low-resolution input (a texture, a face's fine detail), the pixel-distance-minimizing prediction is an average over plausible outputs — and an average of several plausible textures reads as blur, not detail. If the new checkpoint was tuned or trained with more weight on the pixel/distortion loss relative to the adversarial or perceptual loss, it would very plausibly post a better PSNR/SSIM number while producing softer, more "hedged" — and to a human, more obviously fake-looking or unsatisfying — output. This is exactly the perception-distortion trade-off: distortion metrics and perceptual quality are in tension near the achievable frontier, so an improvement in one is not evidence of an improvement in the other, and can coincide with the other getting worse.
2. What should have caught this, and why PSNR/SSIM didn't
LPIPS (a learned perceptual distance trained to correlate with human judgments) and NIQE (a no-reference naturalness score) plus a human MOS panel comparing the new checkpoint against the previous one pairwise should have been run before rollout — these are the metrics that track what a human notices, which PSNR/SSIM structurally cannot, since PSNR/SSIM only measure closeness to one ground truth in pixel space and have no mechanism for penalizing "technically closer to ground truth on average, but reads as blurry and unconvincing to a human viewer with no access to that ground truth." The offline gate that only checked PSNR/SSIM passed a model that was, by the metric that actually matters to users, a regression.
3. A decision rule going forward
Treat LPIPS/NIQE and human MOS as the metrics that gate a launch/rollout decision, and PSNR/SSIM as a secondary sanity check that the model hasn't drifted arbitrarily far from the input content (not as a metric that alone justifies shipping a change). Concretely: a candidate checkpoint should not ship unless it is neutral-or-better on LPIPS/NIQE and human pairwise preference against the current production model, regardless of what happens to PSNR/SSIM in either direction — and any change that improves PSNR/SSIM while degrading LPIPS/NIQE or human preference should be treated as a regression requiring investigation (most likely an underweighted adversarial/perceptual loss term), not a release candidate. This directly matches what actually happened here: the online signal (rising complaints, rising regeneration rate) is the real-world confirmation that the offline gate used the wrong headline metric.
Share this question