Advanced
Open
Pro
Why Pixel-Level AUROC Alone Can Mislead
Part of the ML System Design Interview path →
Part of the MLOps & Production ML path →
Part of the Generative Vision & Image AI System Design path →
Two candidate models for the same SKU report identical pixel-level AUROC of 0.97 on your evaluation set. Digging into per-defect performance:
- Model X localizes large dents (covering hundreds of pixels) very precisely, but almost completely misses hairline scratches (covering a handful of pixels) — its anomaly map barely lights up on them.
- Model Y localizes both large dents and hairline scratches reasonably well, with moderately less precise boundaries on the large dents than Model X.
- Explain, mechanically, how Model X can post the same pixel-level AUROC as Model Y despite missing an entire defect category almost completely.
- What would the PRO (per-region overlap) metric show for each model, and why does it diverge from the pixel-AUROC picture?
- Which model would you actually ship, and what does this imply about reporting pixel-AUROC without PRO alongside it in this case?
Share this question