The Precision Hiding Behind 99% Accuracy
A vendor pitches a fraud-detection model: "99% accurate." Your transaction stream is 0.5% fraud.
What's the lowest the model's precision could be, given only that accuracy number?
Construct the worst case explicitly, and name the metrics you'd demand instead.
D) It could be exactly 0% — 99% accuracy is compatible with a model that has never caught a single fraud.
Worst-case construction, on 1M transactions (5,000 fraud, 995,000 legit): suppose the model flags 5,000 legitimate transactions as fraud and misses all 5,000 real frauds. Errors = 5,000 false positives + 5,000 false negatives = 10,000 = 1% of traffic — so accuracy is 99%, while precision is 0/5,000 = 0% and recall is also 0%.
The base rate is what makes this possible: when 99.5% of labels are "legit," a model earns 99%+ accuracy almost entirely by saying "legit" — the never-flag-anything model scores 99.5% while doing literally nothing. Accuracy above the base rate is table stakes, not evidence.
The rule of thumb to say out loud: under class imbalance, compare accuracy to the majority-class baseline first — if they're close, the number tells you nothing about the minority class.
What to demand instead:
- Precision and recall on the fraud class (or the full confusion matrix at the operating threshold).
- PR-AUC rather than ROC-AUC — precision-recall curves stay honest under heavy imbalance.
- Ideally, expected cost: dollars saved per flagged transaction versus review cost — the metric the business actually optimizes.
Share this question