Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — Model Evaluation Metrics (7 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Intermediate Open Free

Reading a Confusion Matrix Under Imbalance Permalink →

A defect-detection model is evaluated on 20,000 manufactured parts, of which 200 are defective (1% prevalence). At the default 0.5 threshold:

                  Pred defective   Pred OK
Actual defective       120            80
Actual OK              280         19,520
  1. Compute accuracy, precision, recall, specificity, F1 and MCC. Compare accuracy to the "predict OK for everything" baseline.
  2. A stakeholder says "99% accuracy — ship it". Explain, using the numbers, why that statement is misleading in both directions (it overstates and understates the model).
  3. Which single metric would you put on the dashboard for this model and why?

Share this question

Intermediate Open Pro

ROC-AUC vs PR-AUC for a Rare-Event Model

Unlock this question →
Intermediate Open Pro

Choosing a Threshold From a Cost Matrix

Unlock this question →
Advanced Open Pro

Diagnosing Calibration From a Reliability Table

Unlock this question →
Advanced Open Pro

Comparing Two Rankers With NDCG and MRR

Unlock this question →
Intermediate Open Pro

Picking a Regression Metric for Demand Forecasting

Unlock this question →
Intermediate Open Free

The Precision Hiding Behind 99% Accuracy Permalink →

A vendor pitches a fraud-detection model: "99% accurate." Your transaction stream is 0.5% fraud.

What's the lowest the model's precision could be, given only that accuracy number?

Construct the worst case explicitly, and name the metrics you'd demand instead.

Share this question

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.