Practice — ML Monitoring, Drift & Retraining (6 questions)
Compute and Interpret PSI for a Drifting Feature Permalink →
A credit-risk model uses monthly_income bucketed into 4 quantile bins
defined on the training set (each holding 25% of training rows). This
week's production distribution across the same bins is:
| Bin | Expected (train) | Actual (this week) |
|---|---|---|
| 1 (lowest) | 0.25 | 0.40 |
| 2 | 0.25 | 0.30 |
| 3 | 0.25 | 0.20 |
| 4 (highest) | 0.25 | 0.10 |
- Compute the PSI for this feature. Show the per-bin contributions.
- Using the common rule-of-thumb thresholds, what does the value mean and what would you do?
- Your colleague argues that a KS test on the raw values would be
"more rigorous" and proposes alerting on
p < 0.05. What is the problem with that at production scale (millions of rows per day)?
Share this question
PSI Says Nothing Changed. The Model Is Dead Wrong. Permalink →
A pricing model's monitoring dashboard has looked clean for months: PSI on every one of its 40 input features sits around 0.01, deep in the "no significant change" band, and no data-quality alert has ever fired. Meanwhile, conversion rate and revenue per session have quietly declined about 15% over the same period. A manual audit pulls a sample of recent transactions and finds the model's predicted optimal price no longer matches what actually maximizes revenue for the same customer/product feature values it saw a year ago — the inputs look the same, but the right answer for those inputs has changed.
Given that feature-level PSI has stayed flat the entire time, what is actually happening, and why did feature-drift monitoring alone fail to catch it?
Share this question