PSI Says Nothing Changed. The Model Is Dead Wrong.
A pricing model's monitoring dashboard has looked clean for months: PSI on every one of its 40 input features sits around 0.01, deep in the "no significant change" band, and no data-quality alert has ever fired. Meanwhile, conversion rate and revenue per session have quietly declined about 15% over the same period. A manual audit pulls a sample of recent transactions and finds the model's predicted optimal price no longer matches what actually maximizes revenue for the same customer/product feature values it saw a year ago — the inputs look the same, but the right answer for those inputs has changed.
Given that feature-level PSI has stayed flat the entire time, what is actually happening, and why did feature-drift monitoring alone fail to catch it?
Concept drift — P(y|X) shifted while P(X) stayed the same, invisible to feature PSI.
PSI, and feature-drift monitoring generally, measures whether the distribution of the inputs — P(X) — has changed relative to training: are customers, products, and their attributes still arriving in roughly the same mix as before? Here they are; PSI is correctly reporting that nothing about X has moved.
What's changed is the relationship between those inputs and the right answer — P(y|X), or in this case the true optimal price given a customer/product profile. Market conditions, competitor pricing, or customer price sensitivity can shift over months while the customers themselves (and their observable features) look statistically identical to a year ago. This is concept drift, and it is structurally invisible to any monitor that only looks at P(X): by construction, PSI (and any other input-distribution check) cannot distinguish "the world changed in a way that makes our model wrong" from "the world didn't change at all," because both produce the same flat feature distributions. A model can be completely, silently broken while every feature-drift dashboard stays green for months, exactly as happened here.
The fix isn't a better feature-drift metric — it's a monitoring layer that doesn't depend on X alone: performance metrics computed against ground truth once outcomes mature (here, realized revenue per price point, ideally compared against a small randomized exploration/holdout slice so you can measure whether the model's chosen price is actually the best one, not just self-consistent with its own past choices), or business-KPI monitoring (conversion, revenue) with alerting thresholds and correlation to deploy markers. Feature drift and concept drift are complementary signals — one catches "the inputs changed," the other catches "the world changed in a way the inputs don't reveal" — and a monitoring stack that relies on the first alone will always be blind to the second.
Share this question