Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Free

Distinguish Genuine Drift From a Pipeline Bug Before Retraining

The churn model's prediction-distribution panel shows mean score drifting from 0.041 to 0.058 over two weeks. No deploys occurred in that window. Per-feature PSI shows one feature, days_since_last_login, at 0.42 — far above every other feature.

  1. Walk through the specific checks you'd run, in order, to determine whether this is genuine population drift or a pipeline bug.
  2. Your checks reveal that a new mobile app version stopped reporting last_login events for a specific device OS, defaulting the feature to 0 for those users. Which case is this, and what should happen next?
  3. Explain concretely why retraining immediately, before completing this diagnosis, would have been the wrong move.
Solution

1. Ordered checks

Deploy markers first (already ruled out — no releases in the window, so it isn't an obvious release-caused break). Data-quality counters next: null rate looks normal, but the default-value rate for days_since_last_login — distinct from the null rate — is the check that actually matters here, since a feature defaulting to 0 is "not null" but is not real either. Per-feature PSI (already given, and concentrated in one feature) points at exactly where to look next: slice the anomaly by any available dimension (device OS, app version, signup cohort) to see if it's concentrated or diffuse.

2. Diagnosis and next step

This is a pipeline bug, not genuine drift — the world didn't change, an app version stopped emitting an event, so the feature pipeline is silently substituting a default value for an affected slice of users. The correct response is to fix the upstream event pipeline (or the feature job's handling of missing events), backfill the affected window's feature values once available, and only then consider whether any retraining is warranted — not to retrain on data containing the bug.

3. Why retraining first would be wrong

Retraining on data where a real feature has been silently replaced by a default value for an affected slice teaches the model that days_since_last_login = 0 (or whatever the default is) means something it doesn't; the offline evaluation set has the same bug, so it would pass its own gate and look like a legitimate improvement or a neutral retrain, baking the corrupted signal into the new "official" champion. This is the exact anti-pattern this case study names explicitly: a drift-triggered retrain that skips full diagnosis can turn a two-day pipeline incident into a permanently degraded model.

Share this question

← Back to Case Study: Churn Model, End-to-End MLOps practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.