Advanced
Open
Pro
An Automated Retrain That Skipped Approval
A payments-fraud model retrains nightly and auto-promotes to production whenever the challenger's offline precision/recall beats the current champion's on a held-out day. Last week, a nightly retrain auto-promoted a challenger that happened to clear the metric bar but had, unnoticed, been trained on a day where a labeling pipeline bug mislabeled a batch of legitimate transactions as fraud. The bad model ran in production for 36 hours, auto-declining a spike of legitimate transactions, before a business-metrics alert caught it.
- Diagnose exactly which stage-transition gate was missing or insufficient, and explain why "the challenger beat the champion on the metric" was not sufficient evidence to promote.
- Redesign the transition from Train → Validate → Approve → Deploy for this pipeline so this failure mode is caught, while preserving nightly retraining as the normal case (i.e., don't just require a human to approve every single night).
- What would need to be true about the data stage, upstream of training, to prevent this class of bug rather than just catching it later?
Share this question