Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

An Automated Retrain That Skipped Approval

A payments-fraud model retrains nightly and auto-promotes to production whenever the challenger's offline precision/recall beats the current champion's on a held-out day. Last week, a nightly retrain auto-promoted a challenger that happened to clear the metric bar but had, unnoticed, been trained on a day where a labeling pipeline bug mislabeled a batch of legitimate transactions as fraud. The bad model ran in production for 36 hours, auto-declining a spike of legitimate transactions, before a business-metrics alert caught it.

  1. Diagnose exactly which stage-transition gate was missing or insufficient, and explain why "the challenger beat the champion on the metric" was not sufficient evidence to promote.
  2. Redesign the transition from Train → Validate → Approve → Deploy for this pipeline so this failure mode is caught, while preserving nightly retraining as the normal case (i.e., don't just require a human to approve every single night).
  3. What would need to be true about the data stage, upstream of training, to prevent this class of bug rather than just catching it later?

Share this question

← Back to MLOps Model Lifecycle & Governance practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.