Model Release Strategies: Canary, Shadow & Rollback
Ship a bad version of a checkout button and you find out within an hour: error rate spikes, someone gets paged, you roll back. Ship a model whose recall on fraud has quietly dropped from 85% to 60% and nothing spikes. No exception is thrown. Latency is fine. The service returns 200 for every request. The only symptom is that the model is now wrong more often, and "wrong" here is not a crash — it is a probability distribution that has shifted in a way no infrastructure metric detects. This is the central fact that makes releasing a model a genuinely different problem from releasing a normal software change, and it is why this subject exists as its own discipline rather than a footnote on "deployment."
This subject assumes you already know what gets deployed — the artifact, the container, the registry entry (covered in Containers & Reproducible ML Environments, Experiment Tracking & Model Registries, and Model Serving & Deployment) — and focuses narrowly on how you introduce that artifact to live traffic without betting the business on it being right. We cover shadow mode, canary releases, blue/green deployment, gradual ramps, and the rollback triggers that decide when a release reverses itself automatically versus when it waits for a human. The champion/challenger evaluation that decides whether a candidate is even worth releasing is covered in ML Monitoring, Drift & Retraining; here we start from "we have a challenger we believe in" and ask "now what."
Why "Correct" Is Statistical, Not Binary — and Why That Changes Everything
A traditional software release has a mostly binary notion of correctness: the endpoint returns the right shape of response, or it throws an error; the checkout completes, or it doesn't. You can write a unit test that passes or fails. A canary for a normal service watches error rate, latency, and a handful of business counters, and the question it answers is essentially "did we break something."
A model release has no such binary. The challenger model does not "work" or "not work" — it produces a distribution of predictions that is better or worse than the champion's on average, across many requests, in ways you can only observe statistically. Two consequences fall directly out of this, and an interviewer listening for release-strategy maturity is listening for exactly these two points, stated explicitly:
- You need volume before you know anything. A single request scored by the challenger tells you almost nothing about whether the challenger is good — the model could be right or wrong on that one example regardless of its overall quality. Only after enough requests accumulate can you say anything statistically meaningful about precision, calibration, or a business metric. A canary that looks at 50 requests and declares victory is a canary that is measuring noise. This is the single biggest practical difference from software canaries, where a handful of 500 errors is already a strong, immediate signal.
- Correctness can degrade silently and gradually, not just catastrophically. A software bug tends to announce itself (exceptions, timeouts, malformed responses). A model that is 3% worse than the champion on a specific segment announces itself as... nothing, until you specifically measure that segment. This means model release monitoring has to actively watch statistical, model-specific signals — not just infrastructure health — or a real regression sails through a canary that only checks "did the service stay up."
The practical upshot: model releases need more traffic, more time, and different metrics than a typical software canary, and they need a way to observe the challenger's behavior before it is allowed to affect any real decision at all. That last requirement is exactly what shadow mode is for.