Advanced
Open
Pro
Blue/Green Deployment Is Not a Safety Net by Itself
An ML platform team implements blue/green deployment for model serving: the new model version is deployed to a fully separate, warmed-up "green" environment, tested with a smoke test (a handful of manually-crafted example requests), and then a single load-balancer config change routes 100% of traffic from blue to green instantaneously. They tell you: "We've solved the risk problem — rollback is now a one-line config change, so we can revert instantly if anything goes wrong."
- Identify what blue/green deployment, as described, actually solves and what it does not solve, specifically for a model release (rather than a typical stateless-service release).
- Walk through a concrete failure scenario where this setup causes real damage before anyone notices, despite the "instant rollback" capability.
- Propose the minimal change to this pipeline that would close the gap, without throwing away the blue/green infrastructure they've already built.
Share this question