Intermediate
Open
Pro
Diagnose an Incident: Code Deploy or Model Promotion?
At 14:00, your team ships a serving-code release (a refactor of the feature-fetching client) to production. Separately, at 14:10, a model that finished its staging validation earlier that day gets promoted to "production" in the model registry — a routine, independently-scheduled promotion unrelated to the 14:00 code deploy. At 14:20, the on-call is paged: prediction latency is normal, but the positive-prediction rate has jumped sharply.
- Why is it important that the code deploy and the model promotion are independently tracked events, given this incident?
- Describe the diagnostic steps you'd take to determine whether the code deploy or the model promotion (or something else) caused the jump.
- Suppose you determine the model promotion is the cause. What are the mechanics of the fix, and why is it a materially faster and safer operation than it would be if the model were baked into the serving code's container image?
Share this question