Advanced
Open
Pro
Explain the Lifecycle Loop to a Skeptical Engineering Lead
An engineering lead who has shipped plenty of software but never ML says: "I don't get why 'deploy' isn't just the last step, like it is for a normal service. We push code, it runs, we're done until the next feature." They also ask: "Why do you need a human to sign off on a retrain? It's the same pipeline running again — if we trusted it once, why not trust it every time?"
- Answer their first question: explain, using the full lifecycle loop, why "deploy" is not the end for an ML model in a way that it effectively is for typical stateless software.
- Answer their second question: explain why an automated retrain re-enters the validate/approve gates rather than auto-promoting directly, using a concrete example of something that could differ between "the pipeline that trained the first model" and "the same pipeline run again."
- They push back: "Doesn't requiring approval on every retrain just recreate the alert-fatigue problem you'd complain about in monitoring — everyone rubber-stamping because it's always fine?" How do you design around that specific risk?
Share this question