Reconstruct 'What's Actually in Production' After a Bucket-of-Files Incident
You've just joined a team whose churn model lives as .pkl files in a
shared S3 bucket with names like churn_v2_final.pkl and
churn_v2_final_FIXED.pkl. The serving service's config hardcodes a
path to one of them. Nobody can say with confidence which file is
currently live, what data trained it, or what its offline metrics were.
- Design the minimum experiment-tracking + registry setup you'd introduce to prevent this situation from recurring, and explain what each piece specifically fixes.
- Your manager asks "can't we just enforce a strict file-naming
convention instead, like
churn_v{N}_{date}_{author}.pkl? That's much less infrastructure." Explain concretely why a naming convention doesn't solve the actual problem. - Once the new system is in place, what is the first thing you should verify to confirm it would have actually prevented last month's confusion, if it had existed then?
1. Minimum setup and what each piece fixes
- Experiment tracking on every training run (notebook or script): auto-generated run ID, logged hyperparameters, metrics over time, and the resulting model artifact, plus the git commit and data snapshot pointer used. This fixes "what data/code/config produced this artifact" — the provenance gap in the current setup.
- A model registry with a named model (
churn-classifier) that accumulates versions, each version linked to the run that produced it. This fixes "which artifact is which" — replacing human-chosen filenames with an unambiguous, auto-assigned identity plus attached metadata. - Explicit stage labels (staging/production/archived) with the serving config reading "whatever is currently tagged production" by reference, instead of a hardcoded file path. This fixes "what's actually deployed right now" — the specific question nobody could answer in the original scenario — and makes promotion/rollback a metadata operation instead of a manual file copy.
- A gated staging → production transition, even lightweight (e.g. requiring the promoting person to be identified and the linked run's metrics to be visible), so there's a record of who decided a version was production-ready and on what evidence.
2. Why a naming convention doesn't solve it
A naming convention is still a human typing a string under time
pressure, with no enforcement, no structured metadata attached, and no
link back to what actually produced the file. churn_v3_20260815_priya.pkl
tells you a version number, date, and author if the person filled it
in correctly and consistently — but it doesn't tell you the
hyperparameters, the metrics, the exact data snapshot, or the exact
code commit, none of which fit naturally into a filename. It also does
nothing to solve the actual triggering problem: it's still just a file
in a bucket, so "which one is currently live" is still whatever the
serving config happens to point at, with no enforced link and no
transition history — exactly the property that caused the original
incident. A naming convention is a weak substitute for structured
metadata and an explicit stage reference, not an alternative to them.
3. What to verify first
Pick a specific past model version (ideally the one at the center of last month's confusion, if it can be reconstructed even approximately) and actually run the "reproduce last month's number" exercise against the new system end to end: query the registry for what was in production on that date, resolve it to its run, confirm the metrics, data snapshot, and code commit are all present and correct, and ideally attempt to actually re-run evaluation and get a matching result. If any link in that chain is missing or ambiguous, the new system has a gap that would still allow the original incident to recur — verifying this concretely, rather than assuming the new system "must" solve it because it's more sophisticated, is the only way to know it actually closes the gap.
Share this question