Match a job Paths Subjects Questions Quizzes Pricing
Intermediate Open Free

Reconstruct 'What's Actually in Production' After a Bucket-of-Files Incident

You've just joined a team whose churn model lives as .pkl files in a shared S3 bucket with names like churn_v2_final.pkl and churn_v2_final_FIXED.pkl. The serving service's config hardcodes a path to one of them. Nobody can say with confidence which file is currently live, what data trained it, or what its offline metrics were.

  1. Design the minimum experiment-tracking + registry setup you'd introduce to prevent this situation from recurring, and explain what each piece specifically fixes.
  2. Your manager asks "can't we just enforce a strict file-naming convention instead, like churn_v{N}_{date}_{author}.pkl? That's much less infrastructure." Explain concretely why a naming convention doesn't solve the actual problem.
  3. Once the new system is in place, what is the first thing you should verify to confirm it would have actually prevented last month's confusion, if it had existed then?
Solution

1. Minimum setup and what each piece fixes

  • Experiment tracking on every training run (notebook or script): auto-generated run ID, logged hyperparameters, metrics over time, and the resulting model artifact, plus the git commit and data snapshot pointer used. This fixes "what data/code/config produced this artifact" — the provenance gap in the current setup.
  • A model registry with a named model (churn-classifier) that accumulates versions, each version linked to the run that produced it. This fixes "which artifact is which" — replacing human-chosen filenames with an unambiguous, auto-assigned identity plus attached metadata.
  • Explicit stage labels (staging/production/archived) with the serving config reading "whatever is currently tagged production" by reference, instead of a hardcoded file path. This fixes "what's actually deployed right now" — the specific question nobody could answer in the original scenario — and makes promotion/rollback a metadata operation instead of a manual file copy.
  • A gated staging → production transition, even lightweight (e.g. requiring the promoting person to be identified and the linked run's metrics to be visible), so there's a record of who decided a version was production-ready and on what evidence.

2. Why a naming convention doesn't solve it

A naming convention is still a human typing a string under time pressure, with no enforcement, no structured metadata attached, and no link back to what actually produced the file. churn_v3_20260815_priya.pkl tells you a version number, date, and author if the person filled it in correctly and consistently — but it doesn't tell you the hyperparameters, the metrics, the exact data snapshot, or the exact code commit, none of which fit naturally into a filename. It also does nothing to solve the actual triggering problem: it's still just a file in a bucket, so "which one is currently live" is still whatever the serving config happens to point at, with no enforced link and no transition history — exactly the property that caused the original incident. A naming convention is a weak substitute for structured metadata and an explicit stage reference, not an alternative to them.

3. What to verify first

Pick a specific past model version (ideally the one at the center of last month's confusion, if it can be reconstructed even approximately) and actually run the "reproduce last month's number" exercise against the new system end to end: query the registry for what was in production on that date, resolve it to its run, confirm the metrics, data snapshot, and code commit are all present and correct, and ideally attempt to actually re-run evaluation and get a matching result. If any link in that chain is missing or ambiguous, the new system has a gap that would still allow the original incident to recur — verifying this concretely, rather than assuming the new system "must" solve it because it's more sophisticated, is the only way to know it actually closes the gap.

Share this question

← Back to Experiment Tracking & Model Registries practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.