Match a job Paths Subjects Questions Quizzes Pricing

Experiment Tracking & Model Registries

From a folder of notebooks nobody trusts to a system of record for what's actually in production

Overview Read

Experiment Tracking & Model Registries

Friday, 4:50 p.m. A VP asks the on-call data scientist why churn predictions look off this week. The on-call opens the model-serving config and finds it points at models/churn_v_final_REAL_v2_use_this_one.pkl in a shared S3 bucket. Nobody currently on the team wrote that file. Slack search turns up a message from four months ago — "pushed the new churn model, slightly better than last week's" — from someone who left the company in March. There is no record of what hyperparameters trained it, what data it saw, what its offline metrics were, or whether "slightly better" meant 2% better or 0.2% better. A second file, churn_v_final_REAL.pkl, sits in the same folder, untouched since February; a third notebook, churn_model_v3_FINAL_FINAL.ipynb, produces a model that doesn't match either file's predictions on a test batch. Nobody can say with confidence which of the three is actually loaded in production right now, let alone reproduce the number that was reported to the VP's team last month. The honest answer to "why do predictions look off" is "we don't know what's running, so we can't know what changed" — and that answer, delivered on a Friday afternoon, is the entire failure mode this subject exists to prevent.

Contrast that with a tracked workflow. Every training run — from a notebook, a script, or an automated retrain — is automatically logged with a unique ID, the exact hyperparameters, metrics recorded over the course of training, and the resulting artifacts. A registry sits on top of those runs as the answer to one specific, high-stakes question — what is actually deployed right now — with explicit stages (staging, production, archived) and a record of who promoted what, when, and why. Tracing back from a specific production prediction to the exact model version, the training data snapshot, and the code commit that produced it is a query, not an archaeology project. The VP's question gets answered in the time it takes to open a dashboard.

Tool-agnostic framing, stated up front: this subject describes the pattern, not a specific product. MLflow, Weights & Biases, Neptune, and several others all implement roughly the same shape — a run has an ID, parameters, metrics over time, and artifacts; a registry sits above runs with named models and stages — because the shape follows directly from the problem, not from any one vendor's design choices. In an interview, you should be able to describe this pattern without naming a specific tool, and explain why each piece exists by pointing at the incident above.


Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.