Practice — Experiment Tracking & Model Registries (6 questions)
Reconstruct 'What's Actually in Production' After a Bucket-of-Files Incident Permalink →
You've just joined a team whose churn model lives as .pkl files in a
shared S3 bucket with names like churn_v2_final.pkl and
churn_v2_final_FIXED.pkl. The serving service's config hardcodes a
path to one of them. Nobody can say with confidence which file is
currently live, what data trained it, or what its offline metrics were.
- Design the minimum experiment-tracking + registry setup you'd introduce to prevent this situation from recurring, and explain what each piece specifically fixes.
- Your manager asks "can't we just enforce a strict file-naming
convention instead, like
churn_v{N}_{date}_{author}.pkl? That's much less infrastructure." Explain concretely why a naming convention doesn't solve the actual problem. - Once the new system is in place, what is the first thing you should verify to confirm it would have actually prevented last month's confusion, if it had existed then?
Share this question
Answer a Regulatory Lineage Request Under Time Pressure
Unlock this question →Design a Promotion Gating Policy Across Model Risk Tiers
Unlock this question →Use Tracked Runs to Choose Among a Hyperparameter Sweep
Unlock this question →The Registry Says Nothing Changed. Something Changed. Permalink →
Your model registry promotes fraud-classifier/v12, storing the model
binary, training metrics, and a content hash of the weights. It does
not version or pin the upstream feature-computation code. Two
weeks later, unrelated to any model work, a data engineer refactors
the feature pipeline — changing how avg_transaction_amount_30d is
computed (e.g., from a calendar-day window to a rolling 30x24h
window). Nothing about fraud-classifier/v12 changes: same version,
same weights, same hash.
What happens to v12's production behavior?
Share this question