Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — Experiment Tracking & Model Registries (6 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Intermediate Open Free

Reconstruct 'What's Actually in Production' After a Bucket-of-Files Incident Permalink →

You've just joined a team whose churn model lives as .pkl files in a shared S3 bucket with names like churn_v2_final.pkl and churn_v2_final_FIXED.pkl. The serving service's config hardcodes a path to one of them. Nobody can say with confidence which file is currently live, what data trained it, or what its offline metrics were.

  1. Design the minimum experiment-tracking + registry setup you'd introduce to prevent this situation from recurring, and explain what each piece specifically fixes.
  2. Your manager asks "can't we just enforce a strict file-naming convention instead, like churn_v{N}_{date}_{author}.pkl? That's much less infrastructure." Explain concretely why a naming convention doesn't solve the actual problem.
  3. Once the new system is in place, what is the first thing you should verify to confirm it would have actually prevented last month's confusion, if it had existed then?

Share this question

Intermediate Open Pro

Answer a Regulatory Lineage Request Under Time Pressure

Unlock this question →
Intermediate Open Pro

Design a Promotion Gating Policy Across Model Risk Tiers

Unlock this question →
Intermediate Open Pro

Use Tracked Runs to Choose Among a Hyperparameter Sweep

Unlock this question →
Intermediate Open Pro

Compare Rollback Speed With and Without a Registry

Unlock this question →
Advanced Open Free

The Registry Says Nothing Changed. Something Changed. Permalink →

Your model registry promotes fraud-classifier/v12, storing the model binary, training metrics, and a content hash of the weights. It does not version or pin the upstream feature-computation code. Two weeks later, unrelated to any model work, a data engineer refactors the feature pipeline — changing how avg_transaction_amount_30d is computed (e.g., from a calendar-day window to a rolling 30x24h window). Nothing about fraud-classifier/v12 changes: same version, same weights, same hash.

What happens to v12's production behavior?

Share this question

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.