Match a job Paths Subjects Questions Quizzes Pricing
Intermediate Open Pro

Design a Data-Versioning Scheme for a Retraining Pipeline

A retraining pipeline reads from a table user_features in your warehouse, which an upstream job updates in place every night (new rows appended, some old rows corrected or deleted). Three months from now, a stakeholder asks "can you reproduce the exact number you reported for the March 3rd model?" and nobody can answer confidently, because user_features today looks nothing like it did on March 3rd.

  1. Explain precisely why "we'll just rerun the training script" does not answer the stakeholder's question, even with the code and container image both preserved.
  2. Propose a concrete data-versioning scheme that would have prevented this, at the concept level (you do not need to name a specific vendor tool).
  3. What would you additionally need to have logged at training time in order to answer "reproduce March 3rd" requests going forward, beyond just the data version?

Share this question

← Back to Containers & Reproducible ML Environments practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.