Intermediate
Open
Pro
Design a Data-Versioning Scheme for a Retraining Pipeline
A retraining pipeline reads from a table user_features in your warehouse,
which an upstream job updates in place every night (new rows appended,
some old rows corrected or deleted). Three months from now, a stakeholder
asks "can you reproduce the exact number you reported for the March 3rd
model?" and nobody can answer confidently, because user_features today
looks nothing like it did on March 3rd.
- Explain precisely why "we'll just rerun the training script" does not answer the stakeholder's question, even with the code and container image both preserved.
- Propose a concrete data-versioning scheme that would have prevented this, at the concept level (you do not need to name a specific vendor tool).
- What would you additionally need to have logged at training time in order to answer "reproduce March 3rd" requests going forward, beyond just the data version?
Share this question