Advanced
Open
Pro
Find the Broken Links in a Reproducibility Chain
Your team's model registry stores, per production model: the trained
weights file, the offline AUC and log-loss at promotion time, and a
free-text "notes" field where the training engineer sometimes writes
what data they used. Training code lives in a shared notebooks/
folder that everyone edits directly on the main branch; there is no
per-run snapshot of which version of the notebook produced a given
model. Training data is queried live from a warehouse table that is
continuously updated (no partitioning or snapshotting).
- For each of the four things that must be versioned (data, code, model, environment), identify whether this team's setup actually achieves it, and if not, what specifically is missing.
- Six months from now, someone asks "can we reproduce the exact training run that produced the currently-deployed model, to debug why it's behaving oddly on a specific input?" Walk through what would go wrong when they try.
- Propose the minimum set of changes that would close the gaps, ordered by how much risk each one removes.
Share this question