Reconstruct a Denied-Loan Decision for a Regulator
A regulator asks your lending company to explain a specific credit decision made 14 months ago: "User X was denied a loan on 2025-06-20. Explain the decision, including which model made it and why it was trusted to be in production." Your prediction logs retain only 90 days of history, and your model registry only keeps the current production version of each model — older versions are deleted once a newer one is promoted, to save storage.
- Walk through, step by step, what you would need to reconstruct the full chain from "the decision" back to "why the model was trusted," and identify exactly where your current setup breaks the chain.
- Propose a retention policy (what to keep, how long) that would have prevented this gap, and justify the durations.
- Your CTO argues "we can't keep every old model artifact forever, storage isn't free." How do you respond, and is there a legitimate version of the concern worth addressing?
1. The chain and where it breaks
Reconstructing the decision requires, in order: (a) the prediction log entry for user X on 2025-06-20 — model version, feature values as served, score, final decision; (b) that model version's artifact and its lineage — training data snapshot pointer, code SHA, environment digest; (c) the approval record for that version — who signed off, what evaluation and fairness report they reviewed, the model card; (d) monitoring history for that version around 2025-06-20, to show it was not known to be degraded.
Two breaks in the current setup: prediction logs only cover 90 days, so the log entry for a 14-month-old decision is already gone — you cannot even confirm which model version scored user X. And even if the log had survived, the model registry deletes older versions once a newer one is promoted, so the artifact, its training data pointer, and (implicitly) its evaluation report for that version no longer exist. Both breaks independently make the decision unreconstructable; together the answer to the regulator is "we don't know," which is not defensible.
2. A retention policy
- Prediction logs: retain at least as long as a decision could plausibly be challenged — for consumer lending this is commonly multi-year (2–7 years depending on jurisdiction); check with legal, but 90 days is not in the right order of magnitude.
- Model artifacts and their full lineage (data snapshot pointer, code SHA, environment digest, evaluation report): retain for at least as long as any prediction made by that version is still inside its own retention window. Never delete a model version while predictions it made are still retained — the two retention clocks must be linked.
- Approval and sign-off records: retain indefinitely or per legal document-retention policy; these are usually the first artifact a regulator asks for because they show process, not just outcome.
- Monitoring/drift history: retain through the retention window of any prediction made while that history applies.
3. Responding to the storage-cost concern
Storage for model artifacts, evaluation reports, and even prediction logs is cheap relative to the cost of a regulatory finding, forced retraining, or an expensive discovery process in litigation where the answer is "we have to guess" instead of "here is the record" — the asymmetry strongly favors retention. The legitimate part of the concern is that not everything needs the same retention: intermediate training artifacts (checkpoints, scratch data) genuinely can be pruned aggressively, and old staging models that never served production traffic don't carry the same obligation as promoted production versions. The fix is a tiered retention policy — aggressive cleanup for non-production artifacts, strict retention tied to prediction-log lifetime for anything that actually served a decision — not a blanket "keep everything forever" or "delete on promotion" policy.
Share this question