CI/CD for ML & Data Code
A backend engineer opens a pull request that changes a discount-calculation function. CI runs unit tests, they pass, a reviewer reads the diff, it merges, it deploys. Nothing about the data the function operates on changed — the function is the whole artifact, and the diff is the whole story.
A data scientist opens a pull request that changes one line: the training window for a churn model, from 90 days of history to 180. CI runs the same kind of unit tests, they pass, a reviewer reads the diff — and approves it, because the diff looks like a one-line change. But the model that gets trained from that code is not a one-line change. It is a different model, trained on different data, that may perform better or worse on customers who churned early versus late, may shift which features matter, and may regress silently on a segment nobody is looking at. The code review approved a line; nobody reviewed the model. That gap — a change that looks small in a diff but produces an artifact that needs to be re-evaluated, not just re-read — is what this subject is about.
Scope note, stated up front: this subject is about general ML/data engineering CI/CD practice — tool-agnostic, whether you wire it with GitHub Actions, GitLab CI, Jenkins, or something else. It is explicitly distinct from the Headless Mode, CI Automation & the Agent SDK subject, which is about using Claude Code specifically inside a CI pipeline (headless invocation, the Agent SDK as a library). If the question is "how do I get an AI coding agent to run inside my pipeline," that's the other subject. If the question is "what should my ML pipeline's CI/CD actually check, and when should it block a merge or a deploy," that's this one.
In an interview, "how would you set up CI/CD for a model training repo" separates candidates who have only shipped application code from candidates who have shipped ML in production: the former describe lint/test/deploy; the latter add data validation, model evaluation, and a registry-mediated promotion step, and can explain why each exists.
What Makes ML/Data CI Different
General software CI answers one question: does the code behave correctly, given the tests we wrote? ML and data CI has to answer three additional questions, because the thing being shipped is not just code:
| Question | General software CI | ML/data CI |
|---|---|---|
| Does the code run without error? | Unit tests, type checks, linting | Same — necessary but far from sufficient |
| Does the code produce correct data? | N/A (code has no "data shape") | Schema/contract tests on transforms, not just function-level unit tests |
| Does the code produce a correct model? | N/A | Offline evaluation gates: does the new model regress key metrics vs. the current one? |
| Is what's about to deploy actually validated? | Deploy the code that passed tests | Deploy the model artifact that passed evaluation — a different, larger claim than "the training code has no syntax errors" |
The core mental model, worth stating explicitly because it changes what "the diff" means: the artifact is model + data + code, not just code. A pull request that changes zero lines of training code but points at a new data snapshot produces a different model — a different thing to evaluate, review, and potentially roll back — with exactly the same stakes as a code change that changes model behavior directly. Treating "no code changed" as "nothing to review" is the single most common gap between teams that have general software CI and teams that have ML CI.