Containers & Reproducible ML Environments
A data scientist trains a fraud model on their laptop and reports 0.91 AUC. The ML engineer takes the same training script, the same CSV, and runs it on a GPU node in the training cluster. It crashes with a missing shared library. After patching that, it runs — and reports 0.84 AUC, with no error, no warning, just a quietly different number. Nobody changed the code. Nobody changed the data. The laptop had CUDA 11.8 and a newer cuDNN that picked a different (non-deterministic) convolution algorithm; the cluster had CUDA 12.1 and an unpinned numpy that resolved to a version with a different default BLAS backend. Two environments, same code, different answers.
"Works on my machine" is a problem in all of software, and containers are the standard fix everywhere: package the application together with the exact operating system libraries, runtime, and dependencies it needs, so it behaves the same on a laptop, a CI runner, and a production node. Everything in this subject's first half — images, layers, caching, multi-stage builds — is general-purpose software engineering, not ML-specific, and it is worth knowing cold for any infrastructure or backend interview, not just an ML one. What changes in ML is the second half: dependency graphs that include compiled, hardware-specific binaries (CUDA, cuDNN, BLAS libraries), images that are an order of magnitude larger than a typical web service image, and a notion of reproducibility that a container alone cannot deliver, because randomness and data are not "installed" — they have to be seeded and versioned deliberately.
In an interview, being asked "how would you containerize this training job" or "why is your image 6 GB and how would you shrink it" is a proxy for a broader question: do you understand what actually makes ML software hard to reproduce, or have you only ever run pip install on a laptop and gotten lucky?