Practice — Containers & Reproducible ML Environments (5 questions)
Intermediate
Open
Free
Diagnose a CUDA Driver/Toolkit Mismatch After a Deploy Permalink →
Your team ships a new training image: FROM nvidia/cuda:12.4.0-devel-ubuntu22.04,
with PyTorch installed via pip install torch on top. It builds and passes CI
(CPU-only unit tests) without issue. The next morning, every training pod
scheduled on your GPU node pool crashes on startup with:
CUDA error: CUDA driver version is insufficient for CUDA runtime version
The previous image, built from nvidia/cuda:11.8.0-devel-ubuntu22.04, worked
fine on the same nodes.
- Explain exactly what is mismatched here and why the CI pipeline didn't catch it.
- Name two different fixes, at two different layers of the stack, and the trade-offs between them.
- Why can't you simply "put a newer driver inside the image" to fix this once and for all?
Share this question
Intermediate
Open
Pro
Shrink a 9 GB Serving Image Without Losing Functionality
Unlock this question →
Intermediate
Open
Pro
Explain a Silent Metric Regression Traced to an Unpinned Dependency
Unlock this question →
Intermediate
Open
Pro
A 'Reproducible' Training Run That Isn't Fully Reproducible
Unlock this question →
Intermediate
Open
Pro