Diagnose a CUDA Driver/Toolkit Mismatch After a Deploy
Your team ships a new training image: FROM nvidia/cuda:12.4.0-devel-ubuntu22.04,
with PyTorch installed via pip install torch on top. It builds and passes CI
(CPU-only unit tests) without issue. The next morning, every training pod
scheduled on your GPU node pool crashes on startup with:
CUDA error: CUDA driver version is insufficient for CUDA runtime version
The previous image, built from nvidia/cuda:11.8.0-devel-ubuntu22.04, worked
fine on the same nodes.
- Explain exactly what is mismatched here and why the CI pipeline didn't catch it.
- Name two different fixes, at two different layers of the stack, and the trade-offs between them.
- Why can't you simply "put a newer driver inside the image" to fix this once and for all?
1. What's mismatched
The NVIDIA driver lives on the host (the GPU node) and includes the kernel module that talks to the physical GPU; it is not something a container can supply, because a container shares the host's kernel. The CUDA toolkit (here, 12.4) lives inside the image. The compatibility rule is directional: the host driver must be new enough to support the CUDA toolkit version in the image. The node pool's driver was installed for the CUDA 11.8 era and is too old to support a CUDA 12.4 toolkit runtime — hence the startup error.
CI didn't catch it because the unit tests are CPU-only: they never exercise a real GPU driver/toolkit handshake, so a driver/toolkit incompatibility is invisible until a pod actually tries to initialize CUDA on real hardware. This is a build-time-invisible, deploy-time-visible failure by nature.
2. Two fixes at two layers
- Application layer (fastest, no infra change): roll the image back to (or pin future images to) a CUDA toolkit version the current node driver supports, e.g. stay on CUDA 12.1 if the driver supports up to 12.1. Trade-off: blocks you from using any CUDA-12.4-only feature or library until the infra side catches up.
- Infrastructure layer: upgrade the GPU node pool's driver (via the AMI / node image / DaemonSet that installs the NVIDIA driver) to a version that supports CUDA 12.4, then re-deploy the new image. Trade-off: this is a fleet-wide, often manually-applied change (drivers aren't typically part of the app's CI/CD pipeline), needs a maintenance window or rolling node replacement, and needs to be validated against every other workload sharing that node pool, not just this one image.
In practice, teams do both: pin a driver floor in their node provisioning (Terraform/AMI) that comfortably exceeds the CUDA toolkit version their images use, with headroom, and add a GPU smoke test to CI/CD (actually running a trivial CUDA op on a real GPU runner, or at minimum a staging deploy gate) so this class of failure is caught before full rollout rather than after.
3. Why not "just put a newer driver in the image"
The driver includes a kernel module, and kernel modules must match the host's running kernel — a driver baked into a container image would need to be loaded into the host kernel from inside the container, which defeats the isolation model containers rely on (shared kernel, no guest kernel to install a driver into). The NVIDIA Container Toolkit's whole design is to mount the host's already-loaded driver libraries and device files into the container at runtime specifically so the driver stays a host-level concern. This is why the fix has to happen on the host side (or by using an older toolkit in the image) — there is no image-only fix.
Share this question