Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — Containers & Reproducible ML Environments (5 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Intermediate Open Free

Diagnose a CUDA Driver/Toolkit Mismatch After a Deploy Permalink →

Your team ships a new training image: FROM nvidia/cuda:12.4.0-devel-ubuntu22.04, with PyTorch installed via pip install torch on top. It builds and passes CI (CPU-only unit tests) without issue. The next morning, every training pod scheduled on your GPU node pool crashes on startup with:

CUDA error: CUDA driver version is insufficient for CUDA runtime version

The previous image, built from nvidia/cuda:11.8.0-devel-ubuntu22.04, worked fine on the same nodes.

  1. Explain exactly what is mismatched here and why the CI pipeline didn't catch it.
  2. Name two different fixes, at two different layers of the stack, and the trade-offs between them.
  3. Why can't you simply "put a newer driver inside the image" to fix this once and for all?

Share this question

Intermediate Open Pro

Shrink a 9 GB Serving Image Without Losing Functionality

Unlock this question →
Intermediate Open Pro

Explain a Silent Metric Regression Traced to an Unpinned Dependency

Unlock this question →
Intermediate Open Pro

A 'Reproducible' Training Run That Isn't Fully Reproducible

Unlock this question →
Intermediate Open Pro

Design a Data-Versioning Scheme for a Retraining Pipeline

Unlock this question →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.