Intermediate
Open
Pro
Shrink a 9 GB Serving Image Without Losing Functionality
A model-serving image is 9.1 GB. Its Dockerfile is:
FROM nvidia/cuda:12.1.0-devel-ubuntu22.04
RUN apt-get update && apt-get install -y build-essential git wget
COPY . /app
WORKDIR /app
RUN pip install -r requirements.txt
RUN python setup.py build_ext --inplace # compiles one custom CUDA kernel
RUN rm -rf /root/.cache/pip
CMD ["python", "serve.py"]
- Identify at least three concrete contributors to the image's size or layer bloat in this Dockerfile, referencing specific lines.
- Rewrite the build as a multi-stage Dockerfile that keeps the custom CUDA kernel compilation but minimizes the final image.
- Estimate, directionally, how much smaller the result would be and explain where the savings come from.
Share this question