Why does the model data need to be stored in the image? Download the model data on container startup using whatever method works best.
even the smaller nvidia images (like nvidia/cuda:13.1.1-cudnn-runtime-ubuntu24.04) are about 2Gb before adding any python deps and that is a problem.
if you split the image into chunks and pull on-demand, your container will start much faster.