How does it take fifteen minutes to read <100 GB into GPU memory? Shouldn't that be limited by SSD speed with everything slower than a minute being a terrible ssd?
A lot of cloud platforms have terrible slow network storage. They also might need to compile the GPU kernels fresh as they might not have a persistent CUDA cache.
It also depends on the runtime, vllm is unbelievably slow at model loading compared to llama.cpp
this is not right. vllm can easily load safetensors model at >5GB/s if not faster when setup right. did you make the compile cache persist? if you use docker, you should bind mount the kernel compile cache, so they don't need to be recompiled each time vllm restart.