this is not right. vllm can easily load safetensors model at >5GB/s if not faster when setup right. did you make the compile cache persist? if you use docker, you should bind mount the kernel compile cache, so they don't need to be recompiled each time vllm restart.