> We’re releasing Llama 3.1 405B
Is it possible to run this with ollama?
Is it possible to run this with ollama?
Ollama will offload as many layers as it can to the gpu then the rest will run on the cpu/ram.
Can anyone share the cost of their pre-built clusters, they’ve recently started selling? (sorry feeling lazy to research atm, I might do that later when I have more time).
https://smicro.eu/nvidia-hgx-h100-640gb-935-24287-0001-000-1
8x H100 HGX cluster for €250k + VAT