These already exist, they're just disproportionately expensive compared to consumer cards. You can get a Nvidia A100 with 80GB VRAM on eBay for about $14,000.
3.33x the memory for 9x the price.
In Llama V1, 7B was clever but kinda dumb, 13B was still short sighted, but 33B was a big jump to "disturbingly good" territory. 7B/13B finetunes required very specific prompting, but 33B would still give good responses even outside the finetuning format.
16GB is probably perfect for a ~13B model with a very long context and better (5 bit?) quantization, or just partially offloading a ~33B model.
Llama in general is not great for code completion and fact retrieval, from what I have tried
I can't run llama2 70B - but I remember llama1-Guanaco30B was a major leap over the base llama30B model for me.
And you can run this locally with enough combined RAM + VRAM.
First some background: llama is divided into prompt ingestion code, and "layers" for actually generating the tokens.
There are different offloading schemes, but what llama.cpp specifically does is map the layers to different devices. For example, ~7GB of the model could reside on the 8GB GPU, and the other ~9GB would live on the CPU.
During runtime, the prompt is ingested by the GPU all at once (which doesn't take much VRAM), and then for each word, the layers are run sequentially. So the first ~half of a word would run on your GPU, and the last half would run on your CPU, and the alternation repeats till all the words are generated.
The beauty of llama.cpp is that its cpu token generation is (compared to other llama runtimes) extremely fast. So offloading even half or two thirds of a model to a decent CPU is not so bad.
A 3090.
Ampere will be well supported until its very obsolete because of the A100.
The M1 and M2 Apple processors might change that equation.
The 32gb and 64gb MacBooks can run inference on a lot of these big models.
They might be quite a bit slower in TFLOPS than a 4070, but if they can run the model while the 4070 can't run it due to limited RAM, then they come out the winner
The long and short of it is, right now, CUDA beats TFLOPS.
Now for training, $10K for an M2 with 192GB memory could break even somewhere around a dozen models with current costs, iff the compute is not saturated. The limit of memory bandwidth is only a function of the compute; it doesn't scale infinitely.