Llama.cpp AI Performance with the GeForce RTX 5090 Review
phoronix.com
phoronix.com
Then Apple introduced Macs with 128 GB and more unified memory at 800GB/s and the ability to load models as large as 70GB (70b FP8) or even larger ones. The M1 Ultra was unable to take full advantage of the excellent RAM speed, but with the M2 and the M3, performance is improving. Just be prepared to spend 5000€ or more for a M3 Ultra. Another alternative would be a EPYC 9005 system with 12x DDR5-6000 RAM for 576GB/s of memory bandwidth with the LLM (preferably MoE) running on the CPU instead of a GPU.
However today, with the latest, surprisingly good reasoning models like QwQ-32B using up thousands or tens of thousands of tokens in their replies, performance is getting more important than previously and these systems (Macs and even RTX 3090s) might fall out of favor, because waiting for a finished reply will take several minutes or even tens of minutes. Nvidia Ampere and Apple silicon (AFAIK) are also missing FP4 support in hardware, which doesn't help.
For the same reason AMD Halo Strix with a mere 273GB/s of RAM bandwidth and perhaps also NVidia Project Digits (also speculated to offer similar RAM bandwidth) might just be too slow for reasoning models with more than 50GB or so of active parameters.
On the other hand, if the prices for the RTX 5090 remain at 3500€, they will likely remain insignificant for the DIY crowd for that reason alone.
Perhaps AMD will take the crown with a variant of their RDNA4 RX 9070 card with 32GB of VRAM priced at around 1000€? Probably wishful thinking…
There's weight-only FP8 in vLLM on NVidia Ampere: https://docs.vllm.ai/en/latest/features/quantization/fp8.htm...
Granted the 48 GB 4090s started there too before materializing and even becoming common enough to be on eBay, but this time there are more technical barriers and it's less likely they'll ever show up in meaningful numbers.
Each of the cards except the 5090 gets almost exactly 0.1 token/s per GB/s memory bandwidth.
My understanding is that the Macs have soldered memory which allows for much higher memory bandwidth. The M4 has ~400-550 GB/s max depending on configuration[2], while EPYCs seem to have more like 250GB/s max[3].
[1]: https://news.ycombinator.com/item?id=42847284
[2]: https://support.apple.com/en-us/121553
[3]: https://www.servethehome.com/here-is-why-you-should-fully-po...
Your link goes to info on the 2022 EPYC CPUs, the current generation can do 576GB/s: https://chipsandcheese.com/p/amds-turin-5th-gen-epyc-launche...
Intel's current 12ch Xeons should be even faster with MRDIMMs, though I couldnt find a memory specific benchmark.
Thanks for the correction.
edit: found a llama.cpp issue discussing performance bottlenecks on modern dual-socket EPYC here[1]. Also includes single-socket benchmarks, and includes some optimizations. Just thought it was interesting.
[1]: https://github.com/ggml-org/llama.cpp/discussions/11733
I think we should put more effort into compressing the models well to be able to use local GPUs better.
FP4 only helps batched inferencing.
3090 is the last gen to have proper nvlink support, which is supported for LLM inference in some frameworks.
Would 1x 5090 be faster than 2x 3090?
The former can have higher peak throughput while the latter can have lower latency, though it depends on the details[1].
[1]: https://blog.squeezebits.com/vllm-vs-tensorrtllm-9-paralleli...
AFAIK this only happens directly over PCIe if using hacked drivers with p2p enabled (I think tinygrad/tinybox provided these drivers initially.)
Otherwise, data goes through system bus/CPU first. `nvidia-smi topo -m` will show how the GPUs are connected.
It will work, but with less memory bandwidth. If using hacked drivers with p2p enabled, it will depend on the PCIe topology. Otherwise, the data will take a longer path.
Depending on the model, it may not be that big of a hit in performance. Personally, I have seen higher perf on nvlinked GPUs and models that can fit on 2x 3090 24gbs.
But the lower bounds on intermediate tokens on CoT to get close to PTIME expressability + QwQ-32B's verbosity/repeating on those tokens does eat up that pretty quickly.
In theory the DAG walking QwQ-32B appears to be doing should require O(|E| log|V|) scratch space. IMHO they need to train to target that more.
It seems fairly trivial to extend the context loss problem to the public UI, The Ω(n) scratch space may be hard to hit in reality but they seem almost exponential in scratch space right now, even on tasks that a traditional LLM could answer with just approximate retrieval.
I may just be suffering confirmation bias, but the problem with memory in this case seems to be related to making the intermediate tokens palatable to users IMHO.
For text generation, which is the most important metric, the tokens per second will scale almost linearly with memory bandwidth (936 GB/s, 1008 GB/s and 1792 GB/s respectively), but we might see more interesting results when comparing prompt processing, speculative decoding with various models, vLLM vs llama.cpp vs TGI, prompt length, context length, text type/programming language (actually makes a difference with speculative decoding), cache quantization and sampling methods. Results should also be checked for correctness (perplexity or some benchmark like HumanEval etc.) to make sure that results are not garbage.
If anyone from Phoronix is reading this, this post might be a good point to get you started: https://old.reddit.com/r/LocalLLaMA/comments/1h5uq43/llamacp...
At time of writing, Qwen2.5-Coder-32B-Instruct-GGUF with one of the smaller variants for speculative decoding is probably the best local model for most programming tasks, but keep an eye out for any new models. They will probably show up in Bartowksi's "Recommended large models" list, which is also a good place to download quantized models: https://huggingface.co/bartowski
I use ollama for this, and I'm getting useful stuff out of qwq:32b as the architect, qwen2.5-coder:32b as the edit model, and dolphin3:8b as the weak model (which gets used for things like commit messages). Now what that means is that performance swapping these models in and out of the card starts to matter, because they don't all go into VRAM at once; but also using a reasoning model means that you need straight-line tokens per second as well, plus well-tuned context length so as not to starve the architect.
I haven't investigated whether a speculative decoding setup would actually help here, I've not come across anyone doing that with a reasoner before now but presumably it would work.
It would be good to see a benchmark based on practical aider workflows. I'm not aware of one but it should be a good all-round stress test of a lot of different performance boundaries.
That said, it's cool that you can even get an L4 with 24 GB of VRAM that actually performs okay, yet is passively cooled and consumes like 70W, at that point you can throw a bunch of them into a chassis and if you haven't bankrupted yourself by then, they're pretty good.
I did try them out on Scaleway, the pricing isn't even that exorbitant, using consumer GPUs for LLM use cases doesn't quite hit the same since.
I think it's an incredibly tall order to get into the GPU game for a startup, but it should be good entertainment if nothing else.
Sounds like hype
I think there are quite a few constraints you are glossing over here.
That said, it's understandable that companies are sticking with whatever works for now, but occasionally you get an immensely cool project that attempts to do something differently, like Intel's Larrabee did, for example.
https://www.tomshardware.com/pc-components/dram/nvidia-repor...
The CEO of SK Hynix have confirmed that it is in the works:
https://www.mk.co.kr/en/it/11245259
And big companies have started to analyze it, to figure out how it will fit into their domain:
https://x.com/Jukanlosreve/status/1892916771692421228?t=3ikB...
If so that would confirm the notion that they've hit a ceiling and pushing against physical limitations.
The 50-series is made using the same manufacturing process ("node") as the 40-series, and there is not a major difference in design.
So the 50-series is more like tweaking an engine that previously topped out at 5000 RPM so it's now topping out at 6000 RPM, without changing anything fundamental. Yes it's making more horsepower but it's using more fuel to do so.
https://www.hardware-corner.net/guides/gpu-benchmark-large-l...
Power consumption is almost 10x smaller for apple.
Vram is more than 10x larger.
Price wise for running same size models apple is cheaper.
Upper limit (larger models, longer context) is far larger for apple (for nvidia you can easily put 2x cards, more than that it becomes whole complex setup no ordinary person can do).
Am I missing something or apple is simply currently better for local llms?
On apple silicons you can always use MoE models, which work beautifully. On RTX it's kind of waste to be honest to run MoE, you'd be better off running single, whole active model that fills available memory (with enough space for the context).
The OG DeepSeek models are hundreds of GB quantized, nobody is using RTX GPUs to run them anyway…
This just is just a single user chat experience benchmark.
There's no point to performance without context. At its pricepoint, the 5080 is purely a gaming card.
Ampere is the floor, in practice, which is what effectively makes the "buy a Mac Studio" crowd P40 people with 10x the budget.
https://videocardz.com/newz/nvidia-rtx-pro-6000-blackwell-le...
Also, some DeepSeek models would be cool.
He already did the general compute benchmark of the two 9070 cards here[1], and between poorly optimized drivers and GDDR6's lower memory bandwidth, I wouldn't expect any great scores.
In terms of memory bandwidth they're between a 4070 and a 4070 Ti SUPER, and given that LLMs are very memory-bandwidth constrained as I mentioned in another comment, at best you'd expect the LLM score to end up between the 4070 and the 4070 Ti SUPER.
[1]: https://www.phoronix.com/review/amd-radeon-rx9070-linux-comp...
Have a look at the GamersNexus YT rant about it... They make a fair argument in that price range an older model used RTX nvidia card may be a better value.
Depends on your use-case, as rtx 5090 nvidia AI frame interpolation is dog crap hype for CGI or CUDA accelerated ML libraries.
Personally, I would go with the rtx 4090 or even an rtx 3090 with 24G vram for ML and CGI workstation, as CUDA+Optix has better software support. For just gaming, the 9070 XT is a better deal when the MSRP is within range. Depends how willing you are to get ripped off by scalper prices right now. lol =3
It looks to me like you could think about it as a performance/VRAM/convenience stepping-stone between having one 4090 and having a pair.
Paired 5090s, if such a thing is possible, sounds like a very good way to spend a lot of money very quickly while possibly setting things on fire, and you'd have to have a good reason for that.