The fact a laptop can run 70B+ parameter models is a miracle, it's not what the chip was built to do at all.
Those comparisons are unreasonable in a sense, but they are implied by statements like GPs "Hence a lot of people in the local LLMA community are really going after high-memory Macs".
Do you need to do fine tuning on a smaller model and need the highest inference performance with smaller models? Are you planning to use it as a lab to learn how to work with tools that are used in Big Tech (i.e. CUDA)? Or do you just want to do slow inference on super huge models (e.g Grok)?
Personally, I chose the Nvidia route because as a backend engineer, Macs aren’t seriously used in datacenters. The same frameworks I use to develop on a 3090 are transferable to massive, infiniband-connected clusters with TB of VRAM.
- the "hobbyists" with $5k GPUs
- People that work in the industry that never used "not mac" or even if they did - explaining to IT that you need a PC with RTX A6000 48GB instead of a mac like literally everyone else in the company is a loosing battle.
- people that work outside Silicon Valley, where the entire company uses Windows centrally managed through Active Directory, and explaining IT that you need an Mac is an uphill battle. So you just submit your request for an RTX A6000 48GB to be added to your existing workstation
Those people are the intended target customer of the A6000, and there are a lot of them.
Companies that run windows and AD are too busy to make sure you move your mouse every 5 minutes while you're on the clock more. At least that is my experience.
You're certainly right that with a macbook you get a whole computer, so you're getting more for your money. And it's a luxury high-end computer too!
But personally, I've never seen anyone step directly from not-even-having-a-PC to buying a 4090 for $1800. Folks that aren't technically inclined by and large stick with hosted models like ChatGPT.
More common in my experience is for technical folks with, say, an 8GB GPU to experiment with local ML, decide they're interested in it, then step up to a 4090 or something.
Most laptops with 64+GB of RAM can run a 70B model at 4-bit quantization. It’s not a miracle, it’s just math. M2 can do it faster than systems with slower memory bandwidth.
I have a gaming rig with a 4080 with 16GB of RAM and it can't even run Mixtral (kind of the minimum bar of a useful generic LLM in my opinion) without being heavily quantized. Yeah it's fast when something fits on it, but I don't see much point in very fast generation of bad output. A refurbished M1 Max with 32GB of RAM will enable you to generate better quality LLM output than even a 4090 with 24GB of VRAM and for ~$300 less, and it's a whole computer instead of a single part that still needs a computer around it. Compared to my 4080, that GPU and the surrounding computer get you half the VRAM capacity for greater cost than the Mac.
If you're building a rig with multiple GPUs to support many users or for internal private services and are willing to drop more than $3k then I think the equation swings back in favor of Nvidia, but not until then.
Which Mixtral? What does “heavily quantized” mean?
> A refurbished M1 Max with 32GB of RAM will enable you to generate better quality LLM output than even a 4090 with 24GB of VRAM
Not assuming the 4090 is in a computer with non-trivial system RAM (which it kind of needs to be), since models can be split between GPU and CPU. The M1 Max might have better performance for models that take >24GB and <= 32GB than the 4090 machine with 24GB of VRAM, but assuming the 4090 is in a machine with 16GB+ of system RAM, it will be able to run bigger models than the M2, as well as outperforming it for models requiring up to 24GB of RAM.
The system I actually use is a gaming laptop with a 16GB 3080Ti and 64GB of system RAM, and Mixtral 8x7B @ 4-bit (~30GB RAM) works.
Also as an anecdote, my daily driver machine is a bottom end M2 Mac mini because I am a cheap ass. I paid less than it for the 4070 card in my desktop PC. The M2 Mac does a dehaze from RAW in lightroom in 18 seconds. My 4070 takes 9 seconds. So the GPU is twice as fast but the mac has a whole free computer stuck to it.
I'm getting about 7 tokens per sec for Mistral with the Q6_K on a bog standard Intel i5-11400 desktop with 32G of memory and no discrete GPU (the CPU has Intel UHD Graphics 730 built in). 2 year old low end CPU that goes for, what $150? these days. As far as I'm concerned that's conversational speed. Pop in some 8 core modern CPU and I'm betting you can double that, without even involving any GPU.
People way overestimate what they need in order to play around with models these days. Use llama.cpp and buy that extra $80 worth of RAM and pay about half the price of a comparable Mac all in. Bigger models? Buy more RAM, which is very cheap these days.
There's a $487 special on Newegg today with an i7-12700KF, motherboard and 32G of ram. Add another $300 worth of case, power supply, SSD and more RAM and you're under the price of a Macbook Air. There's your LLM inference machine (not for training obviously) which can run even the 70B models at home at acceptable conversational speed.
I think a way for the m series chips to aggressively target GPU inference or training would need a strategy that increases the speed of the RAM to start to match GDDR6 or HBM3 or use it directly.
And still, performance-wise, the 2070 still wins out by a ~33% margin: https://browser.geekbench.com/opencl-benchmarks
The M3 Max has something like 33% faster overall graphics performance than the M1 Max (average benchmark) while the 4090 is something like 138% faster than the 2080Ti.
Depending on which 2070 and 4070 models you compare the difference is similar, close to or exceeding 100% uplift.
As far as desktop products, power consumption is irrelevant.
But of course that’s only helpful for specific workflows.
No idea what specifically everyone is pulling their performance data from or what task(s).
Here is a video to help visualize the differences with a maxed out m3 max vs 16gbm1 pro vs 4090 on llm 7B/13b/70b llama 2. https://youtu.be/jaM02mb6JFM
Here’s a Reddit comparison of 4090 vs M2 Ultra 96gb with tokens/s
https://old.reddit.com/r/LocalLLaMA/comments/14319ra/rtx_409...
M3 pro memory BW 150 gb/s M3 max 10/30 300 gb/s M3 max 12/40 400 gb/s
“Llama models are mostly limited by memory bandwidth. rtx 3090 has 935.8 gb/s rtx 4090 has 1008 gb/s m2 ultra has 800 gb/s m2 max has 400 gb/s so 4090 is 10% faster for llama inference than 3090 and more than 2x faster than apple m2 max https://github.com/turboderp/exllama using exllama you can get 160 tokens/s in 7b model and 97 tokens/s in 13b model while m2 max has only 40 tokens/s in 7b model and 24 tokens/s in 13b apple 40/s Memory bandwidth cap is also the reason why llamas work so well on cpu (…)
buying second gpu will increase memory capacity to 48gb but has no effect on bandwidth so 2x 4090 will have 48gb vram and 1008 gb/s bandwidth and 50% utilization”
That MacBook has an M3 Max and 64GB RAM.
I'd say it does live up to my expectations, perhaps even slightly exceeds them.