I've heard arguments both for and against this, but they always lack concrete numbers.
I'd love something like "Here is Qwen2.5 at Q4 quantization running via Ollama + these settings, and M4 24GB RAM gets X tokens/s while RTX 3090ti gets Y tokens/s", otherwise we're just propagating mostly anecdotes without any reality-checks.
Still, it’s way to early and there are simply way to many hardware and software combinations that change almost weekly to establish “the best practice hardware configuration for training / inferencing large language models locally”.
Some day there will be established guides with solid. In fact someday there will be be PC’s that specifically target LLMs and will feature all kinds of stats aimed at getting you to bust out your wallet. And I even predict they’ll come up with metrics that all the players will chase well beyond when those metrics make sense (megapixels, clock frequency, etc)… but we aren’t there yet!
Saying "Apple seems to be somewhat equal to this other setup" doesn't really contribute to someone getting an accurate picture if it is equal or not, unless we start including raw numbers, even if they aren't directly comparable.
I don't think it's too early to say "I get X tokens/second with this setup + these settings" because then we can at least start comparing, instead of just guessing which seems to be the current SOTA.
But you can likely find similar threads for the llama.cpp benchmark here: https://github.com/ggerganov/llama.cpp/tree/master/examples/...
These are good examples because the llama.cpp and whisper.cpp benchmarks take full advantage of the Apple hardware but also take full advantage of non-Apple hardware with GPU support, AVX support etc.
It’s been true for a while now that the memory bandwidth of modern Apple systems in tandem with the neural cores and gpu has made them very competitive Nvidia for local inference and even basic training.
Still, thanks for the links :)
* hardware spec
* inference engine
* specific model - differences to tokenizer will make models faster/slower with equivalent parameter count
* quantization used - and you need to be aware of hardware specific optimizations for particular quants
* kv cache settings
* input context size
* output token count
This is probably not a complete list either.
What's hard about it? You get the hardware, you run the software, you take measurements.
But how are we supposed to get enough people doing those things if everyone say "There isn't enough data right now for it to be useful"? We have to start somewhere
total duration: 24.919887458s
load duration: 39.315083ms
prompt eval count: 37 token(s)
prompt eval duration: 963.071ms
prompt eval rate: 38.42 tokens/s
eval count: 441 token(s)
eval duration: 23.916616s
eval rate: 18.44 tokens/s
I have a gaming PC with a 4090 I could try, but I don't think this model would fitWhat quantization are you using? What's the runtime+version you run this with? And the rest of the settings?
Edit: Turns out parent is using Q4 for their test. Doing the same test with LM Studio and a 3090ti + Ryzen 5950X (with 44 layers on GPU, 2 on CPU) I get ~15 tokens/second.
Only settings I did were the ones shown in the blog post
OLLAMA_FLASH_ATTENTION=1
OLLAMA_KV_CACHE_TYPE=q8_0
Ran the model like ollama run gemma2:27b --verbose
With the same prompt, "Can you write me a story about a tortoise and a hare, but one that involves a race to get the most tokens per second?"Example: the default model weights for Llama 3.3 70b, after hitting the “view all” have this hash and size listed next to it - a6eb4748fd29 • 43GB
Now scroll down through the list and you will find the one that matches that hash and size is “70b-instruct-q4_K_M”. That tells you that the default weights for Llama 3.3 70B from Ollama are 4-bit quantized (q4) while the “K_M” tells you a bit about what techniques were used during quantization to balance size and performance.
total_duration: 10530451000
load_duration: 54350253
prompt_eval_count: 36
prompt_eval_duration: 29000000
prompt_token/s: 1241.38
eval_count: 460
eval_duration: 10445000000
response_token/s: 44.04
Fast prompt eval is important when feeding larger contexts into these models, which is required for almost anything useful. GPUs have other advantages for traditional ML, whisper models, vision, and image generation. There's a lot of flexibility that doesn't really get discussed when folks trot out the 'just buy a mac' line.Anecdotally I can share my revealed preference. I have both an M3 (36gb) as well as a GPU machine, and I went through the trouble of putting my GPU box online because it was so much faster than the mac. And doubling up the GPUs allows me to run models like the deepseek-tuned llama 3.3, with which I have completely replaced my use of chatgpt 4o.
total duration: 10.5922028s
load duration: 21.1739ms
prompt eval count: 36 token(s)
prompt eval duration: 546ms
prompt eval rate: 65.93 tokens/s
eval count: 467 token(s)
eval duration: 10.023s
eval rate: 46.59 tokens/sThe same on Nvidia (various models) https://github.com/ggerganov/llama.cpp/issues/11474
[1] this is a the model: https://huggingface.co/unsloth/DeepSeek-R1-GGUF/tree/main/De...
I'm not sure I'm reading the results wrong or missing some vital context, but that sounds unlikely to me.
Where the a100 and other similar chips dominate is in training &c, which is mostly a question of flops.
I don't think they do.
From Wikipedia:
> the M2 Pro, M2 Max, and M2 Ultra have approximately 200 GB/s, 400 GB/s, and 800 GB/s respectively
From techpowerup:
> NVIDIA A100 SXM4 80 GB - Memory bandwidth - 2.04 TB/s
Seems to be a magnitude of difference, and that's just the bandwidth.
Around half that price tag was attributed to the blogger reusing an old workstation he had lying around. Beyond this point, OP slapped two graphics cards into an old rig. A better description would be something like "what buying two graphics cards gets you in terms of AI".
Meaning what? This is largely what you do on a budget since RAM is such a difference maker in token generation. This is what's recommended. OP could buy an a100, but that wouldn't be a budget build.
Why do you say this? I thought the p40 only had a memory bandwidth of 346 Gbytes/sec. The m4 is 546 GB/s. So the macbook should kick the crap out of the p40.
I've heard that Macs are pretty slow with XL and borderline unusable for flux requiring minutes at a time to generate a single image - whereas an RTX4090 can generate a 1024x1024 image with the higher quality Flux Dev model (not schnell) in 14 seconds.
OP is probably correct that if you want to branch out of just strictly LLM's, cuda is the way to go. I've never heard of anyone getting LTX or hunyuan running on a Mac for example.