A 4-bit MLX quant with 128k window should fit perfectly, in the 50-70 tok/s range.
A 4-bit MLX quant with 128k window should fit perfectly, in the 50-70 tok/s range.
I’m personally considering retiring my MBP for a Studio + 15" Air whenever this MBP ages out.
I’m doing that. Mac Mini M4 Pro with 48G RAM as a headless llama.cpp server.
I much prefer using " thin clients " as the interface to the big VMs running in my homelab
qwen3.5:122b-a10b is significantly faster at around 60-65.
Does it work? yeah... But I'd pick a subscription anyday...
I've seen benchmarks that show 4-5x faster of M5 Max vs. M4 Max.
For local models you're likely using M5 Max, prefill is low thousands of tokens per second, as opposed to, say high hundreds with M4 Max.
For larger dense models, some fraction of that, but similar multiple.