Unified memory is the only reason Macs are so coveted right now for local AI. A single 192 gb ram Mac costs less than the equivalent in standalone GPUs.
Unified memory is the only reason Macs are so coveted right now for local AI. A single 192 gb ram Mac costs less than the equivalent in standalone GPUs.
What are the good use cases for very large memory amounts?
DeepSeek v3 for instance has 671B params, but should have the memory bandwidth of a 37B dense model with a batch size of one.
You could also try Qwen 2.5 32b, which you should just work with ollama or LM Studio with no config changes.
I've got a 32gb M1 Max and a 24gb 4090, and I barely ever run models on my Mac, as the memory bandwidth and compute for prefill is much better on the 4090. But I'm essentially locked out of Llama 3 70B class models, which I only use via API.
[1] See: https://www.reddit.com/r/LocalLLaMA/comments/186phti/m1m2m3_...
I also have a home server with 2x3090 and 2xA4000 (80GB vRAM) - yes it's a lot faster, but it's a pain in the ass to build, it takes up a lot of space, uses 10x the power, and honestly - cost about the same as my MacBook Pro.
The software will still see a single memory pool
I know im disagreeing with the article
I hope Apple sticks with the architecture. Even if its not very practical, its great to have it as possible.