The current 128GB (e.g. M3 Max) and 192GB (e.g. M2 Ultra) Macs run these large models. For example on the M2 Ultra, the Qwen 110B model, 4-bit quantized, gets almost 10 t/s using Ollama [2] and other tools built with llama.cpp.
There's also the benefit of being able to load different models simultaneously which is becoming important for RAG and agent-related workflows.
[1] https://www.macrumors.com/2024/04/11/m4-ai-chips-late-2024/ [2] https://ollama.com/library/qwen:110b
I consider "heavily" quantized to be anything below 4-bit quantization. At 4-bit, you could run a 110B model on around 55GB to 60GB of memory. Right now, Llama-3-70B-Instruct is the highest ranked model you can download[0], and you should be able to fit the 6-bit quantization into 64GB of RAM. Historically, 4-bit quantization represents very little quality loss compared to the full 16-bit models for LLMs, but I have heard rumors that Llama 3 might be so well trained that the quality loss starts to occur earlier, so 6-bit quantization seems like a safe bet for good quality.
If you had 128GB of RAM, you still couldn't run the unquantized 70B model, but you could run the 8-bit quantization in a little over 70GB of RAM. Which could feel unsatisfying, since you would have so much unused RAM sitting around, and Apple charges a shocking amount of money for RAM.
96GB RAM might be a good compromise for now. 64GB is cutting it close, 128GB leaves more breathing room but is expensive.
As far as I know, the most RAM you can get in a MacBook Pro (which is what you said you're shopping for) is 48G.* The base price for new one with that much unified RAM is around $4000.
The Mac Pro towers (not MacBook) have up to 192G unified RAM. The base price for that configuration is around $8600.
The smaller LLMs are getting quite good. A lightly quantized Llama-8B should comfortably run on a MacBook Pro with with 16G of RAM which you can get for around $2000. The money you save on a cheaper machine will go a very long way renting compute from a datacenter.
If you need to run locally, then high end Macs are excellent machines. Though at those prices you might get better value buying a second hand crypto-mining rig with multiple Nvidia 4090's.
EDIT: I was wrong about the MBP unified RAM. You can get an M3 Max with 128GB for around $4700.
My Macbook Pro has 2.5x that (128GB), and I run models that use 2x that RAM (96GB) with no impact to my IDE, browser, or other apps running at the time, they act like they're on a 32GB machine.
LM Studio makes it easy for newcomers to on-device LLMs to dip your toe into this, both for turning on Metal and helping suggest which models will fit entirely in RAM.
My main point is that if your objective is dipping your toe into this, you can do it with smaller models for far less. That is a really sweet machine, but for the amount of money involved you should be clear about what your needs are.
Huh? They have options up to 128GB…
https://www.apple.com/shop/buy-mac/macbook-pro/14-inch-space...
You can basically just divide by the multiple as you scale up parameters. Since this is all with a 7B model, just multiply memory by 10X and divide speed by 10X . For batch size=1 (single user interactive inference) if you can fit the model you're basically going to be memory bandwidth limited, but pay attention to the "PP" (prompt generation) number - this is the speed for how long it will take to process any existing conversation. If you're 4000 tokens in, and you are prompt processing at 100 tokens/s, that means you will wait for 40 seconds before any text even starts generating for the next turn.
If you're not in a rush, I'd wait for the M4, it's rumored to have much better AI processing (the M3 actually reduced memory bandwidth vs the M2...)
As far as quantifiable results in terms of perplexity go, q4+ quants are generally considered OK. (eg. https://arxiv.org/abs/2212.09720 )
Soldered RAM, no real upgrade path - M2/M3 is cool, but not for this.