Is there a rule-of-thumb estimate for how much RAM this would need to be used locally?
Is the RAM requirement the same for a GPU and "unified" RAM like Apple silicon?
Is there a rule-of-thumb estimate for how much RAM this would need to be used locally?
Is the RAM requirement the same for a GPU and "unified" RAM like Apple silicon?
When the model gets quantized to say 4bit ints, it'll be 22B params * 0.5 bytes = 11GB for example.
B: number of parameters
Q: quantization (16 = no quantization)
Running LLM models on a MacBook Pro with Apple Silicon vs. a PC with an Nvidia 4090 GPU has trade-offs. My 128GB MacBook Pro handles models using up to 96GB of unified memory, running at a little under half the speed of a 4090. If you use a quantized version of full floating point model, you can run the largest open models available.
While the 4090 has 24GB of dedicated memory and higher bandwidth (1000 GB/s vs. 400 GB/s on M3 Max), the Mac’s unified memory system (up to 128GB) is flexible and holds smarter models (8 bit and 6 bit models act still mostly all there, 4 bit is so so, 2 bit is brain damaged).
The M2 Ultra in Mac Studio offers even more (800 GB/s bandwidth and 192GB memory). So, ok, 6 or 8 of 4090 cards or 4 x A6000 cards excels in raw performance, but Apple’s unified memory in a laptop fits in your backback.
It's not clear to me why Macbooks and Mac Studio Ultras with maxed out RAM aren't selling better if you look at the convenience and price relative to model size. Models that fit in one 4090 or even a pair of 4090s are toys compared to what fits on these, so for the big models you're comparing a laptop to a minifridge.
It's a bit slower perhaps than the mac, but i get the best of both worlds. That is I get a lot of RAM to hold the model and I can offload as much of it as possible to the GPU. This works especially well with models like mixtral 8x22, but also models like llama3 and the old large bloom model.
I also get the utility of running Linux instead of the closed up mac os.
But running large models locally is not exclusive to mac studio, you can do the same on PC for a much lower cost.
> closed up MacOS
https://github.com/apple-oss-distributions/distribution-macO...
curl https://alx.sh | sh
https://asahilinux.org/I prefer the "utility" of BSDs, but that's just a preference.
Have you ever seen the inside of a datacenter? Why is it that surprising to you that nobody perks up when you start waxing on about battery life? Even terms of power-to-performance, Apple's latest chips get ethered by Nvidia's server offerings.
This "Apple for Inference" meme is so dead that I can only feel sad when I see people unironically promoting it. You actually think serious customers are going to load up Asahi (even funnier, MacOS) on their Mac Pro... so they can inference half as fast as a single Blackwell GPU? You think the industry is doing this shit? I don't even think the Steve Jobs apologists are dumb enough to fall for this one, you must be a particularly aspirational shareholder.
Aren't these machines extremly expensive and generally not upgradable?
you need enough RAM and HBM (GPU RAM) so it’s a constraint on both.
Though realistically for code completion smaller models will be better due to speed