How far are we from running a GPT-3/GPT-4 level LLM on regular consumer hardware, like a MacBook Pro?
How far are we from running a GPT-3/GPT-4 level LLM on regular consumer hardware, like a MacBook Pro?
Llama 3.3 70B and Qwen 2.5 72B are certainly comparable to GPT-4, and they will run on MacBook Pros with at least 64GB of RAM. However, I have an M3 Max and I can’t say that models of this size run at comfortable speeds. They’re a bit sluggish.
https://github.com/ggerganov/llama.cpp/discussions/4167
OMM, Llama 3.3 70B runs at ~7 text generation tokens per second on Macbook Pro Max 128GB, while generating GPT-4 feeling text with more in depth responses and fewer bullets. Llama 3.3 70B also doesn't fight the system prompt, it leans in.
Consider e.g. LM Studio (0.3.5 or newer) for a Metal (MLX) centered UI, include MLX in your search term when downloading models.
Also, do not scrimp on the storage. At 60GB - 100GB per model, it takes a day of experimentation to use 2.5TB of storage in your model cache. And remember to exclude that path from your TimeMachine backups.
The problem is that often you can't run anything else. I've had trouble running larger models in 64GB when I've had a bunch of Firefox and VS Code tabs open at the same time.
Qwen2.5 Coder 14B at a 4 bit quantization could run but you will need to be diligent about what else you have in memory at the same time.
Yes, it is not "local" as I will have to use the internet when not at home. But it will also not drain the battery very quickly when using it, which I suspect would happen to a Macbook Pro running such models. Also 70B models are out of reach of my setup, but I think they are painfully slow on Mac hardware.
If only those models supported anything other than English
All of the Qwen models are basically fluent in both English and Chinese.
Did you mean _external gpu_?
Choose any 12GB or more video card with GDDR6 or superior and you'll have at least double the performance of a base m4 mini.
The base model is almost an older generation. Thunderbolt 4 instead of 5, slower bandwidths, slower SSDs.
For $500 all included?
Here's a config for around the same price. All brand new parts for 573. You can spend the difference improving any part you wish, or maybe get an used 3060 and go AM5 instead (Ryzen 8400F). Both paths are upgradeable.
https://pcpartpicker.com/list/ftK8rM
Double the LLM performance. Half the desktop performance. But you can use both at the same time. Your computer will not slow down when running inference.
You'll need a mini-pc with two M.2 slots, like this:
https://www.amazon.com/Beelink-SER7-7840HS-Computer-Display/...
And a riser like this:
https://www.amazon.com/CERRXIAN-Graphics-Left-PCI-Express-Ex...
And some courage to open it and rig the stuff in.
Then you can plug a GPU on it. It should have decent load times. Better than an eGPU, worse than the AM4 desktop build, fast enough to beat the M4 (once the data is in the GPU, it doesn't matter).
It makes for a very portable setup. I haven't built it, but I think it's a reasonable LLM choice comparable to the M4 in speed and portability while still being upgradable.
Edit: and you'll need an external power supply of at least 400W:)
Phi-4 is yet another step towards a small, open, GPT-4 level model. I think we're getting quite close.
Check the benchmarks comparing to GPT-4o on the first page of their technical report if you haven't already https://arxiv.org/pdf/2412.08905
The Qwen2 models that run on my MacBook Pro are GPT-4 level too.
Some people do place value on running locally, and I'm not against then for it, but realistically no 70B class model has the amount of general knowledge or understanding of nuance as any recent GPT-4 checkpoint.
That being said these models are still very strong compared to what we had a year ago and capable of useful work
At the end of the day, there's only so much you can cram into any given number of parameters, regardless of what any artificial benchmark says.
The comparison is between something you can buy off the shelf like a powerful Mac, vs something powered by a Grace Hopper CPU from Nvidia, which would require both lots of money and a business relationship.
Honestly, people pay $4k for nice TVs, refrigerators and even couches, and those are not professional tools by any stretch. If LLMs needed a $50k Mac Pro with maxed out everything, that might be different. But anything that's a laptop is definitely regular consumer hardware.
If I'm running a software business selling software that runs on 'consumer hardware' the more people can run my software, the more people can pay me. For me, the term means the hardware used by a typical-ish consumer. I'll check the Steam hardware survey, find the 75th-percentile gamer has 8 cores, 32GB RAM, 12GB VRAM - and I'd better make sure my software works on a machine like that.
On the other hand, 'consumer hardware' could also be used to simply mean hardware available off-the-shelf from retailers who sell to consumers. By this definition, 128GB of RAM is 'consumer hardware' even if it only counts as 0.5% in Steam's hardware survey.
For AI and LLMs, I'm not aware of any company even selling the models assets directly to consumers, they're either completely unavailable (OpenAI) or freely licensed so the companies training them aren't really dependent what the average person has for commercial success.
Today, lots of people spend far more than that for gaming PCs. An Alienware R16 (unquestionably a consumer PC) with 64 GB of RAM starts at $4700.
It is an expensive computer, but the best mainstream computers at any particular time have always cost between $2500 and $5000.