i have an m4 studio with a lot of unified memory and i’m still no where near running a 120b model. i’m at like 30b
apple or nvidia’s going to have to sell 1.5 tb ram machines before benchmark performance is going to be comparable
Plus when you use claude or openai, these days it’s performing google searches etc that my local model isn’t doing.
I'm running a 400B parameter model at FP8 and it still took a lot of post-training to get an even somewhat comparable performance
-
I think a lot of people implicitly bake in some grace because the models are open weights, and that's not unreasonable because of the flexibility... but in terms of raw performance it's not even close.
GPT-3.5 has better world knowledge than some 70B models, and a few even larger.
"the hacker news dream" - a house, 2 kids, and a desktop supercomputer that can run a 700B model.
I'm on a 128GB M4 Max, and running models locally is a curiosity at best given the relative performance.
I'm not saying those machines can't be usefull or fun, but it's not in the range of the 'fantasy' thing you're responding to.
Without constantly refreshing the underlying LLM and the expert system layer, these models would be outdated in months. Language and underlying reality would shift from under their representations and they would rot quick.
That's my reasoning for considering this a bubble. There has been zero indication that the R&D can be frozen. They are stuck burning increasing amouts of cash for as long as they want these models to be relevant and useful.
I'm more than a bit overwhelmed with what I've gotten on my plate and have completely missed the boat on ex. understanding what MLX is, really curious for a thought dump if you have some opinionated experience/thoughts here. (ex. never crossed my mind until now that you might get better results on the NPU than GPU)
I agree with other comments that there are productive uses for them. Just not on the scale of o4-mini/o3/claude 4 sonnet/opus.
So imo open weights larger models from big US labs is a big deal! Glad to see it. Gemma models, for example, are great for their size. They’re just quite small.
I should try Kimi K2 too.
You get the picture. Sure, even last year's local LLM will do well in capable hands in that scenario.
Now try pushing over 100,000 tokens in a single call, every call, in an automated process. I'm talking the type of workflows where you push over a million tokens in a few minutes, over several steps.
That's where the moat, no, the chasm, between local setups and a public API lies.
No one who does serious work "chats" with an LLM. They trigger workflows where "agents" chew on a complex problem for several minutes.
That's where local models fold.
If you asked "What's the best bicycle", most enthusiasts would say one you tried, works for your usecase, etc.
Benchmarks should be for pruning models you try at the absolute highest level, because at the end of the day it's way too easy to hack them without breaking any rules (post-train on the public, generate a ton of synthetic examples, train on those, repeat)
You should also remember that there's no free lunch. If you see models below a certain size fail consistently, don't expect a model that is even smaller to somehow magically succeed, no matter how much pixie dust the developer advertises.
I suppose it's an open question whether there is another free lunch or whether the 30B models in a year will be not much better than our current ones.
Many small models are supposedly good for controlled tasks, but given a detailed prompt, I can't get any of them to follow simple instructions. They usually just regurgitate the examples in the system prompt. Useless.