Just writing this down so I can be praised/mocked in 5 years.
Just writing this down so I can be praised/mocked in 5 years.
I'm fatigued by it all at this point. It's streamlining the interesting and fun parts out of my job (by practical necessity of use there), and if I used it half as much outside of work I'm sure it'd do the same there too.
This is the prevailing opinion of people even outside of tech.
That said, I think it's a good thing that this sentiment is coming to the forefront.
He was complaining that he would ask how to perform a certain repair on a car, and the LLMs he tried (ChatGPT & Grok) would give him a long involved process and he'd ask why not do it this simpler way and it would say, oh you're right! He just found it gave bad advice and realized (rightly) that in areas he has less expertise in he has no way to judge how good the outputs are.
This is from a guy who loves tech, historically worshipped Elon, loves his Tesla, and (rightfully again) didn't buy into SpaceX because he thought it was overvalued.
In the past when I visited for holidays he was liable to have a positive outlook on LLMs and their utility. Seems telling that he's starting to see the cracks.
This is easily the biggest problem with the current models. The models are just way too eager to please / say yes to the point that the models are happy to lie/make shit up if it means it can say yes.
Interesting. For me it's streamlining the tedious and attentionally taxing parts of my work tasks. I love solving problems, I don't particularly love shaving yaks.
1. Most useful LLM work is done in parallel. A Mac Mini can run one LLM inference thread at a time. The cloud can spool up dozens and spread that inference across efficiently batched operations over a fleet of hardware.
2. Faster inference hardware such as the chips from Cerebras and Groq cannot be run locally. But the advantages of running >5x the token throughput per thread can’t be overstated. Add in the multi-threading advantage and it’s a knock-out punch for local LLMs.
Local inference has a role: if you’re working with extremely private matters or you want an uncapped model that will talk dirty or generate NSFW photos, local is the only option. I think Apple and others will continue to also run a lot of useful workloads locally such as text editing suggestions, speech to text, text to speech, and image manipulation. As local hardware improves, these capabilities will get better too.
But, for most LLM work, the cloud will continue to dominate for a long time to come, if not forever.
There is so much happening in that scene, where tokens/sec double or 10x
So I could see the same hardware doing 20 tokens/sec on a large model suddenly doing 200 tokens/sec in the future, a better device in the future doing 500 tokens/sec, while having vision models baked in, audio models etc
Users wont consciously switch to local, they will just have it and use it
you’re far far far more likely to see a camry or equivalent in an americans driveway than you are a super car.
you’re also more likely to see an enthusiast with a corvette or equivalent muscle car spending way more than it’s worth to tinker on that car in their garage than you are a super car.
That’s not accurate. With MLX, at least, parallel inference is both possible and useful. Model serving tools like LM Studio and oMLX support parallel generation with continuous batching, and the total throughput increases with it.
Can I run a few inferences in parallel on my Mac Mini? Yes. But put 1,000 Mac Minis in a datacenter serving 1,000 copies of myself? That's going to be more efficient.
I guess what I'm doing is not considered that useful then? I usually only have zero, one, or occasionally two things actively doing inference at a time, be it claude code sessions or one of the chatgpt/claude web interfaces, and i bet that's true for like 95% of people using llms. And anyway i bet even the hardcore people using a bunch of parallel agents would appreciate having access to local, private inference for some things.
You're obviously right though that cloud inference isn't going away anytime soon
And yes, I agree. I find the experience better less for speed and more for context management. But it's far from necessary.
There is probably a middle ground though. Something like a main thread agent (say a local model for Hermes) orchestrating cloud based sub agents.
I don't want to run any workflows on someone else's computers.
Well just because they don't sell them. Doesn't mean that will be the case forever.
We are both late and early.
Buy an Nvidia Spark, then whatever cheap Mac you want to use as a thin client. There's no reason to force Apple Silicon's round peg into a square hole like AI inference.
The other half of that equation is latency, predicated on prefill performance which needs a powerful GPU and ideally ALU-level optimization to build larger KV caches quickly. Even the M5 gets smoked in this department, the M5 Max has a 50% longer TTFT on Qwen's 27b dense model at only 16k of context, which is a pretty typical starting context to use for agentic editing in normal apps like OpenCode/Claude Code: https://raw.githubusercontent.com/Osmantic/MMBT-Messy-Model-...
For agentic, 50-256k token on-device coding sessions, the Spark will be faster and consume less power running larger models. Without an external GPU (which Apple doesn't support), Apple Silicon will always be bottlenecked during prefill. Apple's failure to address this with their GPU architecture is a big reason why Apple Silicon viewed as a waste of time and money for professional datacenter deployment.
I keep hearing people make this claim that TTFT is a problem, and… it just isn’t, if you’re running oMLX.
Folks in my camp keep saying this, and folks in your camp keep beating a drum we tell you isn’t resonating. Not sure why I keep bothering to argue; you can’t buy a high RAM Mac Studio like mine anymore.
They're not usable for deployment. They're perfectly fine for "enthusiast" low-end usage with 10-30B models, but the same goes for almost every dGPU made in the last 10 years. Your Mac Studio cannot run frontier LLMs at an interactive speed, even Apple has given up on using it as an inference backbone.
Apple is doing something very different. Their AI experience for end users definitely has been a little behind.
Apple Silicon, however, has been quite unique for the last 4-6 years and it's increasing overlap with LLMS.
The model/chip optimizations are definitely improvements, the thing that is really standing out the past 2 years is how much the open source model community has been making possible, especially when you know a group of use cases.
Apple's shovel (ahem, Mac mini) is the highest quality.with Companies burning money left, right and center, Apple can dispense with advertising altogether
The one thing that is marginally exciting: the Apple SoC or M series chips.
It's unfortunate they are locked behind crappy macOS and other proprietary apple crap.
Unsurprising. Apple seriously thought the iPad would replace computers and usher in a "post-PC" word during their "What is a computer?" ad campaign era. Now they are sticking phone chips in laptop chassis.
...for some users. See their "Mac is a truck" analogy.
And it has. My parents haven't owned a Windows or Mac machine in six years, since they got rid of the one I gifted to them a decade ago. Its all iPad and iPhone.
You think this is a mistake...
Of course. Do you think this was on purpose? All part of Apple's brilliant master plan?