Apple caught off guard by AI demand for Mac Mini and Mac Studio
macrumors.com
macrumors.com
The same thing happened with Mac Mini's and OpenClaw. Nobody cared about or was using Mac Mini's for OpenClaw, but there were all of these very suspicious posts from accounts that were clearly Apple marketing bots (you could tell by looking at their post history, where they would drop "Mac Mini" into every conversation they could across all different subreddits and unrelated topics). Then it became fairly common.
So Apple's marketing department is seemingly using the same strategy again. Because still, it's impossible to find a reputable source for this claim.
I can see them not releasing an M6 Ultra but no reason not to see Pro and Max versions of the M6.
I wouldn't have a problem with Apple fanboys if they just stated that they like Macs, but for some reason you have no problem with just straight up LYING. Or at least being so deluded that you state things that are so easy to prove false as fact.
I.e, 600 gb/sec is dogshit slow compared to vram speeds.
Being able to run frontier models at <10 tok/sec is just an excersize in showing that you can afford a Mac. Its useless for any real tasks. Gemma4 is on par with a lot of the older frontier models like Gemini Flash, and you can run that at 100+ tok/sec on a 2x3090 rig which costs way less than even top of the line Mac Studio.
* I say apparently because The Information wants you to sign up for access to the article, but I found this mentioned in multiple summaries of the article like this one [1].
[1]: https://tech-insider.org/openai-mac-buying-apple-supply-shor...
For mine I was on the fence - waited an hour or so and the delivery date went from early oct to 10-12 weeks so i decided to wait on the 512gb version
People really were talking about and buying Minis due to OpenClaw, because they wanted something that was on all the time and had access to all their stuff in macOS. Occam’s Razor.
If you’ve hot any occam’s razor left in you, consider:
1. Apple is a masterful media manipulator, carefully planting stories like this by gradually seeding fringe media first, then more mainstream, with a clever and subtle goal of making themselves look dumb in order to trick people into, uh, ok, I’m lost here, but something. Maybe drum up demand for products they can’t keep in stock anyway?
2. Media outlets chase clicks, and when fringe media finds a hit, mainstream media follows because you can always write a story about the fact that other people are writing stories. Low cost, proven clicks, nobody yells at you for missing a trend.
Which is more likely? I’ll cheerfully stipulate Apple is pure evil, serves cooked babies in their cafeterias, whatever it takes to discuss based on reason rather than big-evil-corpos.
And, this isn't making them look particularly dumb. It is making them look dumb the way that most people make themselves look "bad" when they list their biggest weakness in a job interview. Silly Apple, not seeing how fantastically great their product is and how much people are dying to have it.
The most serious point in favor of this being organic nonsense is just the sheer lack of any actual evidence that it isn't, but the case is being wildly overstated, in part because we're using the non-standard term "psyop" to describe what I'd consider to be mundane marketing practices.
At this point I feel like there is no reasonable way to draw a conclusion. However, viral marketing like this is common and in fact, incredibly easy. This is done through PR services (and even shadier sources) all the time... I'm a bit surprised people on HN are unfamiliar with how utterly common the practice is, it wasn't new when I first came into contact with the industry over a decade ago. The hallmark of this is very similar articles appearing suddenly through the Internet and in the news. It's often not just random plagiarism. This is why the news is constantly running stories about how suits are back in vogue and other vain crap like that.
It's not a “complicated psy-op”, promoting viral stories is the basics of marketing in the age of social media…
”Oh wow our product is so amazing we can’t keep up with demand you better buy before we run out”
You simply promote or finance marketing materials (blog posts, podcasts, youtube videos) with a clear idea of the kind of people you want to reach, and if your "Persona" (that is, ideal customer or evangelist) model is accurate enough, they will talk endlessly about your product for free without having to directly connect with them.
> Nobody cared about or was using Mac Mini's for OpenClaw, but there were all of these very suspicious posts from accounts that were clearly Apple [...]
It has nothing to do with "big-evil-corpo", it happens at all levels. For some reason you're trying to claim that "big-evil-corpo" are the rare exception, which is ridiculous.
I think the Mac Mini is legit popular and is actually being used for local AI use cases, especially Open Claw.
Beyond that, Apple is at capacity for SoCs and ram, and the Mac Mini is going to have a ceiling of how much of Apple's allotment they are willing to give it.
Most companies have that division. It's called marketing.
I mean, it's pretty much just blind conjecture that they may be marketing in stealth in the first place, so really who knows on that front. But you gotta be smoking something to think Apple doesn't want their low margin products to sell like hot cakes. I mean sure, they'd love even more for their high margin products to sell like hot cakes (like the Mac Minis that are useful for AI like in this story), but the low margin ones literally make up for it in volume so you'd sure hope they sell well.
They are a trillion dollar company, that wasn't by accident. They are marketing and sales monsters. Arguably of the big companies, they are the ones that have mastered consumer perception the best of any company on Earth. They would like all of their products to like crazy always.
Practiced by Apple? Just because someone is doing it doesn’t mean that every company does it.
Even though they can and do pay for the highest-profile placements, they also do shady guerrilla marketing "just in case," because they have a whole lab that just analyzes channel efficiency, and they've discovered that anonymous comments on Reddit in random channels have 8.7x the effectiveness of YouTube ads, especially among the coveted Pale Incel demo.
This is how they got to be a trillion dollar company after all! It was marketing, not product.
[0]: Obligatory https://xkcd.com/386/
As you say, Apple would not be a trillion dollar company without a great product, but you can build something great and also aggressively market it. Nobody is saying they got where they are just with marketing.
Oh shit, I didn't know Marketing exists! Different thing from what OP described.
The smart companies out source their gorilla marketing to a third party, having realized their own inadequacies.
Wake me up when an extremely capitalistic company, in the past caught with dirty hands on numerous topics, will reject increasing massively a legal revenue stream with no strings attached. Right, exactly how to become a trillion dollar company and keep it.
I grok that specifically for apple their echo chambers do majority of certain work themselves, but claiming with 100% confidence thats all is a bit of a stretch. If I was a personal close friend with their head of marketing, I would maybe believe that. In any other case, hard logic prevails over anonymous online claims.
Can confirm though, being acquainted with PR people in the field for >20 years; everything on TV is scripted these days.
Chatted with some of my PR friends coworkers over the years at Christmas parties and the like. They claimed stunts during MTV specials, other "candid" moments from the last 20-25 well known to pop culture were actually organized by a PR company they worked at (never that they worked on that project; such work was reserved for whatever PR rockstar PR firm hired)
But all this too could be PR sales pitch to impress; society is full of hearsay to titillate
https://www.forbes.com/sites/jaymcgregor/2017/02/20/reddit-i...
“I have worked over 100 of these kinds of campaigns and never had it come back on the client. I’ve been doing viral marketing and reputation management since 2005. =In the past year I’ve worked for a major entertainment network to magnify a rumor within sports entertainment, as well as damage control on a rumor that came out of an actor being hired on a film before the production company was ready to announce that casting.”
Thus this could be Apple marketing, or it could be someone independent.
But not for such niche products and such niche use cases.
To promote the latest airpods or iphone, yes.
No, Apple is not content with however much it makes from the iPhone. That’s why it stuffed cringeworthy ads in Apple News, filled the App Store screen with ads and is now doing the same with Apple Maps. It’s a company focused on making money through every avenue possible, even if it means being shortsighted and losing the goodwill of developers and users.
That… isn't how capitalism works…
I don’t think there’s a direct guerilla marketing division, but I would absolutely bet money that Apple’s marketing/comms team is encouraging creators, influencers, etc to promote Mac Mini and Mac Studio for AI.
And when those teams brief and speak to people, even on confidential terms, they almost expect some things to get leaked; or will keep sharing it with more people until a third party leaks what they want leaked.
So they may seed things like “super confidential, we were internally really surprised by the demand for AI”. Intentionally.
Why would they produce "low-end, low-margin Macs" then?
This sounds exactly like the kind thing that someone running a psyop for Apple would say.
Honestly, I'm an Apple user since 1988 and shareholder. I use my full-ass name on hn. I wish they'd give me money to promote the brand, but in reality they have something even more valuable than paid shitposting: authentic word of mouth and devoted users. That comes out of the product, not the marketing department, and it takes years.
This sounds exactly like the kind thing that someone running a psyop for Apple would say.
You seem to have product marketing all figured out!
I think you're radically misunderestimating the power of the retail front lines. Folks leave the Guru Bar with what they see as 'bleeding edge ways to use Apple tech', and Apple know it ..
Every large corporation has a psyops division. It's called a marketing department.
Of course not, they outsource that.
Wouldn’t put it past them to pay strategic influencers to feed social media fake information to generate hype. As a way to make the leak ecosystem work in benefit of the marketing department (or psyops division if you want to call it that).
No, but they have marketing people, and marketing people in my experience (even in "big" corporate America) have personal (anonymized, but just created themselves) accounts they do stuff like this on all the time.
Your argument is basically: Why would they try and make more money when they can just sit back and get rich off iPhone's.
That's certainly not a tenable position.
This is absolutely how how corporate America works. I don't know about Apple specifically, but most in companies with a marketing departement, you can bet eating your hat that they do.
None shipped anything of value. Same with anyone claiming to be running dozens or hundreds of agents 24/4.
He started that whole trend!
Also the demand is really there. Or do you think Apple just removed certain options from sale for fun in stead of RAM being out of stock?
This is rather revisionist. A lot of the early OpenClaw publicity was about connecting it to iMessage (which was really only possible on Mac hardware at the time)
People don't fact check in the first place, they don't check that facts came from a trusted source, they don't check if information has changed the last time they got it from a trusted source, etc.
Is it just me or is it scary that it seems to me as if the average person just takes everyone's word for most things, as long as the person telling them is a "trusted source" based only on tribal lines?
I'm convinced that this is just guerilla marketing from Apple.
I'm not. I have an M4 Mac Mini running 24/7 so I can use it as my Claude Code machine when I'm on the go.Why would they do a guerilla marketing campaign now if their Macs are generally sold out despite nearly doubling the Mac Mini price in the last few months?
The former honestly sounds like a nightmare because the Mac Mini can run some big models but not that big, and Claude Code will guaranteed stuff your context with 35k tokens, and all the Anthropic proprietary tool-deferral and tool-search stuff is meant to be handled inside the inference engine in a way that I don't think any open-source engine like vLLM can support right now (much as I wish it did). So you have a huge context on startup, broken functionality, and an underpowered LLM struggling through it all. Not to mention I have no idea how any of the sidecar LLM stuff is meant to work in that context (e.g. auto-mode classifier and even things like dynamic session naming). Am I missing something?
I don't want to assume it's the latter, I want to be charitable to other people. But.
Either they're running Claude Code with a local LLM (technically possible via dev tools albeit unsupported), or they genuinely have no idea why the Mac is good for AI and just bought it because they were told it was good for AI.
Neither. I use Claude Code subscription through my Mac Mini when I’m on the go.I’ve found a new use for it which I’m assuming others have as well.
You could make your case easily if you had some links to these.
(I understand though it's not likely something one remembers at the time to cache away. Still, without that I have little to go on.)
can you approximate plausibility by the wait time for a maxed-out config?
I follow a lot of people on Twitter who were one-upping each other about how much hardware they had bought (now gathering dust0 to run their Claw assistants. Those same people are now buying some toy robot duck?
I realize I’m somewhat limited (16GB RX 9070), but still, it seems really far off from the kind of experience even a basic $20/month subscription gets me.
Any tips anyone might have are appreciated! I’d love to be local first and would be willing to buy hardware to get there.
The $20/month subs are much stronger than the local models you can run, even with how far local models have advanced lately.
The appeal of local models is that the data never leaves your network so you can feel safer putting sensitive content into it. It also feels “free” to use when you’ve already paid for the hardware.
But it doesn’t perform better and if you do the math you’re probably not saving money either. It’s helpful for things that you can’t or don’t want to outsource to a 3rd party.
IDK if that might be a concern for Apple or their AI partners.
There's cloud hosts out there with unused compute. Wasted cycles are all over businesses running unused cloud apps and subscribed to services they don't use.
Still need a local computer to access the cloud; so a barely used gadget still exists. And this creates duplication of effort; we built RAM for servers AND the edge devices.
Seems redundant when tech nerds and corporations are really the only people that care.
And all that data in the web is meaningless yet we create a supply crunch storing it in servers.
This an out of touch nickel and dime perspective given the big picture to say nothing of the mess of strip mining and manufacturing pipelines that go into every screw, wire, and such
All of this is private, but not necessarily sensitive. You never know what is happening with this data. They might say they don't log it or don't sell it, then few years later you'll find it all online or read a book that has a story eerily similar to what you chatted about with GPT a year ago.
- translations: cloud providers can bowdlerize (censor) bad words/content; also, if you want to do a translation for personal use of copyrighted materials, cloud providers may block it
- image generation: generating drawings with a style that even just resembles a copyrighted one (ie. Disney) may be blocked by cloud providers - for example, generating old cartoons style with GPT may not be possible.
In _a_ cloud: their own virtual private cloud. They also have enough power to negotiate contracts with strong privacy provisions.
What privacy provisions would you want to add to AWS? Most of the reasonable strong privacy provisions you'd want are already there and/or available if you want to sign up for it, even including US Govt Top Secret data if you meet some approval.
If they violate their privacy agreement with me there's effectively no punishment I can get that they will actually feel. There would have to be a class action law suit and they always just settle those for some small amount and admit to no wrong doing
To be honest, they'll just send a few million and some gold plated trinkets to the right people and it'll go away. You know that's how things work now in the US. Even if you're an actual convicted drug dealer you can have your mother say a few nice things and whoop there's your get out of jail free card. This stuff is peanuts compared to that.
If not they'll just stall in court and settle for peanuts after everyone is exhausted battling it out. That's what they've always done before the current government made it even easier. Fines are not a penalty, just a business expense.
Nobody important is going to do any time for this. Not in this America.
I know what I'm talking about here - trust me when I say that the difference between a "government" or "large corpo" account and "some dude" is literally a few flags to prevent resources being created in the wrong regions. It's nothing that a random customer can't do themselves.
- I don't trust them with my data
- I don't trust the data they send back to me
The first is about my privacy, the second is about not being manipulated by what is considered the politically correct answer to my questions.
but depending on what you're doing, you may not need the "bleeding edge" performance
In my opinion, 98% of the work most devs would send to an AI can be capably achieved with a local model and a frontier-level model is overkill.
The goalpost moving feeds right into Anthropic and OpenAI's interests.
It is really a once in a lifetime deal.
Once the deal is over the local models will be better than what I am using now anyway and the hardware will be all the better than what I can get now for the price.
I've been running into annoying limits with Claude recently. It gives me like 5 questions over the course of 15 mins and then tells me to wait 5 hours. When companies can change things up to make the base subscription nearly useless (the last question always gets messed up, too), then you realize the value of owning your own infrastructure.
$200/month is vastly cheaper than owning and operating comparable hardware.
And all this before you get into privacy/security/compliance stuff.
What we're paying now is just akin to the drug dealer's free starting shot.
I started tasking fable with huge projects over the weekend and now I hit at very least fable limit by monday.
My understanding would be that if you’re interested in this sort of card for AI that you should go with the AI PRO R9700, which is basically the professional version of the RX 9070XT but with 32GB of memory.
It’s significantly more money but not crazy like a 5090.
I just happen to have the 9070XT primarily for gaming purposes.
I’m not quite sure how to describe my experience using it other than “rudimentary,” and a lot of that is on me for not really understanding the best way to set it up.
So yes, they are genuinely very useful, but they are not yet a full replacement unless you have more powerful hardware and or don't need more intelligent ai.
It's still frustrating as hell to come down in the morning, having given it a list of tasks to do overnight, with tests to pass before they're "done" and find that it worked for about 20 minutes after I went to bed, and decided that it would stop at "3am" (it wasn't) and "not do significant work this at this late hour". Like WTF ? You're an LLM. You don't sleep.
Bloody training data full of humans demanding sleep. I tells ya...
I think that's Anthropic trying to get you to not extract as much value out of that subsidized subscription as possible.
Is this Claude code? Or your local? I assume Claude? I'm more than a little staggered by this, like, it makes no sense! It doesn't even serve Anthropic's interests (surely better for them if it burns your token quota so you have to buy more the next morning.) The LLM just... decided? I'd be so mad.
WTF indeed. Can one even file bugs?
Parent mentioned their $20/month subscription. It's definitely in Anthropic's interests for you to not use it.
Local llms don't suffer from cloud availability issues. Anyone that used Google models know that sometimes they just don't have capacity whatsoever, at least that was the state of things some months back when I used them. Just bear in mind if needed, cloud providers will prioritise API and corporate customers over subscriptions if availability degrades more.
Also they don't have the same guardrails as the other models, so for hacking, reverse engineering and black coding (piracy etc...) these local models might be the only options.
But currently it's really hard to beat anything offered by the cloud companies. And the cost and complexity of setting it all up, just to barely (if at all) touch on Opus-level intelligence makes it seem like we're not quite there for the common man (enthusiasts are a different story.)
I am very excited for open source local models, and we're nearly there, but it's still too complex and expensive to be my daily driver (yet).
What you'll learn pretty quickly from said engineering is that there's a lot more to a good LLM than just the weights themselves. You need a good search provider (also self-hostable, but sounds easier than it really is). You need (well, it's debatable) a memory system. You need a good system for up-to-date library references like a Context7 (also self-hostable but the options are surprisingly not that good). You need a good set of specialized subagents that can perform various tasks well -- for the sake of "doing things well" but also managing context efficiently.
When you've got all that, local models can be _extremely_ useful. But there's one other important thing and that's decent hardware, unfortunately. A lot of people try out local models using small consumer GPUs or Macs and are rightfully unimpressed with the performance. And if the performance doesn't get them, usually they have expectations that they'll perform at Claude levels out of the box. Getting in that neighborhood, like I said, definitely requires some work.
I keep hoping that one day some comment is going to paste a link to some kind of idiot-proof guide or piece of software that’s “90% as good as Claude but running local.”
And by 90% I don’t mean that the model is 90% as good or runs 90% as fast, more like all the other stuff you mentioned is set up out of the box.
Though keep in mind not being beholden to shenanigans from said cloud companies (and interference from government entities!) is definitely worth something intangible.
I just ordered a new Mac Studio M5 Max 128GB $5899 ($6400 with tax) to be able to run the bigger "consumer size" models in the 70B parameter range (~96 GB). That said, I have no illusions that this expensive setup with a Qwen Flash coding LLM will be comparable to a $20/month subscription. Even upgrading to an even more expensive Mac Ultra 256GB for $10000 to hold a bigger model still won't be comparable. Apple hasn't shipped my Mac yet and I'm still considering cancelling it and downgrading to a smaller 64GB RAM config ($4299) to save $1600.
Why did I initially spend the extra $1600 if I knew ahead of time that it wasn't as good as cloud AI? Because I thought I could use some local LLM for the easy tasks or when I hit cloud rate limits. No issues with privacy so that wasn't part of the motivation at all. I just wanted some local AI capability to augment a subscription. I've not totally convinced myself of the cost/benefit of this.
Based on today's consumer hardware landscape, you're paying very high prices for crippled capability compared to the cloud AI subscriptions. We're also in a transition period where the next iteration of hardware improvements have some compelling features for local AI. Apple's upcoming M7 (2027 or 2028) is anticipated to have better GPU and neural engine to help with prefill TTFT. AMD Strix Halo is about to release 192GB system which is a big upgrade to their current 128GB ai pc. Maybe apply my $1600 savings towards those newer products. Those future products will still be very expensive but maybe the cost/benefit will be better.
The maths don't check. With Deepseek Flash one goes a very long way with 1600$ - even 10$/month, for easy jobs, are more than 13 years, and at a higher quality.
The low hanging fruit stuff for me is more something I use it for because I have the local LLM setup running anyway. It wasn't the reason I bought it, but now that it's there I might just as well use it as much as I can.
No, this is a big misconception, and part of the cargo cult.
Use cases like the parent's are essentially about having a local LLM handle the leftover tasks. By that point, a lot of bits (main/big tasks) have already left home anyway.
- There's no guarantee of the $20/month service, and it likely has some limits compared to dedicated hardware token wise.
- Model are becoming more and more efficient, in many cases an M1 Max Mac Studio is still capable with 32 GB. 128 GB ram may not be the necessary baseline.
- Folks may think they want to only have a general model running locally (it's the comparable after all from the cloud providers), but we have to remember if the tasks we're trying to do ultimately are more specific than general and if there's space for the smaller models to do that.
Big upgrade to memory capacity but memory speed is only going up by a few percent, so its still going to be slow with more than a few B active params (I have one)
I think that some of the hardware design folks have been blindsided by AI demand and we haven’t really gotten that next generation AI hardware yet, to the point where buying M5 isn’t going to make sense in a couple of years.
Rumors seem to be that the M7 is the generation that Apple is looking to push AI performance much further.
I’m not sure that Apple anticipated this specific route that computer hardware has gone and I don’t think M5 and previous iterations were really specifically architected for local AI performance, more like they happened to be pretty good at it.
What specifically does that mean? The M5 series has 10 cores per 128 bits of memory bus; are the cores unable to keep up with the RAM? I thought they did and memory bandwidth was usually the bottleneck. But the memory bus is already very highly clocked and goes up to 1024 bits wide so it's hard to picture memory bandwidth having a huge leap.
[1] "LPDDR6 is coming." - https://news.ycombinator.com/item?id=49436849
50% more pins... we'll see. If they can reasonably make that fit then even more shame upon the traditional desktop CPU makers for sticking with 128 bits for so long.
I have been eyeing a 512 GB Mac 5 Ultra to run full DS4 pro locally, which I expect would be pretty amazing as far as quality/recall. The only downside is that the speed is a lot slower than something like 27B on the 5090.
What I noticed is that (1) the great local models are optimized run inference (diffusion & LLMs) well on 32GB VRAM <= GPU's because that that's what the target has ...
(2) The quality of local models (esp. in diffusion) is increasing faster than the need for more VRAM - additional reason for the value of these FAST GPUs to increase!
(3) RTX PRO 6000 96GB is really great for fine tunes (ai-toolkit) :) but doesn't outperform my RTX 5090 with inference by anything significant on the good local models.
I have never run an AI job on a Mac, i also have doubts about performance and compatibilities - since the reviews almost never compare directly.
What's an RTX 9070? Do you mean the RX 9070 or RTX 5070?
We also built some QA agents that are always playing our games from the same builds a player would and flagging things to fix/improve; that alone needs the game focused and front-and-center so it can properly screen-capture for deciding what inputs to take next (and for screenshots/replays), which also means we can't really do any hands-on work at all on the machine when it's running.
Having a separate (and tiny) machine for all of this has been great. We don't bother with local models because, you're right, the $20/month sub is way better than anything that can run on small consumer hardware atm.
I'm curious about your setup. I've been tinkering with the idea of setting up Blender (cli use) in a container to allow agents to verify the scripts they are generating compile at a minimum. One thing I've found extremely helpful was generating a RAG of the current version of Blender.
For anyone wondering, I'm running Gemma4 26b A4B on a mini PC with 32 GB of DD4 and a Vega 7 iGPU (llama.cpp w/ Vulkan).
That said I have an RTX 5090, not a Mac Mini, so it's not exactly the same level of performance... The latest open models run at 200 tpm at around 30B params.
I have an older M2 Mac mini that does the OCR and visual description of all my screenshots. Screenshots are stored on my NAS.
I like to screenshot things as a quick way to remember. They are things that I would not be comfortable sending a cloud provider (customer data, prototype screenshots, bank dispute details).
It runs Qwen3.5:9b and glm5.2-ocr with Ollama and uses about 10GB of RAM. It automatically releases the models from RAM after 5 minutes of inactivity so it is pretty seamless to leave running in the background.
All the details are stored in a simple webapp with a SQLite db that I can search through.
> They are things that I would not be comfortable sending a cloud provider
It's also an old machine that the commenter already has; it's intellectually dishonest to compare it to the price of a brand new, 4-iteration-newer machine.
Neither openai or xai. Anthropic maybe but not likely. Mistral is the most likely one because they're under the EU laws but I think they focus more on commercial these days.
a) model I pick will not 'suddenly' go away
b) I am sure my data stays where I want it
c) my inference mac can run other things if I need to
I pay for that.
Yes, of course the most economical path is to hand over all your data and become fully dependent on a cloud provider who is already operating as scale, hoping that they won't change/remove models, hamstring capabilities, or raise prices.
If this were a thread about hosting your own email or blog or cloud photos, you'd have plenty of people out here telling you how easy it is to do it yourself instead of relying on Gmail for email or WordPress/Medium/Substack for blogging, or iCloud for cloud photos.
And yet, without fail, every single thread about self hosting local models seems to have some copy/paste form of this cost-savings argument.
Where is the appreciation for this cool thing GP built? Where is the appreciation for the desire to figure out how to host your own version of the incredible capabilities that were not available merely a few years ago? And why, on this site of all places, would someone advocate trading all of the knowledge and independence gained from learning how to host something like this ourselves in favor of throwing it all over the wall to Google?
Come on.
It will be horrible to be dependent on an AI who is also be trying to sell you various goods and services.
We're going to need AI whose loyalty is to us and only us.
Because there are more privacy guarantees there, depending on the provider. "But what if they violate their contract!" is some pretty tin-foil hat stuff.
How is this any different than a business running their website out of the cloud, assuming you are using a provider with appropriate contractual terms?
You can care about tracking and ads but still be comfortable storing your backups in the cloud, and many have been for quite awhile, even sometimes without encryption - that is totally different than e.g. Meta actively trying to track you and understand your relationship graph and your purchases etc.
But what if they violate their contract!" is some pretty tin-foil hat stuff.
The foundation of these businesses is stealing IP in bulk.So don't use them for inference.
I’m pretty sure that all of those disclaimers that all the AI model makers have for you to sign off on to say that they’re not responsible for anything that might go wrong if your work gets copied accidentally and used someplace else wink wink?
You know the lawsuits for that particular aspect are incoming in the future…
But no way whatsoever I'm uploading my personal files, photos, emails, chats into cloud AI. No way.
There is no reason to believe that equivalent level model output will be more expensive in 12 months, let alone almost 4 years from now.
Of all the good reasons to use local AI (privacy, etc), worrying about not having access to cheap models in 4 years is not one of them.
It's almost never a drop-in replacement, and having to check and adjust integrations and workflows with new models gets old fast. My task was perfectly solved by the old model, I don't need a newer, "better" one - especially at higher prices ("more cost-effective" my foot). Local models lets one choose a model and freeze the downstream integrations forever, without being forced on the 6/8-month upgrade treadmill by aggressively short, scarcity-driven hosted model-deprecation schedules.
The frontier labs have very high prices for inference. The prices are actually going down, not up.
Others running open models in the cloud is a nice third alternative, but not a solution for those that need frontier models, which I believe continue to need training as well as inference. These are currently heavily subsidized and hardware constrained several years out.
And this sub-discussion is more about whether there will continue to be some cloud model that's both more capable and cheaper than local, not so much about what kind of cloud model it is.
Don't get me wrong: I hope you are right, and I am generally optimistic about the future of AI.
But do you really think, in a world filled with examples of big software companies repeatedly taking away or hamstringing capabilities we've taken for granted, that you can just count on a big tech company hosting cheap inference on incredibly powerful models forever? Surely we have learned by now that these companies do not exist to provide a public service to us, and the government cannot always be counted on to have the best interests of the citizens in mind.
I mean, how many times have we seen this in just the past decade or two?
- Consistent attempts to pass legislation weakening or banning the use of encryption
- Exorbitant Reddit API pricing (still salty about the death of the amazing Apollo app)
- Google fighting against sideloading on android
- US gov't issuing export control directive to suspend access to Fable/Mythos
- US lawmakers considering ways to regulate adoption of open weight models
- Chinese officials considering restricting overseas access to their most advanced models
I can absolutely see a much more restricted, closed down, and expensive future due to a combination of government regulations (regardless of which nation is doing it) and big companies rug-pulling as the check comes due on all the billions of dollars spent to get here.
I'm not saying that local AI is necessarily economically justified right now; but it's certainly quite reasonable to think that in 4 years the subscriptions will be significantly more expensive.
In general, I'm a big believer in doing more with fewer resources, within reason, and think having local setups really helps me be mindful with what's happening under the hood with these systems and managing context efficiently to get high quality results.
A $20/month Gemini subscription is truly all you need, then yeah, sure.... obviously a homelab setup is a ridiculous alternative on a pure cost basis. For most people doing "real" work with LLMs 40+ hours per week, a more apt comparison would be one or multiple $200/month subscriptions. At which point the break-even point of a homelab is much sooner.
However, most people running homelabs are doing it for other reasons. Independence, learning, and/or privacy issues.
- Is that even worth the electricity price compared to api? - We don't ask that here
I have not spun up my own homelab, but, I have researched it and the electricity costs can be mitigated to a large extent.
One... the GPUs can be massively clocked down during idle state, to the point where the fans can be shut off as well. The machines themselves can be shut down and wait for a magic wake-on-LAN packet if needed.
Two... even under load, the GPU cores can be significantly underclocked with very little performance loss. The GPU and VRAM/HBM clocks are independent, and the GPU is largely bottlenecked on VRAM/HBM so you can just drop the GPU speed. A commonly reported figure I saw was, basically, 350W nominal cards being underclocked to consume "only" 200W under load with ~5-10% perf loss.
This all assumes you're using discrete GPUs and not AIO systems like a Mac Studio which is going to be pretty efficient just by design; they idle at 35W or so and in practice max out at a few hundred watts. I believe DGX Spark and Strix Halo are similar.
I must stress that this is second-hand anecdata here, admittedly, but my understanding is that it can be pretty manageable.
Doesn't Apple do this already within it's OS all locally? It certainly does it for OCR and categorization.
EDIT: Also, no reason to use a generic LLM for this. This functionality exists in something like Immich (both OCR and 'context categorization'), and doesn't tie you into the Apple ecosystem either.
Works out really well.
You can grab the code here: https://github.com/mcotton/listing
I run it in Docker on my laptop or Synology NAS. It uses the OpenAI/Ollama URL schema for processing.
For OpenAI and Anthropic, the $100 subscriptions cost 5x the $20 subscriptions and give you 5x the tokens. And the $200 subscriptions are 10x the cost for 20x the tokens. (Tokens cost 50% as much.)
I've spent $2 in the last 2 weeks on OpenRouter. I've been trying to only use the medium sized models that I would otherwise be able to run on a nice local setup. That nice local setup would cost ~$4k. I don't know what the operating cost would be, but I would be concerned that my home electricity would cost more than at a datacenter. It just doesn't make sense right now except for privacy reasons.
I'm probably going to hoarde open weights models in the ~31B range until memory costs fall in a few years. Then, I'll buy some hardware to run at home just so I feel more sovereign over my stack regardless the cost/token speed.
But I am looking forward to lower hardware costs!
I realistically costs $5-10k to replicate a ChatGPT like agent. And it doesn't scale.
That's still really close. And models and quantization etc keep improving.
I'm absolutely positive that I'll be switching to mostly local AI in the next 5 years.
If you have a real product and can actually sell it, youre taking a largish risk relying on the cloud.
From model changes, alignment, to enshittification and the natural cognitive offloading, you could be one day removed and ROI tanked.
Think of AI like a mafia boss who helpfully supports you untill they need a favor. Thats all cloud AI is in America.
Fast 150+ tg and 2-8K pp Qwen 3.8 27B nvfp4 is about 8K (5090 +PC) Gives really only one concurrent session that flies because kv caching is not perfect for ninfer https://github.com/Neroued/ninfer
both are very serviceable, I prefer FP8 on 2xR9700
But, yes it doesn't scale that well but in 5 years the same hardware should still be very capable of running some great MoE models, for example Qwen 3.6 35BA3B on 5090 can fly at 600 tg
E.g., the M7 chip is rumored to be the one where Apple has poured really serious effort into local AI performance where previous generations seem to have mostly been coincidentally good at it.
Maybe this is an incorrect opinion but I don’t personally think that the M1-M3 or maybe M4 or even M5 chips were designed with LLM inference in mind at all. These were designed with things like video rendering, image/video ML, and rasterization performance in mind.
I recently built a minimal Dark Software Factory out of an N150 Mini PC. It uses three models; Sonnit, Sol, and Gemma.
But, I have a LOT of instructions about how I prefer the software it builds. Gemma doesn’t handle all my instructions very well. But it’s close!
I’m running gemma-4-12b because I have limited RAM and larger models were too slow.
I do two types of jobs: planning and prototyping. It has done fine at some of my planning rounds.
I still consider it experimental and don’t use it a lot but I think we’re getting there.
The model I’m specifically targeting to use at high speeds is Qwen 3.8 27b @q4ks. This model actually proved to be good at coding (it sits somewhere between Sonnet 5 and Opus 5 capability). M1 got 10 tok/s, Ryzen Halo 20tok/s, and Radeon 7900 XTX 50tok/s (can only do 128k context window in Radeon card).
The prefill gets extremely slow around 50k tokens in context window (whatever prompt processing stage entails could be wrong about phases here). It takes about 2 hours to fill the context.
Even with a drafter model intended for speed instead of mtp, I can’t get past 70tok/s, still is extremely slow to process prompts as context grows, and drops down to 40-50tok/s anyway making this config still moot for improvement on my Radeon card.
The only thing I can point to slowing me down is bandwidth of the card itself.
I am waiting to actually get my 5090 right now and I am betting that the 1700 Gbps of capacity will fix my prompt processing speeds. I don’t need full PCIe lane bandwidth to serve my house I just need to load the full model into vRAM and let the GPU do its thing.
Additional benefit to the external enclosure route is being able to migrate the inference between devices more easily. I can develop out the infrastructure then migrate the card to be hooked up to a shared node in the house with all the tools necessary for my family to take advantage of the privacy enhancement that comes with local inference.
Prior to two weeks ago, I was just using Pi and Ollama.
I have tried my hand at putting together a few harnesses and I finally landed on what I like. Been working on this small app to handle running llama-server for me from any device that has the llama-cpp stack setup: https://github.com/SamInTheShell/loom
Qwen 3.8 is the first model I've been using that hasn't been having issues doing edit calls. Here are my llama server settings and GUFF that I use: https://gist.github.com/SamInTheShell/0bf838e8dc5093583b688e...
Initial results boiled down as follows.
# lmstudio-community/qwen3.8-27b@q4_k_m decode falloff 104.3 tok/s @ 12,683 -> 55.8 tok/s @ 240,755 (53% retained) prefill falloff 3,274 tok/s -> 1,059 tok/s (32% retained)
# qwen3_8_27b_nvfp4.ninfer decode falloff 173.3 tok/s @ 11,867 -> 139.3 tok/s @ 225,710 (80% retained) prefill falloff 8,726 tok/s -> 2,816 tok/s (32% retained)
I should still have room for more performance on the table. I've not even touched the overclock settings on the GPU.
This is a really cool project, I'm going to have to get into what those 3 guys are doing... assuming it can be done with what I got.
I'm not surprised at all.
Context: I have a farm of DGX Sparks and several RTX 6000's, and can run very close to foundational models with ~2 sparks
So if I stay within 35B, especially MOE, my M5 Pro 64GB MBP can also run them well, and it can do plenty of other stuff too including gaming. While 256 GB with such RAM bandwidth and powerful GPU sounds like fun on paper, it doesn’t seem to be the next level compared to 64 GB
Really curious what people run on 256 GB Macs
But in the meantime I get dishes done, vacuum, flip laundry... etc etc
Frontier models also seem in such a rush to emit anything they produce a mess that needs steering all day anyway
While I have not tested it, it feels like my local setup going slower is better at producing code that works the first time as its not trying to look fast for marketing sake
if they are a heavy user, perhaps they string 4x together.
LLMs are harder: not much useful below 12B, and the 700B+ ones are really much better. Models like Qwen 3.8 27b show promise: in a few years pretty good local AI should be in reach for anyone willing to buy a $1000 computer (but who knows what your $20 sub buys you then).
My Mac can do ~200x realtime (1 hour takes 20s or so). I can do several thousand hours per day. Its pretty incredible
Not sure how much that qualifies as AI vs LLM usage, but it seems to work pretty good
Check out the fluid-ml library which packages this up for ANE very nicely.
Of course, 131k context at 4-bit quant is a trade off, but even then, it's VERY capable. It doesn't feel that far behind something like GPT 5.6 Luna.
Local coding is a step down but good enough for a lot of things if you have privacy concerns.
Wild oversimplification, and benchmarks vary widely, but I've read a lot of benchmarks suggesting that Qwen3.8-27B (xhigh effort) competes with near-frontier models at a lot of coding tasks. To the best of my understanding it's not going to run very feasibly in 16GB of VRAM at usable quants however.
r/LocalLLM and r/LocalLlama are noisy, but valuable sources of anecdata if you have the time (or the tokens, hah) to comb through them. You are going to see a lot of modest setups there, and also guys with $20K+ of hardware.
The two things (besides my bank account) that keep me from investing heavily in local are (1) we are not guaranteed to get a steady release of open models in the future (2) a lot of the "fun" stuff LLM stuff that interests me involves orchestrating lots of parallel agents, which of course multiples the hardware you need to achieve it.
For example, I've been having good results having both Sol and Opus review the same PR, and then I have them cross-review each others' PRs. A next step I'd like to consider is maybe having a swarm of Luna agents review the same PR and have them fight it out... maybe with Sol doing final arbitration? I suspect 5-10 Lunas might outperform a single Opus. Or maybe not. But at any rate, that would be impractical in a homelab without a pretty big hardware (or time) budget.
It's about not having to accept any of the stupid "terms" of the corporations. It's about doing things the big labs don't allow you to do, like cybersecurity stuff, or even just chatting with the AI about some wrongthink.
I don't really see how they couldn't see the Local AI demand or demand for a cheaper Macbooks. This just reads like marketing imo.
this article's also about enterprise demand specifically. That's a bit surprising to me as well frankly. I'd have thought the primary market for mac studios would be hobbyists/enthusiasts with a bunch of disposable income who are willing to pay 18k for a 512 gb machine to run glm 3.5 flash or 9k to run deepseek v4 flash locally. It's competing with a $200/mo subscription or renting server gpu time for open source models during a memory shortage - and idk if it's going to be powerful enough to train or fine tune so it's really just inference. seems reasonable to be surprised
A big enterprise can drop one (or 4) of these on someone's desk and let them go nuts.
The target audience didn't evaluate that as a limitation - the majority of the market for Apple devices does trust that they will not produce and sell a computer incapable of support their use.
Professionals know there are tasks that a baseline computer cannot handle, and even common tasks that a more powerful computer does better, but those people weren't really the target demo for the laptop.
Well, kinda the point. We know this to be fully true in retrospect (and many people correctly predicted it) but decent arguments existed against it at time of release.
the problem with that argument is that the vision pro exists, where they clearly overestimated demand, and where even among people with interest in VR and disposable income, the compromises on battery life and weight were actually too much to bear. Forecasting is just hard and you always need to be especially skeptical when you yourself are doing the forecasting for something you want to succeed. All those arguments for the neo line up in retrospect hindsight is always 2020
I’d be willing to pay more to own vs subscribe, but the gap is currently far too large where buying a Mac Studio for AI is a straight up terrible investment.
My biggest problem with buying a Mac Studio is that even in the case of the M5 Max models, I can’t think of any non-AI macOS applications in my creative life that have anywhere near that level of hardware requirements, and nor can I forecast that I would within its ordinary supported lifespan as a macOS machine. These machines left high-end stills and most video work behind generations ago, for example.
So while I would like a local machine that I can leave running in a way that I wouldn’t want to do with my old secondhand M1 Max laptop, it’s going to have to be something a little more pragmatic. Probably based around the Radeon R9700, since the AMD/Nvidia gap for LLMs is beginning to close, and even a single R9700 runs the main models I am interested in at speeds that are acceptably faster than what I have here.
At the moment, speccing out a local box more powerful than the machine I am using is largely an intellectual exercise, but it does have some value for understanding what a client might be able to use if they absolutely require on-premises inference, and I do have a couple who fall into that zone. So I am trying to keep up to date in principle, but the built machines don't get beyond the shopping cart.
The M5 Ultra has probably smoothed out the major issues of AI on Apple Silicon (prefill doesn't suck anymore) but the way Apple marketed that chip strongly suggests they are now fully aware of the strengths but also the limitations of their architecture and will get a lot more done.
It's properly put the (fairly equivalently priced) DGX Spark in the shade, though.
Not if you're an enterprise that wants or requires on-prem inference.
i come from a ultrabook with 2 gigs of ram and the neo takes 2 vscode instances, one antigravity and firefox while keeping telegram and whatsapp on the background. that device can be used as a workhorse, pagination is very good.
can even run ios/android emulators but then you have to only have that project open, which is a non-issue for me. and the battery literally keeps working all day. good screen, good keyboard, good touchpad, operative system is close to linux. if youre thinking about getting one for programming and you're scared about the 8 gigs, hope i gave someone some light about it.
Basically you can have your own 24/7 AI Employee. At-least thats the appeal and folks were buying mac minis massively.
Just my gut feel. Everybody moved on from that now. But it was a big deal back then.
Now, I use claude code directly on my machine with the remote control flag.. so its like open claw exactly. I can be outside and give commands to it via mobile app
then Control + Command + Q ftw
(keep system from sleeping if plugged in, keep system from going idle even if not plugged in, for 2 hours, start claude code w/ remote control enabled, lock my mac w/ the keyboard shutcut)
It's quite expensive from a value per token perspective. It's quite cheap from a "send one message from my phone and it does something probably productive" perspective. No major security breaches yet.
AI for developing software that runs predictably seems more useful than AI performing actions like a program would.
I think AI is an incredible competence magnifier. If someone knows what they are doing that competence can get magnified massively. If they do not know what they are doing, that gets magnified just as much. The tricky part is getting people with a low level but real degree of competence, which is all of us when we're starting out, to learn how to reliably magnify that.
It was able to create an application that Claude created, using the same instructions file. Considerably slower but I was surprised it could do it.
The main place they are a bit behind is in the number formats they support natively. Iirc M5 doesn't have native FP8 support, so you will take a speed penalty on quants where other architectures get better acceleration.
Macs don’t have very good prefill, though. So it’s important to use a model serving stack that has excellent prompt caching and use a harness that won’t bust the cache.
I’m cross-shopping DGX Sparks and M5 Studios, and having a hard time deciding because they have exactly opposite characteristics for prefill and decode.
The first time around they almost went broke because people no longer were paying Apple tax, and they did not had any other product to save them.
Nowadays they have the iPhone piggy bank.
Even though you saw this coming 4 months before the launch, at the launch time itself you're still caught off guard. Demand is still dramatically higher than your capacity to meet it, and that's still going to be true for at least another few months, and there is nothing you can do about that. What you can do, you did 4 months ago to increase parts orders, but you're still caught flat footed right now.
Under Apple's reality distortion you instead must say that they were surprised by the demand.
Projecting much?
We don't know if Apple misjudged the demand or couldn't source enough parts to match the initial demand. Another explanation could be that they decided it wasn't worth it to increase the production capacity to match the higher demand at launch because that capacity would go unused later down the line.
You could be right, but I don't see being caught off guard as the kind of praise you seem to think it fosters for Apple. If I were in marketing I would be looking for a message that showed Apple in more of a leading, not following, position.
The way you get people to believe a lie is to put just a bit of truth in it. The way you get people to accept marketing as not-marketing is to put just enough not-marketing in it.
I'm training a model using reinforcement learning with self-play. I can and do use vast.ai when scaling but for experiments it's far faster, and cheaper, to run it locally until the bugs are all figured out. Just provisioning a new instance and copying the relevant checkpoints and things can take 25 minutes. It's zero locally.
Apples ProRes codec is only licensed to run in high quality mode on a Mac, and so my Nvidia PC can’t do what I need. Thus, I own the beefiest Mac Studio you can currently buy. I would pay more for more TFlops.
I have done local LLM on there but it wasn’t interesting. Far worse performance and intelligence per dollar than the cloud boys.
There is no cloud offering for my video needs though.
It’s probably files that, over the course of a month, add up to multiple TBs.
Which would suggest that a 1Gbit fibre connection would be adequate. For serious commercial usage, multi Gbit fibre is available in many places around the world.
Although being video files, they could easily be in the TB range. In which case, it would be interesting to know!
I had to start with some heuristic-based bots that played the decks very simply just to get to the point where the was some signal to learn from. I did behavioral cloning on the bots as a foundation, then self-play.
If you can fit it on a GPU, and especially for training, it is so much quicker than a Mac.
Inference with MLX is surprisingly zippy. I’m running a classification task on the entire HN comment dataset and it’s projected to take about two and a half days, which is not bad considering we’re talking about tens of millions of comments.
Yes, I could do it much more quickly by throwing Modal GPUs at it but this is low-priority work. I might as well throw my M4 a bone.
Modal significantly improves this. Highly recommend.
Do you find that CoreML manages to fill up your drive with so many tiny files that a reboot takes hours to clean them up? I keep meaning to get my friends still inside the spaceship to file a radar about that.
What game are you building?
I'd suspect that agentic coding has birthed so many new effective engineers, that the entire dynamics of demand for high-end machines have been upended.
nothing like 600mb of for type check server another 500 for webpack and another 600 for the inevitable electron wrapper you didnt know about. python isnt as bad but still not great.
I'm pretty sure the C-level people don't want us programmers faffing around with their roles. Why are they encroaching into ours? Don't they already have enough to do; setting the direction of the company, making sure it's profitable, selling things to people, etc.?
My team is using coding agents to develop analytics and reporting applications, and an engineering automation application, specific to our needs, that simply would not exist otherwise. The cost and time would be far too great. I can easily imagine there are plenty of cases like this, for all sorts of roles, including managerial ones, where only the person doing the job understands the job well enough to specify a requirement and guide an agent to code exactly what they need.
[0]https://pmarchive.com/guide_to_startups_part4.html: "In a great market—a market with lots of real potential customers—the market pulls product out of the startup... The product doesn’t need to be great; it just has to basically work."
They have been taking NPU's seriously on their phones, tablets and laptops since the M1.
Then they enabled fully-connected RDMA for 4 x 512GB MacStudio's = 2TB RAM. Perfect for a large Mixture-of-Experts model.
It would be very strange if they didn't notice their product line had landed in a new sweet spot.
While the company I am in is embracing AI the disconnect and delay between what is available and possible versus what is approved and permitted is a three month window. The State employees I speak to are just now getting around to writing their usage policies for internal AI usage.
All the while the top of the company is of course praising AI and saying we should do everything with it right away.
It's a joke really. We had to go through all this rigmarole just for an agent that does some preliminaries on a service support ticket (making sure all data was filled in correctly and contacting the requester of not) before sending it to a human for final review. It couldn't action anything, not a single thing. Just assist with the prerequisites.
And yet they call our company 'innovative'. We're certainly innovative at inventing bureaucracy.
This story turned out to be false but I think smart, reasonable people a couple years ago could have believed it with conviction. It doesn’t really seem like “completely asleep” to me.
But "oops, we missed that people are interested in AI work on our machines" seems like a really fucking big myopia. But then again, Tim's off to retire on a bed made of cash this week, so...
My vibes were that Apple wound down the “actual work” side of their operations (including machines like Xserve), because Ives couldn’t handle the unsexiness and unpredictability of business requirements in hardware.
He was self-indulgent and only wanted to work on things that “vibed” with him, rather than what the customers needed. It’s easy to be creative when you get to do what you want to do, it’s hard when you have hard constraints.
Enterprise sales sucks. There's infinite amounts of politicking and glad handing and buyers will get all sorts of sweet brib..."sales dinners" then go with the cheapest option. Margins on hardware sucks and the only money is in support contracts. Apple instead invests in consumer sales/support primarily and all the other channels are side businesses.
Stuff like the Xserve existed mostly for Apple internal purposes and ended up being sold externally to goose the scale enough to make them not a huge loss. At one point a large percentage of the offices on Bubb road were packed with Xserves running portions of the iTunes Music Store and the Apple online store. More offices were packed with Xserves doing media ingest and encoding for iTMS. Just about every building had racks of them as build and file servers.
However, I don't think they expected the level of Enterprise interest they saw.
“Not as fully staffed as some people might hope” or “Developer Relations isn’t as responsive as I’d like” are both at least not obviously false.
e.g.: previous crypto hype-cycle
https://www.pcgamer.com/nvidia-cmp-graphics-card-availabilit...
If they had a choice, they would have sold every GPU to gamers instead because a crypto boom was always followed by a crash which flooded the market with used GPUs and crashed Nvidia's stock price.
Sure they did and still do. They did a good job of managing the bubble but participate in the market nontheless. Read the above article: Nvidia has promised that its new lineup of cryptocurrency mining GPUs, called CMP for short, won't impact the supply of GeForce graphics cards for PC gamers. . . Since these cards are destined for mining rigs they will also lack video outputs. That may mean the resale value is diminished, which has been one reason why miners prefer gaming graphics cards.
Or take it from the horse's mouth [0]:
NVIDIA Cmp Hx
Dedicated GPU for Professional Mining
0. https://www.nvidia.com/en-us/cmp/They also nerfed GPU hash rates through software to prevent miners from buying them.
[1]: https://www.youtube.com/watch?v=fYuH2Kl_b98 [2]: https://scholar.google.com/citations?view_op=view_citation&h...
Yes, the whole Deep Learning thing was luck, but as with most lucky things, they ensured they were positioned to capitalize on it.
AlexNet kicked off a new wave of research around neural networks by demonstrating they could be scaled well and trained on GPUs.
This is clearly a mis-statement, they have a whole annual conference for developers. Maybe they mean specifically AI devs.
Frankly even the entry price is a bit high - I remember buying one for my son a few years ago (M1 mini) and it was a few hundred; now we're up to $900 for the base model.
The base price of a mini has only gone up $200 from $699 in 2020 to $899 today, and for $699 you only got 8GB of RAM instead of 16. Yeah the price has gone up but not nearly as much as you seem to be remembering...
The only reason for this huge speedbump is that chip makers have been dragging their feet for the last 10 years with "just enough" memory.
Stuff is crazy expensive.
I also bought 96GB some while ago but after the initial increases, thinking I'll wait it out. Now 128GB is far more expensive than it was when I first looked. Luck has it I want DDR5 RDIMM as well, which seems the hardest hit when it comes to RAM prices, fun stuff.
I've been using one for about a decade as a media server.
It just sits in the cabinet happily running the macOS TV program with the video files on an external hard drive. Playback on the TV is handled by the AppleTV's built-in Computer app. Works beautifully.
I have more movies and TV shows on that box than I could watch in my lifetime — a combination of ripped DVDs (Netflix, public library, and purchased) and OTA recordings.
When the cable goes out in my neighborhood (frequently), or a big storm screws up satellite reception (seasonally), I just don't care because I'm all localhost. As long as the lights stay on, everything is fine.
No ads. No privacy violation. No fees. No bandwidth congestion. No buffering. No subscription rate increases. All I pay for is electricity.
What planet do you live on?
1. Running an agent like OpenClaude. The $599 Mac Mini was an insanely good deal for this. I happened to buy a M5 Pro Mac Mini for $999 last year for other reasons. The equivalent is now almost $2000; and
2. Hardware for running inference on local models. This to me is the far more interesting market because Apple has a real opportunity to disrupt NVidia's stranglehold on the market.
With current architecture, the largest model you can reasonbly run is the amount of memory on the GPU and is a function of the quantization (eg int4, int8, fp8, fp16, etc) available and the number of parameters. NVidia aggressively segments the market. The most VRAM on a "consumer" card is 32GB on the 5090, which allows you to run ~31B parameter models.
In comparison, the RTX 6000 Pro has only slightly more CUDA units than a 5090 but has 80GB of VRAM. A few months ago they were $10-11k. Now they're ~$16k.
Macs use a shared memory architecture. Apple has previously sold Mac Studios with up to 512GB of RAM. Almost all of that memory can be used to hold much larger models without taking a penalty for interconnections between different GPUs or machines. Plus Apple interconnects between computers are actually relatively good by chaining TB5. It's still slow but it's about the best non-enterprise option available.
But the previous Mac Studios just didn't have the raw FLOPS and memory bandwidth. The M5 Ultras are up to 1.2TB/s of memory bandwidth. M3 Ultra had ~900GB/s. RTX 5090s and RTX 6000 Pros are 1.8TB/s. The current best HBM3 NVidia DC GPUs are at 3.2TB/s IIRC. But the M5 Ultra has a claimed ~4.5x the FLOPS of the M3 Ultra.
We don't have our hands on these yet but it probably means they are going to be much closer to a 5090. I expect ~50% of a 5090's inference speed. That may sound bad but it's actually really good because a 256/512GB Mac Studio can probably locally run the best Flash models. With NVidia hardware you'll need to spend many tens of thousands for that.
We'll see what the inference speed is but I expect it to be usable. DeepSeek v4 Flash, for example, will be entirely runnable. We're not at DeepSeek v4 Pro local yet.
I still have zero clue how "Buy a $599 Mac Mini to have a sandboxed LLM API caller" became the default. If you're not doing local inference and don't need to inject into iMessage or iCloud, all you need to run openclaw-style harnesses that call external APIs is a Raspberry Pi 4B, an N100, an HTPC, or that 10 year old laptop sitting in your desk.
Incidentally ios27 has a user control slider for glass frostiness.
I’ll take significantly removed.
What’s there to refine - Liquid Glass is the single most pointless, wasteful, accessibility destroying ‘feature’ I have ever seen
https://www.macrumors.com/2026/08/30/apple-unexpected-mac-mi...
https://www.macrumors.com/2026/05/01/apple-was-caught-off-gu...
https://www.macrumors.com/2026/01/29/apple-on-airpods-pro-3-...
Recent work in this space has got me looking at using local models for daily use. I’m waiting for people to start dumping some of the previous generation minis on Facebook marketplace or eBay so I can pick one up. Of course there’s also other options. I’m still hammering out my requirements and what I want to do besides putz around.
I like to buy high end and then keep it for a long time. I'd love local AI But even with a fairly large spend, these machines won't run that much local AI well and that's now, I'm not sure if in the future I'd want even more RAM, seems to be it is better to budget for an AI sub + a less insane machine.
I hope Apple can take all this cash and do some stability releases like they used to do, bugs around things like Family Sharing, the painful "update" to Settings App, etc could all use a lot of love.
It's a niche market but it's a market that overlaps heavily with professionals in the AI space and lead developers, so it's a market that gets them customers in those roles.
If I were running Apple I'd call the RAM price bubble for what it is and temporarily eat some margin to offer machines with more RAM than competitors, especially these models that are great for edge AI, and capture market share.
Nvidia has CUDA, AMD has CDNA, and Apple has... compute shaders, I guess?
How much it matters in inference? Most GPUs have enough computing for that and the bottleneck is the RAM speed and size. And M5 Ultra is becoming to challenge this.
It's kinda why memory bandwidth is an enormous red herring, even for datacenter applications. Nvidia's huge advantage is a compute-optimized GPU architecture and their Infiniband networking, their memory controllers aren't really the star of the show.
They're letting everyone else pay for model training, with the obvious end of the road being open weights.
They're letting everyone else overspend on first-generation AI data centers before the chip industry has truly optimized for today's AI designs (which will reduce data center footprint).
They're letting everyone else play with and pioneer UI ideas.
Meanwhile all they've done is put a few toes in the water: "Apple Intelligence" which is barely anything, and adding AI acceleration to their GPUs and doing a bit of up-market marketing to AI.
So yeah they'll probably follow with a second generation Apple Intelligence that incorporates everything everyone else pioneered that worked, and an M7 or M8 line of chips with a GPU augmented with whatever approaches the other chip companies found worked best for running models... and built around the state of the art model architecture the industry finally converges on. By then RAM will probably be cheap again, so you'll get a 14" MacBook Pro with an M8 with a tensor-GPU and 512GiB of RAM.
How so? In tokens per second when running major open-weights models, or something else?
I use a DGX Spark with NixOS on it as my daily driver and it's fantastic.
It's true, most people don't run models, but being the default platform for running open weights seems like it has plenty of advantages right now. Just like sales benefited from developers defaulting to MacOS for most open source languages like Ruby, Go, Rust, and TypeScript.
32GB is not enough RAM. I don't even own a device with less than 36GB at this point, and that device I only have because my employer is being cheap. 64GB is a reasonable starting point for running local LLMs + normal tasks. 128GB let's you really run most smaller models like Qwen 27B and 35BA3B with good context. Even Qwen3.8-Flash-Next runs in 128GB with a 4-bit quant.
32GB would be limited to running models like Gemma4 12B and smaller dense Qwen versions like 9B unless you were using very small quants which damages quality of response.
Is this the best? No. That's why I said the sweet spot. Getting from 16GB macs to 32GB is perhaps possible. Jumping to 64GB or 128GB as the default is simply unreasonable right now.
What sorts of things are you doing with the local LLM? Anything interactive? Should I take another look?
EDIT to add that you need to reserve 8GB for the system if you don't want to cause problems on macOS, which means 32GB RAM = 24GB max for model + context. It takes 18-19GB to load a 4-bit quant of Qwen3.8-27B, so I'd be really surprised if you can actually get a 200k context window. You need to fit within a 24GB WSS (which is generally a more constrained RSS) to get stable performance on 32GB RAM.
And I don't remember to have been able to have pushed to 200k context Qwen 3.6. 3.8 is running on my RTX 5090.
64GB+ or dedicated 48GB (2x24 on GPUs) is IMHO absolute minimum.
- 16GB for the weights at Q4
- 9GB for the full 256K context at Q8
- 7GB spare for overhead and system.
The problem is that these Macs have 32GB of slow unified memory.
Edit: I'm thinking of a headless Mac mini, if you meant running it on the same machine you're using of course you'll need more memory, but LLMs are best served from a headless server so that's what I'd recommend.
32gb of unified memory is enough enough for system to be used for anything other than LLM generation.
What? LLMs are best served from a massive PD disaggregated cluster of B300s connected via NVLink.
If you're running LLMs on a Mac Mini, it's because you want to run local, not because it's the best setup.
So a headless server.
Macs were mentioned because that's what the post is about. It could be a PC (I use a 2x3090 PC). The point is that it's a better experience to have a box dedicated to the LLM than running it in your system. Obviously in your home, so local.
If someone wanted to use an 8x B300 as their daily driver...go ahead.
It would still be the best way to serve a given model.
Single user conversation spawns multiple parallel backend conversations, you need extra room for it as well, not just single context.
This plus usual apps like Mail, Spotify, iTerm2, SourceTree etc. also fill in memory.
Also 4 bit quantization is already quite aggressive compromise (measurable but sometimes acceptable loss, compared to ie. 8 bits which are often practically lossless) – for weights it's ok'ish, sometimes (especially if model was trained as 4 bit quants aware), but activations need to stay at higher bits taking more memory, otherwise quality degrades a lot.
For a dedicated headless setup, I’d probably use something like NVIDIA DGX Spark rather than a Mac (to be more precise NVIDIA GB10 Grace Blackwell from other suppliers than directly NVidia, they are much cheaper and have same insides). Linux is much better for running headless server, you also get standard NVIDIA/CUDA ecosystem instead of being tied to Metal/macOS.
For people who are interested in buying IMHO I'd wait a bit – next generation of Spark and/or Macs that are going to come out next year will be much better / will cross the line of being actually useful, not just a toy with goldfish LLM.
https://www.canirun.ai (five months ago: https://news.ycombinator.com/item?id=47363754 377 comments)
One very efficient option today is to have the cheapest Jetson (Orin Nano) run the classical robotics stack, then have a base mac mini run nothing but the VLM. The Mac mini is considerably cheaper and faster at these workloads than the mid-range Jetsons.
I think this wonky situation is because Apple us under immense consumer pressure to absorb the ridiculous memory prices, while the Jetson is aimed at "business" and much more likely to fluctuate with the market. Last year I bought a Jetson Orin Nano 8GB for $375CAD, today that official nVidia Amazon page is out of stock and other sellers have it listed for $900-$1100CAD. Absolutely bonkers pricing.
Everyone else keeps buying PCs as always, even more so that Apple doesn't have any proper desktops any longer.
Also if there is demand for something, are CUDA cards, being used from Windows and Linux distributions.
I always thought Apple was in a good position because unlike their peers they didn’t burn billions of dollars trying to build an AI model that has no financial moat around it.
I still think they’re in a good position in comparison to their tech peers and we will know even more when some of the new computers get into the hands of some of the tech reviewers.
I believe the new computer’s will be pretty good hardware wise what I’m interested in, is the Apple software support for connecting several Mac computers together, and some of the other (new?) software that Apple may have written in house to support those who want to run AI software locally that is just as important as the new hardware.
It's old Steve Jobs logic. Works backwards from the customer experience to the technology (they're the only big player I see doing this).
A year ago you could get an M4 Mac mini for $399 on sale and now the same one used goes for over $700. The general AI RAM/SSD spike is part of that but there was also a huge demand spike for small, powerful, desktop machines that could be easily configured with these workflow tools.
They did have one... Xserve. Or the macOS Server package you could buy on the App Store. But there hasn't been an actual device that people could buy for many a year, and most of what enterprises need (aka fleet management, software deployment etc) has been left to third parties like JAMF.
I would wait till the ram crisis is over to fetch a future 64gb ram gpu to run Q8 models. Cloud inference until than.
What makes you think that? There's a lot of data centers that sell you access to colocated Mac Mini's, they have added FileVault unlock via SSH in the boot process which also makes things easier. There's not that many reasons to run a Mac in the cloud unless you have some very specific Mac related workload.
https://wccftech.com/apples-private-cloud-compute-server-m5-...
There’s articles about them, Apple uses them internally for AI services. https://forums.macrumors.com/threads/photos-of-apples-own-ne...
Here's a leaked / rumor image of Apple servers themselves.
https://www.macrumors.com/2026/08/26/leaked-images-of-apple-...
https://www.macrumors.com/2026/08/26/leaked-images-of-apple-...
https://www.reuters.com/business/apple-begins-shipping-ai-se...
> [still claims there is no reason]
https://forums.macrumors.com/threads/photos-of-apples-own-ne...
Only if by "secret" you mean "announced in multiple press releases and a public event with federal, state, and local officials at its new sever factory in Houston."
https://www.apple.com/newsroom/2026/08/apple-opens-advanced-...
Plus they did sell rackmount servers for some time.
And?
[1] https://tech-insider.org/ca/steam-deck-price-increase-2026/
I can't even name one person who fits this mold, let alone a non-insignificant proportion of people. Who are you thinking of?
One thing to watch for when Apple introduces the new phones coming up shortly is whether or not Apple has replaced Qualcomm in their flagship smart phones because that is coming up soon Qualcomm has given warning to their investors.
Obviously; no one else can justify the expense.
I have no idea why. They could be so successful if they leaned into local AI.
But apple have to own everything they do so they’re in a bind
I also don't think they do "technology". For example, containers have been around for a long time, and apple didn't show up. (I know they have some support now). Imagine an apple-native docker/podman doing something like FROM macos:10.12
I was actually surprised when they did their own chips. I figure it was about control.
Classic monopoly move by who?
Apple created MLX as an open source framework to allow users to run any open model locally.
MLX was just Apple's bridge to what already ran in other hardware.
You can take advantage of larger amounts of memory, higher memory bandwidth, and clustering multiple systems.
How is using an open source framework to run open models a monopoly move?
Local inference solves so many of the privacy and inconsistency problems with these frontier subscriptions.
Memory probably is the biggest holdup/obstacle at this moment.
It's all marketing and you're being manipulated by it.
Edit: Some of y'all are naive.
Is that what others read ?
This unified architecture makes Apple hardware very good for AI work, where latency between RAM and CPU is very important
Even the SSDs (which are normally more strippable) use a proprietary hardware form factor.
I think the article mentioned the real draw:
> Apple noticeably promoted the ability to link multiple Mac Studios together into a single, more capable system for running large frontier AI models, a feature aimed at business and developer customers rather than everyday consumers.
(Couldn't read the source cited, might have more info: The Information)