Apple Silicon Exec Explains Mac Mini AI Demand and On-Device Future
macrumors.com
macrumors.com
Just writing this down so I can be praised/mocked in 5 years.
The one thing that is marginally exciting: the Apple SoC or M series chips.
It's unfortunate they are locked behind crappy macOS and other proprietary apple crap.
Unsurprising. Apple seriously thought the iPad would replace computers and usher in a "post-PC" word during their "What is a computer?" ad campaign era. Now they are sticking phone chips in laptop chassis.
...for some users. See their "Mac is a truck" analogy.
And it has. My parents haven't owned a Windows or Mac machine in six years, since they got rid of the one I gifted to them a decade ago. Its all iPad and iPhone.
We are both late and early.
Buy an Nvidia Spark, then whatever cheap Mac you want to use as a thin client. There's no reason to force Apple Silicon's round peg into a square hole like AI inference.
The other half of that equation is latency, predicated on prefill performance which needs a powerful GPU and ideally ALU-level optimization to build larger KV caches quickly. Even the M5 gets smoked in this department, the M5 Max has a 50% longer TTFT on Qwen's 27b dense model at only 16k of context, which is a pretty typical starting context to use for agentic editing in normal apps like OpenCode/Claude Code: https://raw.githubusercontent.com/Osmantic/MMBT-Messy-Model-...
For agentic, 50-256k token on-device coding sessions, the Spark will be faster and consume less power running larger models. Without an external GPU (which Apple doesn't support), Apple Silicon will always be bottlenecked during prefill. Apple's failure to address this with their GPU architecture is a big reason why Apple Silicon viewed as a waste of time and money for professional datacenter deployment.
I keep hearing people make this claim that TTFT is a problem, and… it just isn’t, if you’re running oMLX.
Folks in my camp keep saying this, and folks in your camp keep beating a drum we tell you isn’t resonating. Not sure why I keep bothering to argue; you can’t buy a high RAM Mac Studio like mine anymore.
They're not usable for deployment. They're perfectly fine for "enthusiast" low-end usage with 10-30B models, but the same goes for almost every dGPU made in the last 10 years. Your Mac Studio cannot run frontier LLMs at an interactive speed, even Apple has given up on using it as an inference backbone.
You think this is a mistake...
Of course. Do you think this was on purpose? All part of Apple's brilliant master plan?
I'm fatigued by it all at this point. It's streamlining the interesting and fun parts out of my job (by practical necessity of use there), and if I used it half as much outside of work I'm sure it'd do the same there too.
This is the prevailing opinion of people even outside of tech.
That said, I think it's a good thing that this sentiment is coming to the forefront.
He was complaining that he would ask how to perform a certain repair on a car, and the LLMs he tried (ChatGPT & Grok) would give him a long involved process and he'd ask why not do it this simpler way and it would say, oh you're right! He just found it gave bad advice and realized (rightly) that in areas he has less expertise in he has no way to judge how good the outputs are.
This is from a guy who loves tech, historically worshipped Elon, loves his Tesla, and (rightfully again) didn't buy into SpaceX because he thought it was overvalued.
In the past when I visited for holidays he was liable to have a positive outlook on LLMs and their utility. Seems telling that he's starting to see the cracks.
This is easily the biggest problem with the current models. The models are just way too eager to please / say yes to the point that the models are happy to lie/make shit up if it means it can say yes.
Interesting. For me it's streamlining the tedious and attentionally taxing parts of my work tasks. I love solving problems, I don't particularly love shaving yaks.
Apple is doing something very different. Their AI experience for end users definitely has been a little behind.
Apple Silicon, however, has been quite unique for the last 4-6 years and it's increasing overlap with LLMS.
The model/chip optimizations are definitely improvements, the thing that is really standing out the past 2 years is how much the open source model community has been making possible, especially when you know a group of use cases.
1. Most useful LLM work is done in parallel. A Mac Mini can run one LLM inference thread at a time. The cloud can spool up dozens and spread that inference across efficiently batched operations over a fleet of hardware.
2. Faster inference hardware such as the chips from Cerebras and Groq cannot be run locally. But the advantages of running >5x the token throughput per thread can’t be overstated. Add in the multi-threading advantage and it’s a knock-out punch for local LLMs.
Local inference has a role: if you’re working with extremely private matters or you want an uncapped model that will talk dirty or generate NSFW photos, local is the only option. I think Apple and others will continue to also run a lot of useful workloads locally such as text editing suggestions, speech to text, text to speech, and image manipulation. As local hardware improves, these capabilities will get better too.
But, for most LLM work, the cloud will continue to dominate for a long time to come, if not forever.
I don't want to run any workflows on someone else's computers.
That’s not accurate. With MLX, at least, parallel inference is both possible and useful. Model serving tools like LM Studio and oMLX support parallel generation with continuous batching, and the total throughput increases with it.
Can I run a few inferences in parallel on my Mac Mini? Yes. But put 1,000 Mac Minis in a datacenter serving 1,000 copies of myself? That's going to be more efficient.
There is so much happening in that scene, where tokens/sec double or 10x
So I could see the same hardware doing 20 tokens/sec on a large model suddenly doing 200 tokens/sec in the future, a better device in the future doing 500 tokens/sec, while having vision models baked in, audio models etc
Users wont consciously switch to local, they will just have it and use it
Well just because they don't sell them. Doesn't mean that will be the case forever.
you’re far far far more likely to see a camry or equivalent in an americans driveway than you are a super car.
you’re also more likely to see an enthusiast with a corvette or equivalent muscle car spending way more than it’s worth to tinker on that car in their garage than you are a super car.
I guess what I'm doing is not considered that useful then? I usually only have zero, one, or occasionally two things actively doing inference at a time, be it claude code sessions or one of the chatgpt/claude web interfaces, and i bet that's true for like 95% of people using llms. And anyway i bet even the hardcore people using a bunch of parallel agents would appreciate having access to local, private inference for some things.
You're obviously right though that cloud inference isn't going away anytime soon
And yes, I agree. I find the experience better less for speed and more for context management. But it's far from necessary.
There is probably a middle ground though. Something like a main thread agent (say a local model for Hermes) orchestrating cloud based sub agents.
Apple's shovel (ahem, Mac mini) is the highest quality.with Companies burning money left, right and center, Apple can dispense with advertising altogether
It would need a path to a $2,500 machine, I think. But this is a niche I don’t think another consumer-facing brand could do like Apple.
Like, take out the price sheets for the Apple Car. Then sell me an AI tower at those price points.
Those Macbook Neo users would be very reliant on Apple intelligence, enough maybe to pay for a service with it. I think Apple's much happier going this path.
If it's an "or," absolutely. But if it's an or, they should be prioritising Macbooks over the Mac Mini Doug Brooks is discussing.
When we breach the "and" of memory supply sufficient to allow for more Mac minis (and Mac Studios), I think it would make sense to consider relaunching Xserve (with new branding, of course) as a consumer/small business product.
The writing has been on the wall since 2019. Apple doesn't like the old way of computing, their goal is to expand the ecosystem by prioritizing install-base and then pushing first-party service offerings like they did with the iPhone. And like they did with the iPhone, Apple is great at ignoring power users to focus on features that make them more money.
You may be waiting a few decades for this type of product, memory supply be damned.
Now that the Mac Pro is depreciated, Apple's plan to pivot to service offerings seems set in stone. That's the "want it all" attitude they've adopted with the App Store.
At the $150 mark (which is probably accurate factoring in lifetime service spend), that's a $10,000 minimum return on the 64x Macbook Neos. Apple can charge that type of premium on consumer hardware, but they're in no position to command $10,000 margins on professional hardware. They're not Nvidia, Apple has always been LARPing as an HPC vendor.
Apple won't subsidize these low-margin enthusiast products with the profits made from services and higher-margin hardware. Tim Cook would much rather ship the 64 Macs, and get ~15-20 school-age kids hooked on Apple One or the App Store for the rest of their life. There's understandably not much patience for catering to people that want to opt-out of the Apple Intelligence service ecosystem, effectively leeching off of more successful products. The volume and opportunity cost kills the concept in the cradle.
The hard part is the GPU architecture. Apple Silicon was designed with a laser focus on raster efficiency (similar to AMD's GPUs) which makes a lot of sense for highly mobile hardware, but is a crippling mistake for high-performance compute. Apple's largest Ultra chips are hamstrung with SOC-tier GPU performance, their highest-end desktops are outperformed by Nvidia's laptop offerings. Apple has to find a way to scale upwards without imposing too much architectural strain on their cheaper hardware like the iPhone and Macbook. Nvidia has already solved this issue; full CUDA compute stacks are usable on extremely cheap GPUs like the Nintendo Switch's Tegra SOC, or the Mac Mini-sized Jetson boards.
In terms of "who needs to redesign more to address the market", Apple has a lot of technical debt to unearth before they catch up to Nvidia. And if they do catch up, Nvidia will still support Linux and other differentiating features that Apple refuses to implement. It definitely feels like Nvidia is closer to a winner with the Spark than Apple is with the Mini or Studio.
It assumes that RAM remains supply constrained and that none of the existing RAM contracts are cut short.
But Meta and xAI putting A TON of AI compute onto the market. OpenAI and Anthropic are raising the costs of inference (by reducing how much inference users get via subscriptions). And we haven’t seen Oracle / CoreWeave struggle to pay their debts yet, but they will be selling assets once they get close to that point.
Edit: Okay, this doesn’t mean that that’s actually possible in the short-term, so I think you’re right. But that means as the silver lining, in the medium term horizon there’ll be enough supply again? :’)
Memory is a cyclical market that has historically rewarded conservatism [1].
Counterpoint: there is enough demand from enough capital-rich customers that they may be willing to shoulder the capital risk.
[1] https://www.ldeepai.com/tech-hub/dram-industry-consolidation... Sorry for the slop link, it has a good chart from a solid source
So essentially, due to technological progress and other factors inducing price collapses (or at least cycles), you can’t start stockpiling insane amounts of finished-product semiconductor, which means you can’t scale production at current technology levels to infinity either?
I also believe there would be one or two tech companies that will get into memory by taking it in-house to make sure that they won’t have this problem again in the future.
This is a good hypothesis. Curious if anyone has data on the failure rates of new entrants in semiconductors based on how frothy it was on founding.
On one hand, more demand makes selling easier. On the other hand, a shortage makes your input costs (consumable and capital) pricier.
EDIT: It seems like the 2 to 3 year lead time and a crowding effect from new entrants historically made booting up a fab into a boom a bad bet [1]. (The article argues, convincingly, that this time may be different.)
[1] https://www.uncoveralpha.com/p/every-memory-cycle-ends-the-s...
https://www.techspot.com/news/112502-memory-prices-tipped-fa...
If a place can do it, another place, with a huge track record on manufacturing and lately expanding all kinds of tech, can.
Whether or not you feel like those are good overall (I do), they do actually also slow things down.
Yes, like how it helped western industry early on. Or, well into the 70s for the most part.
Unfortunately its not so cheap anymore as everyone ramped prices up of course.
Last year I could still get 32GB of DDR4 for under $60 from chinese brands.
I just upgraded my 2008 Thinkpad R61i to 8GB of DDR2 a few months ago while I was also upgrading to a core2duo.
DDR2 and DDR3 are still in active use by SBC manufacturers.
Uhhh… ok… good for you I guess…
Chinese fabs might not be so tied with red tape and regulation upon regulation (which is a funny reversal, in terms of "communism vs capitalism" bureucracy/inefficiency cold war thinking)
All of their fabrication ability is based on old processes.
2.) Authoritarianism can move faster than anything. They can just say "wipe out that village, build the coal plant there, data center here, fab here.
3.) If it's red tape and regulation holding the US back, then that's clearly not "capitalism."
Except in the actual historical sense. They appear to enjoy all sorts of freedoms, increased prosperity, even have elections at different levels but under a single party system. Which is not necessarily that different than a effectively two party system.
>2.) Authoritarianism can move faster than anything. They can just say "wipe out that village, build the coal plant there, data center here, fab here.
Now that China is more effective, "it's easy because they're authoritarian". Before the argument was "authoritarianism can never be as effective as free-market democracy".
>3.) If it's red tape and regulation holding the US back, then that's clearly not "capitalism."
It's real world capitalism, not some fantasy some guy imagined removing all warts.
And that's somehow fascist, as if the problem with fascist Italy or nazi Germany was that they didn't hand out citizenships?
what if demand keeps rising faster than production capacity is deployed?
We are in a bubble which will be burst the moment the world starts retaliating against the US' 20+ year history of supporting genocide and committing war crimes unabated.
Buy the AI toys while you still can.
> Multiple motherboard and PC component makers move forward with Chinese-made memory validation
https://www.pcgamer.com/hardware/memory/multiple-motherboard...
The only thing standing in the way of a major Chinese DDR5 ramp up is money.
There are lots of signals that the sector has been overinvested and that corporate customers are pulling back on spending as the cost of the APIs is revealed.
Once the hyperscalers start struggling to bay their debts (it will happen, just a question of time), there will be a supply glut.
So the only question is: do we share the same definition of “short term”.
The issue with the M3 chip is the compute performance, as it doesn't fit well the transformer architecture. This changed with the M5 (apple baked their own matmul into the chip), which would significantly speed up PP (and video/image generation btw), making the M5 Ultra significantly faster than the M3 Ultra and in practice much more usable. You can try to load Kimi or GLM on M3 Ultra, but it's not usable. Now the M5 Ultra is not out yet, but undoubtedly it will be a superior offering, and shilling 15k on 512GB version is actually reasonable (if it's every priced remotely around that tag).
My money is on small companies or affluent programmers experimenting with some new hobby / business model.
For me, the privacy pitch wins. I have a friend visiting, however, who spends like $2,400 with Anthropic every year. That's a solid ROI even if the thing becomes obsolete after a couple years. (I'm still on my 2020 MacBook Pro. I love it and will be sad when I have to replace it.)
How can that be a solid return on investment? There's no model you can run locally to have frontier model level performance. Also who spends 2.4k yearly for personal AI usage, like what's the usecase? If your friend is spending that money for his business then it's not personal computing.
I'm betting he doesn't need a frontier model. Sonnet, today, is likely good for 80% of his tasks, which largely involve repretitive, tedious work.
> Also who spends 2.4k yearly for personal AI usage, like what's the usecase? If your friend is spending that money for his business then it's not personal computing
Combination of business and personal.
I do the $100/mo for myself, then about ~$200/mo for startup.
Apple makes product lines with assembly lines, its not a hand fab or custom build type of place.
I think "buy this $10,000 box and to easily grant every Macbook Neo on your team safe, private, free AI" could be a real winner.
I could absolutely justify having that machine for the actual work I do, and I'm not even doing any of the really hard AI things. If I were, I might prefer a Thelio Mega workstation from System76, which is $90,383.00 fully loaded.
Alas I can not afford a 10K personal computer right now, which is not the same thing.
If Apple released a machine that would let me run, say, DeepSeek V4 Pro or GLM 5.2 locally at 100 tok/s for $10k, I might hurt myself running to get my credit card.
But then, I'm also posting this sitting in my truck while my family attends an event, with a Vision Pro on my face so I can monitor 7 Claude Code sessions without constantly switching screens on my laptop.
It’s not necessarily out of reach or unusable over the course of time, the thing I keep hearing over and over is how usable many people find Mac’s very useable over the course of time, obviously software support plays a big part of that.
On top of that, AI models are getting more useful despite their smaller size. The future is a personal computer future, not a mainframe computer one. Yes, larger computers are useful, but most people will not be using that larger computers, which will be relegated to universities and larger companies.
I am unsure that apple themselves understand why their hardware (top end & bottom end) has been so successful, without this understanding leaning into these use cases isn't really going to be possible.
You have a bold career as a technical journalist ahead of you!
With their apple finger right there on the pulse, they are going hard on the VR/AR glasses (following the lead of the visionary CEO of facebook), cars and folding phones. By the end of the year (tm) we 100% will have all the features that were showcased and demonstrated 2 releases ago.
I trust they know more about their business model than some rando on the internet, sorry.
How you hold it is, of course, up to you.
It's just that Apple isn't really focused on software development professionals, and it's still fashionable to throw shade on them, so we hear a lot of kvetching about it, in communities like this.
I've been using Macs for all kinds of stuff, since 1986, so I can definitely state they get work done.
But I still strongly believe that Apple hates pro users because they don't make as much money and because they get in the way of serving laymen. The Aperture fiasco, the Final Cut Saga, the Xcode war of attrition and the never ending chain of failures with MacPro - all suggest that I'm right.
(Developer of a major plugin for DaVinci here)
As an Apple developer, since then, I have been incandescent with rage at Apple, many times.
Guess I’m a walking demonstration of Stockholm Syndrome (at least, that’s what I’ve been told).
And I definitely feel the pain with Aperture, not making it to Apple Silicon.
There are many larger software companies that can support all three operating systems, they just make up excuses/reasons as to why they can’t perform over the years.
Apple simply cannot comprehend the ask.
Apple knows the market demand for this type of device.
You may have paid $50,000 for it, but you’re only one customer. At Apple scale they need to focus their finite resources on the products that serve the largest market demand.
$50,000 rack mount servers are not a large demand.
From a historic standpoint though Apple came back from near-death because they differentiated by focusing on the consumer first.
While other companies were recycling the same beige boxes meant to be tucked under desks for home use, Apple came out with products in candy colors.
While Microsoft was rolling out new business process tool SKUs, Apple came out with GarageBand and bundled it in for free.
Apple is not the company prioritizing going after a Fortune 500 company to replace their fleet with Macs. So their focus isn't going to be to design products and features to try to get that deal closed.
No peripherals except Ethernet, integrated compute (cpu+gpu+mem) and secondary storage (+mobo, psu). No accoutrements, just the minimum amount of hardware to run a model as a utility.
Even the appliance faceplate would be a display showing stats like an old HiFi stereo.
Edit: something like a series of modules consisting of a RISC-V CPU + Vortex GPGPU + memory
> Unified memory in Linux creates a single address space accessible to both the CPU and GPU, eliminating the need to manually copy data between system RAM and video memory. It is enabled via NVIDIA's CUDA, AMD's ROCm/HIP, or generic kernel-level Heterogeneous Memory Management (HMM).
So it does exist and is available for platforms that matter.
Intel and AMD had been doing this for years already, and had linux support for it from day 1.
Otherwise, AMD is quite close to what Apple has, and Strix Halo is honestly incredible.
Not sure what RDMA brings to the table.
All have unified memory. Linux runs just fine on all of those.
A bit too expensive for a home appliance though, isn't it?
95% of the price is going to be in GPU+CPU+RAM
Unfortunately their chatbot, while amazingly fast, doesn't know anything about the company running it.
Anyway I wouldn't mind an ASIC running a diffusion language model locally. Even if eventually it would become dated. Beats outsourcing all that to a company that's running on VC money which in the future might either perish or worse - dominate the market and charge whatever they wish.
That's a new one.
Definitely on the edge of what would make sense at home, but its interesting.
Unless you go for the very expensive options, most of the Mac Minis really aren't suitable for running local LLMs, they're painfully slow with prefill/processing input, and the models you are able to run don't handle long context very well, which these sort of long-running agents perform very differently with when you can.
I'll agree with your latter point, hard to beat the value of using something like OpenRouter or similar remote inference.
Even with local models, you can run the agent software and the inference workload on different hosts, which is what I'm doing at home. Beefy server responsible for inference, tiny VM on other server is running the actual agent software + RPC + bridges and what not.
I use Claude Pro ($20/m) as a glorified search engine (no ads/SEO) plus simple hobbyist dev things (shell scripts, managing my Mac, apps etc.
I also use it for tasks like - “search the web for top ten selling EVs, put them in a table” and then iterate - pivot tables, charts, additional research”. It could be cars, it could be broccoli. Code Work has facilities to streamline this type of work, but I usually drop into the CLI.
How much if any functionality would I need to recreate if I switch to OpenRouter and would be match my costs with the API approach. I don’t want any cost overruns. With Codex or Claude, if I run of tokens, no big deal, I can wait.
Thanks!
OpenClaw supports all the mainstream (and free) chat apps like Discord, WhatsApp, Signal, Telegram... None of them requiring a MacOS machine.
Is it a lack of knowledge from the users or do they really value iMessage integration that much?
The relevant questions here are: will the person using this machine also conceivably be wearing a pair of $549 AirPod Max? Or a $399 base Apple Watch? Does that person expect to pay more or less for their largest-screen computing device than their headphones?
Framing that way points toward a $350 price point being a laptop for young children (younger than Apple Watch age, so lower elementary). That's a whole different software experience beyond just the hardware.
Anyone who wanted the OpenClaw use case that is comfortable with Linux probably already has several Linux machines (including a few Raspberry Pis) on-hand.
My understanding is that the barrier to entry to using iMessage makes iMessage a LOT more secure from spam. If you want to do mass iMessages you have to register as a business with Apple, go through all sorts of checks and attestations, etc.
At any rate, iMessages are a lot more trustworthy than SMS. So being able to spam people via iMessage is very desirable. I recall a few months ago a guy posting his little spam-iMessage-as-a-Service product here on HN. You could build your little iMessage spam army using a bunch of Mac Minis...
It seems like it's driven either by 1) people hearing Macs are good for AI, buying one, and using Claude for inference, not realizing that you interact with the anthropic API from an internet connected hair dryer. Or 2) people want their agents to have blue bubbles.
I find it hard to believe that enough normal people are doing on device inference is driving Mac Mini's out of stock. And even if they were the Mac mini is not actually a very good platform for it.
Neo-Siri in iOS 27 removes the need for a lot of this, but before then, if you want to ask a robot about information that is stored in Apple notes, or to send an iMessage, a Mac mini is your only practical option.
It has nothing to do with Macs being especially good at AI. It has everything to do with being one of the last 'cheap' devices being sold with that much unified RAM.
The second is that the puck is heading towards local models. The people running their own 'Claws are usually experimenting running their own services either to save money or to explore the future where 95% of requests are handled on device.
You can sort of justify it by assuming it will last a long time and they’ll use it for other things, too.
Apps like LMStudio, Ollama, Draw Things, etc do a great job of simplifying it but it's still a pain.
MLX is fine, but the cross platform alternatives (ONNX) are terrible, GGUF often lacks and doesn't have the "easy convert" that some of the commenters below say MLX has.
I don't think I'm taking this out of context when I say this is unintentionally correct. Apple still doesn't know what to do about AI.
Luckily, it doesn't matter because it's a solution in search of a problem. Most consumers aren't using AI apart from google search.
Everyone else is using it as a content scraper and praying nobody will step in to end the piracy/fraud.
Others running the beta now on newer iPhones and enjoying it more so?
Although given how effed up the voice for chatgpt is now with the latest updates I might talk with siri more.
Because I use carplay in tandem with my phone where the map is on the carplay screen and turn by turn directions are on my phone, it's always unlocked so I haven't run into whatever lock screen issue you brought up.
On my Pro 16 it has its ups and downs - I still can't get it to "play my running playlist on shuffle" whilst running (this is the only thing I used Siri for before the beta and it would improve my life immeasurably if it worked). But it responds to things like "how long will it take to drive to the AirBnb booking in my inbox", and "when is X playing a concert in Y - add a calendar entry with details" perfectly.
This is a beta and I have hopes, but I can imagine it will run better on a 17 and later
This is... a view.
Maybe I live in a strange sphere of strange ("normie"-ish) people, but the people around me are for sure using AI. Mostly chatgpt to be fair. They use it to compare products that they intend to buy, identify plants in nature, create travel plans, find interesting places to visit nearby, give movie suggestions based on what they have previously enjoyed and so on and so forth. AI is becoming a very integrated part of their reality. To "google" something and digging through the search results manually is very rapidly being replaced by asking chatgpt, for better or worse.
SOTA AI for "serious" work is in a different position, used by fewer people but with big pockets and sometimes a pathological dependence on it.
I can see this kind of low level usage as being perfect for local LLMs... So I can't see a market there for openai etc forever.
Are they using the free chatgpt or a paid one?
>"I can't imagine where we're going to be a year from now, three months from now, or even a month from now,"
I'd say he's making an accurate appraisal of his abilities
It’s not a huge niche but it’s an influential one. They’d get the engineers and CXOs of AI ventures and a lot of academics and hobbyists.
For the platform it would keep them cemented as the high end vendor. In the long term it would position them to take advantage of any software or training breakthroughs that deliver frontier model performance at that scale.
This is mostly an US phenomenon, no Mac mini nor Mac Studio around here.
Only Thinkpads and Macbooks laptops talking to hyperscalers.
People are buying apple unified as electricity costs in many countries are very high, so cheaper to run than Nvidia setup.
As non-apple unified memory options increase, many people will have more choose those
> Many AI tools are also Mac-first or Mac-only
I fail to recall AI tools Mac-only general purpose AI or agentic tools. Most of the claws, harnesses, studios and inference engines seem to be multiplatform. You can say you can run then in a Mac with a nicer UI wrapper or whatever, but "Mac-first" or "Mac-only"?
[1] https://omlx.ai/
https://www.thedeepview.com/articles/how-apple-s-decade-long...
They would even sell less than Windows Server licenses.
By the way, they are down the same path with the workstation market, now that they only top level answer is the Mac Studio.
Workstation market wants flexible towers that they can customise to their own liking and special use cases.
The main reason Swift exists for Linux, is that app developers need to have servers somewhere, and if they want to share Swift code with the backend, well it isn't going to be on macOS Server.
but Apple needs to change the licensing model, currently you are allowed to run only 2 macOS VMs for every physical one you buy
So the ad free Apple on device experience will be welcome.
For example Apple radio is a free product with no ads, Apple TV and Apple Music don’t have an ads supported tier.
I said as they don’t push ads as much as others, the customer may be better
As ads focused services are designed to keep you there as long as possible rather than delivering what you are actually want quickly
1-2% is rounding error.
That said, Siri seems a little bit better now - my subjective opinion. It is a little bit less frustrating.
And the voice is still a poor text to speech model, very far behind GPT live.
For as much as I dont like this aspect of modern computing, I understand why it is done from a technical perspective. Power, heat, and performance are all "better" when ram is on the motherboard vs in a "stick".
These execs are so out of touch they believe Apple hardware to be "a system that's under their control", how does it come to this? Besides, a VM without bi-directional sharing of data gives you pretty much the exact same thing.
Did hundreds/thousands of developers really go out there and bought Mac Minis just because one prominent technology semi-celebrity happens to have used a Mac Mini for the development of their thing? Seems bananas people would spend hundreds on monies on something they barely grasp how it works.
And all of that because Tim Apple fears any feature that could mean people could have less than one iDevice per person.
Host resolution automatically matches that of client, image quality is great, framerate is decent, latency is minimal. The host creates virtual screens for the connection so connected screens don’t light up and the machine remains locked to anybody accessing it physically too, which is a nice privacy assurance.
> “He also described a shift toward running AI locally rather than in the cloud – a move motivated by privacy, security, and the rising cost of inference as agents consume more tokens.”
Classic Apple. No more just beating the “security and privacy” drum, now its “tokens are expensive!”
<neanderthal voice/> Cloud scary. Cloud expensive. Mac good. Buy Mac!
> “He also singled out what he calls ‘transparent AI’ on iPhone and iPad, referring to features scattered throughout the operating system and third-party apps that work quietly without announcing themselves as AI.”
<neanderthal voice/> Apple use AI, Apple just not say it. Apple smart, not lagging behind industry! Buy iPhone!
How about you invest in developing your own models, correctly? And provide a secure and private inference cloud service on your fancy Apple silicon? And integrate that into your platform so Siri gets smarter without you farming queries out to Google Gemini? Bill me for it in iCloud+ I’ll probably pay for those tokens.
Was that so hard?
Or phrase it in a very similar ask, why don't they invest in power plants? The model space is truly crowded, what do they gain or recover suppose they are SOTA? Across the Pacific they are pumping out free models that are only 6-12 months behind. What business sense does it make for Apple to develop their own models?
https://machinelearning.apple.com/research/introducing-third...
they just suck
I agree, apple shouldn't invest in their own models. But they should have close to the best inference + end user design.
AI features not being constantly shoved in my face and just selectively silently integrated where it’s most useful is preferred to what the rest of the industry has been doing, too. I think most of us are pretty sick of AI getting tacked onto things that don’t need it and then given prominent promotion and UI positioning, potentially at the cost of features we actually use.
They could be doing more, sure, but directionally this all seems fine?