Apple wants AI to run directly on its hardware instead of in the cloud
arstechnica.com
arstechnica.com
Yes, OpenAI gets a lot of attention to the extent LLM's could displace search - except search is based on ads from links and selling data, which are even less promising for AI.
"Big" AI now is stuck with organization-defining cloud bills for training and huge players are scrambling to push software into hardware. They are barely getting started on the integration.
OpenAI has been talking about AGI as it lines up commercial partners across the globe and up and down the stack. But AGI is no more realistic than crypto taking over from central banks. It's possible in theory, but not in fact.
Meanwhile, Apple has had neural processors on devices for 4+ years; AI features have figured in every subsequent marketing campaign. And VisionOS augmented reality provides a whole new domain for AI utility, explicitly targeted not just for play, but also for work, as remote work becomes the rule instead of the exception. Plus, Apple is the only safety- and privacy- enforcing ecosystem in existence.
At some level it feels like the pearl clutching about GPT4 as it relates to AGI is astroturfed.
“How could you trust random people to do this safely? It needs to be illegal and something only we can be trusted to make money o… I mean use responsibly.”
It’s an attempt at regulatory capture.
Well, who said they're just gonna abandon it altogether. I'm not going to. Those of us that make use of all the google operators sure as hell ain't abandoning it. If anything, it will provide augmented results where appropriate and explicit results where appropriate. Otherwise, we'll switch to a competitor that does.
Apple’s so good at pushing out these kind of “little” features that immediately feel like a kind of bare minimum a system should do. Having to use a system where these two things weren’t automatic, quick, and practically flawless would feel like taking a big step back.
See also: Live Photos. I wouldn’t even consider any kind of phone or camera without that feature, now that I’ve experienced it.
[edit] context of how good it is: some days back there was a link here that went to an image with a bunch of text on it. I followed the link, read it, and did my usual compulsive-text-selection thing I do when reading on the web.
I didn’t realize I’d been looking at an image, rather than html, until I came back to the thread and read complaints about it. Double checked, yep, it was an image.
Either way I also get that feature on the laptop, where it's more useful to me overall.
Works with Japanese on my MacBook Pro. I use it quite frequently.
Just like Siri/voice recognition I’m sure Apple will add additional scripts over time.
https://www.apple.com/ios/feature-availability/#live-text-li...
A recent iOS update greatly improved identification. It would often get the breed of our dogs wrong. Recently it was not only was more accurate, it started being able to identify both dogs in a single image, which I'd never seen before two days ago.
A Flickr contact posted a picture of an animal I'd never seen or heard of earlier today. I took a picture of my monitor, and it was able to identify it as a coati.
Have you had non-Apple devices?
Apple is always "late to X". Their approach to all things ML so far has always been "can we run this on device". And very few, if any, of the latest LLMs can be run on the device reliaby, and without sacrificing device resources.
> OpenAI has been talking about AGI
Of course they have. Precisely because they "line up commercial partners"
---
Edit: previous version said, confusingly, "AI is always late to X". Also fixed typos and cleaned up sentences
You mean the company that wanted to ship CSAM detection for iCloud uploads with iOS, which was only abandoned because of the huge backlash? Or the company that's implementing privacy features which neatly don't affect its own apps, effectively boosting their ad platform? [1]
[1]: https://adguard.com/en/blog/apple-tracking-ads-business.html
Apple wasn't required to implement any CSAM detection, so why did they do it? Saying it was pushed onto big tech by governments when there is no law demanding it feels apologetic to me.
As things were, Apple was able to serve up peoples images when presented with a warrant.
If they went to end to end encryption for photo libraries they wouldn’t be able to do that. And I can basically guarantee something would happen and there would be claims about how “Apple protects child molesters”. It’s just too juicy, and we saw the ones about Apple protecting terrorists by not unlocking phones.
So I think they tried to come up with a solution. Something that would let them do end to end encryption but be able to show they’re not protecting child pornographers.
If someone wants to store stuff on YOUR servers YOU get to inspect it for whatever you want for whatever reason you want, arbitrarily and/or capriciously within the terms of whatever agreement you enact with your users. And the law, of course. If you don't want to inspect it, that's fine and that's your business. If you do want to inspect it, that's fine and that's your business.
If users want to keep stuff on their device you don't get to inspect it.
This isn't complicated.
If it's E2E encrypted, I don't think it'd be able to identify that during transit, so it had to have been locally done.
Thought it was interesting, and a nice QoL update. Excited to see how this can grow if they introduce more things to be processed with AI on the hardware.
Sometimes my lowly iPhone X gets super hot during charging. Next morning I either get the facial recognition updated, a new memory generated or something else has happened. It's nice, confidence evoking even. The features you like are run on your device, with your data not leaving your device, on its free time.
This is the AI I like. Personal, confined and private.
If we can combine this with a neuro bill of privacy/rights [0] then the future might not be as bleak as it seems.
[0] https://www.cirsd.org/en/horizons/horizons-winter-2021-issue...
> Sean Carroll & Nita Farahany on Ethics, Law, and Neurotechnology
https://www.preposterousuniverse.com/podcast/2023/03/13/229-...
They do stuff on-device because they want to keep selling expensive devices.
It's fantastic for them that they sell the hardware, unlike Google/OpenAI that have to spend on hardware.
Yes. Image description has been done on-device for quite a while now, and it is great. You can look for pictures using a description, and you can look for any text in any picture as well. It works very, very well.
https://findthatmeme.com/blog/2023/01/08/image-stacks-and-ip...
I had always assumed that search box in Photos.app would be useless, but once in desperation I typed something in and it found exactly what I was looking for (a photo of one of my kids' passports I knew I'd taken but could not for the life of me remember where).
Now I use it more and more, it makes my tens of thousands of photos dramatically more useful to me. That time we at at the restaurant with the big shrimp platter? The airbnb we stayed at where I took a picture of the cool coffee table? All right there.
I’ve used the feature to find my cars VIN, passport numbers, random places etc.
It also has other accessibility features like helping you find doors.
You may need to turn something on in settings before all of this shows up though, I don’t remember.
It’s new for iOS 17.
I've been running it on an iPhone 15 using this app: https://llm.mlc.ai/#ios - App Store link: https://apps.apple.com/us/app/mlc-chat/id6448482937
The same team have an Android version too which I've not yet tried myself: https://llm.mlc.ai/#android
On my iPhone it works, and provides very decent performance. The only catch is that it needs pretty much the entire phone's memory to work - so if you switch out of it to another app and back again it resets the state and has to load the model from scratch.
This has been my experience anyway ever since upgrading to iOS 16 (and now 17) anyway. Everything is always paging, even on the latest hardware.
I've tried really hard to switch to Apple Maps but even after three months it simply can't seem to remember I basically drive to 5-6 addresses on a weekly basis. It's always to the parking lot near the office on Monday and Wednesday morning, daycare to pick up the youngest on the way back, swimming pool on Saturday, etc. etc. but navigating to those addresses is always a pain (even selecting it from recent addresses in CarPlay) and it did sent me to the same street as the daycare in a completely different town several times now.
While I value not giving more data to google, after three months I decided it simply wasn't worth it. I found the experience (especially the first time it sent me to a completely different town I never go to and I didn't catch it) pretty infuriating.
One nit with this is that I haven't been able to make it automatically transition into navigation.
Based on this demonstration of a local LLM, with optimizations, it definitely the way of the future on apple devices at least.
I see this also as a privacy win as I don't want any of my data being used for training. I also like to have control over what models I use and the ethics the models follow.
Compare that to cloud, you can run the latest and greatest model and effectively unlimited storage.
Hell, that’s likely what will happen anyway. The models run on device with cloud assistance.
Welcome to the future.
Edit: I see a few comments about hundreds of GB. That works for storage, but DRAM consumes battery just by existing, and that's been keeping RAM sizes from increasing the way storage has.
A copy of GTA V takes like 100 GB.
These models all have to be fully loaded in memory to run. I think the max RAM on an iPhone is 8GB, max storage is 1TB, so we're talking 2 orders of magnitude more ability to store models vs actually running them.
And of course you'd really rather not use the full RAM for inference.
It thinks about 5-10 seconds before answering and writes the answer about the speed I can read it.
The current top of the line iPads are more powerful than this and an optimised modern iPhone implementation using the ML hardware should be faster than my current setup.
Just like there is 1 app store and 1 assistant, Apple will fight forcefully to make sure there is 1 AI model.
If EU app stores fill up with pornography and terrorism apps, you can imagine their move will be decried as a debacle and used by Apple to demonstrate how the EU failed by forcing them to open up
Apple has no right to decide what users do with the hardware they sell.
I really want to start playing hardcore porn games on my PS5 as soon as possible. Sony has no right to decide what I do with my hardware, right?
Sure.
I really want to start playing hardcore porn games on my PS5 as soon as possible.
And if somebody figures out how to do that, they should be able to.
Sony has no right to decide what I do with my hardware, right?
Pretty much, yes.
Absolutely.
> Sony has no right to decide what I do with my hardware, right?
That is correct.
Or is there a common API for llm constitution (or however it is called)?
They’re not doing anything to censor any of this.
Nothing would need to change in the existing terms and conditions with developers: we only accept binaries that use our official APIs.
The Core ML API for using the neural processing features of Apple's chips has been in there for a while and plenty of developers are using it.
It doesn’t benefit Apple to restrict access to silicon.
- have “Hey Siri” send audio to your code
- have your code reliably listen in to the microphone so that your code can listen for its own trigger phrase (AFAIK you can start an audio recording session, and have that continue while your app is in the background, but that requires users to open your app first. I also am not sure that recording will reliably survive device sleep and automatic app shutdown because of memory pressure)
I think #1 is fine. It doesn’t seem fair to me that Apple should be forced to allow other developers link “Hey Siri” to their voice assistant because that assistant may be detrimental to the value of the “Hey Siri” mark or Apple’s values, for example because it is racist, sexist, homophobic, etc.
The second IMO isn’t fine, but I can see Apple argue they are in their rights to have some control over how that rigger phrase gets detected. “Hey Siri” uses a small on-device network specifically trained for that, in order to conserve battery life (https://machinelearning.apple.com/research/hey-siri), and they may argue having good battery life is essential to their brand.
My iPad Pro has 16G RAM and I run a 13B model on it.
Things are moving fast. About 6 weeks ago I bought a 32G Mac Mini and the models that I can now run 6 weeks later are a huge improvement (e.g., mistral, solar, mixtral:8x7b-instruct, Phi, Yi, etc.) and for local RAG applications and various experiments I feel like I could live with what I have now for a while. My expectations for the future are high!
Interesting read. "Create a new upgrade cycle" is a phrase that's almost too on the nose, but will be interesting to see if the AI-powered feature set genuinely justifies an upgrade. I have a hard time imagining someone wanting to shell out 4 digits to be able to run ex. stable diffusion locally if they weren't already in the market for a new phone, but maybe I'm underestimating the AI hype wave right now. I could see predictive typing or improved voice assistants being the two biggest opportunities for a killer feature, but even then not sure if the value proposition would change that much, especially if the battery life takes any hit whatsoever.
"Hey Siri, good evening" - "I turned on the `good morning` scene."
"Hey Siri, timer five minutes" - "I set an alarm for seventeen hundred hours"
"Hey Siri, weather" - turns on the lights in the bathroom
I wish I was joking or exaggerating... The last thing I need is an LLM, just please fix the obvious problems and then maybe iterate on that
Dude. How about optimizing for the option that I've actually communicated with this decade, and 4 times already today?
In the HomePod settings there is a section for "Personal Requests". When enabled, adding reminders will then go to the device / account of the person speaking. I'm not sure what the behavior is when set to off. You can change the setting for each HomePod.
I think I figured out the issue with this one, finally: Siri understands “Set a timer for 6:30” as a request for a timer for however far in the future 6:30 is. And it also understands “set a timer for 6” similarly.
And Siri doesn’t seem at all intelligent about when to stop transcribing.
So Siri has a rather large chance of interpreting “set a timer for 5 minutes” as “set a timer for 5” and ending up with utter nonsense.
A LLM that jointly handles language and audio might do much better.
I hadn't noticed the improvements until I started using my apple watch a lot... I had given up on it a few years back, I guess.
https://www.notebookcheck.net/MrWhosetheboss-video-reveals-G...
> Google's Pixel 8 Pro Tensor G3 off-loads all generative AI tasks to the cloud
I’d like my data to stay local.
I probably don’t need cloud processing for any regular usage.
The privacy to life benefit equation is broken in society atm.
This will change.
Honestly, Apple's business model is harder if everyone wants to upgrade to a certain new version. The M1 Macs were so good, everyone upgraded their Intel Macs and then mac sales dropped. It's better logistically if the sales are more predictable and consistent year-to-year.
If you can use this tech to make a feature that they really want you can trigger upgrades that otherwise may not happen for many more years. We’ve seen that before with things like when Apple first released a bigger phone. That triggered a LOT of “early” upgrades that wouldn’t have otherwise happened.
I'm not saying Apple couldn't do the same thing, but the model alone would take up so much space on the phone that it would not be practical. I can see Apple making something closer to a very fine tuned mistral 7B than GPT4.
But I'm sure in a year this comment will be wrong.
But 1, it’s still a profitable business to be in, so why shouldn’t OpenAI go there? And 2, more importantly, as hardware and software become more efficient, in the future on-device might be good enough to put up a fight and start eating the cloud. Seems risky not to go there.
1) They think the existing hardware is powerful enough.
2) The utilization of the ANE doesn't justify increased resources.
3) They plan to re-generalize AI compute through things like vector operations.
4) (the most pessimistic) They are saving large increases for future releases when they need to force upgrades.
I could see the math for on-device just not working for things like massive LLMs. The amount of silicon you'd need to make it possible would be large, and the frequency of use low. The same silicon for one person's phone could likely support dozens if it were in a datacenter instead.
Trying to map neural nets onto graphics seems reasonable, but personally I would bet on neuralizing graphics to be the better strategy long term.
So today's on-die AI accelerators were, as far as I can tell, a cautious bet. It turns out caution was warranted, because LLM's need large amounts of memory bandwidth and capacity far more than they need a specialized neural compute engine.
However, a big factor motivating these companies to promote "on-device" inference is the operating cost.
While the compute and data requirements for training many models exceed the resources of almost everyone but these titans, it is still expected that when these models will be used for inference, and the total costs of inference over time, will far exceed the cost of training.
LLM in a Flash: Efficient LLM Inference with Limited Memory - https://news.ycombinator.com/item?id=38704982 - Dec 2023 (51 comments)
That eats into Tim's precious margins, but it seems like a price worth paying if they want to do exciting, useful things with this tech.
Adding more would certainly be a price/margin issue. Plus DRAM uses battery even if sitting unused, and battery is precious on a phone. Unless they came up with some sort of hot-plug style option to be able to turn off unused RAM chips completely, which would be cool.
Plus it would mean Apple could improve it with OS updates without app makers having to update and ship app updates every single time.
They probably make peanuts on advertisement and it's not their core business at all. They have such good PR potential. They can easily paint Google/Meta as blood sucking data vampires that read all your emails and look at all your dick pics. They could paint Android as a data hoovering OS that only works for the benefit of advertisers.
Meanwhile Apple.. The good guys.. Encrypt all your data, don't inject JS on every web page , protect you from ads with Safari, and do compute on-device
As a user you can choose to lock (almost) everything in iCloud so that Apple can’t see it, but that comes with the trade-off that they can’t help you if you lose your access. It will be permanently gone.
Isn’t there supposed to be a version of Gemini specifically for Pixel phones? It seems like Google see this as a USP for their hardware as well. I wonder how companies like Samsung are thinking about it.
Just my naive thoughts
As AI keeps catching up more and more with human intelligence, will it also need more and more hardware? Can it achieve human and superhuman intelligence and still run on a device like the iPhone, which weights about 170g? The human brain is 8 times heavier.
Maybe in 24 years the iPhone 39 will have the computing power that OpenAI has today? Hopefully our compression gets much better in the meantime.
But they also have to carry around their resource gathering and self-repair systems, something phones don't have to worry about.
Why are you comparing the weight? I'm sorry but this is a bizarre comparison. This isn't even apples to oranges, this is apples to a telephone pole.
I'll also throw in that a single Nvidia H100 is 1200 grams. Unless you have a "Bracket with screws" which will add 20 grams (who wouldn't want an extra 20grams of intelligence?).
Like 70%+ of the human brain is water. The human brain needs a massive network of systems to transport nutrients/oxygen which is irrelevant to logical processing.
Similarly the majority of the iphone weight is the battery and frame. The weight of the processing chip is _grams_.
Besides the conflicting variables with the weight, the way that ML works on a physical level is completely different from the human brain.
If you build "AI" into an OS, and that "AI" is on the cloud that means you are requiring your users to upload their data in plaintext.
PyTorch came from PCs running on GPUs with their own pool of dedicated memory and a slower link to main memory.
I know people have been speeding it up a lot on Apple processors, but I’m not sure it’s a good indicator of how ready the iPhone is for running models for end user programs.
obviously its augmented by any Galaxync.
Other kinds of transformers could too
Fix Apple autocorrect and detect replacements for misspelled words better
Anyways, i had fun building a Linux server to stick an nvidia GPU in. I’m sure it’s going to be pretty common over the next few years as devs realize they can’t just do everything with a MacBook. (Or maybe apple will add support for 3rd party GPUs like they had with the intel MacBooks)
This was my take away from the apple AR headset; it doesn’t matter how nice the UX is if it lacks the compute to do anything interesting… and AR/VR takes a LOT of compute.
But you don’t need all that. There’s must be tons of great uses that would benefit me as a user that don’t have requirements that steep.
My devices will never compete with a data center, there will always be things that have to be hosted elsewhere. But there’s still benefit to be had.
And seems to be aligned well with Apple too: their focus on privacy, their hardware performance gains, their ai assistant products.
Apple is good at launching new product categories, ahead of competition. The apple watch had no competition for the first couple of years after launch, ipads are a class of its own, I wonder if they’re planning a similar move with ai too.
Apple has been using their Neural Engine for a couple of years now, and it’s obviously been noticed. Other companies are now starting to include similar blocks in their CPUs.
They could be a match or even better, or they could end up looking like an old integrated graphics card against a discrete GPU.
Time will tell us how much of a difference all of this makes.
I do wonder what foundation model Apple will eventually use.
Both training their own as well as licensing a 3P one seem at odds with their claimed data privacy stance.
My bet though is a licensed 3p model that’s fine tuned locally with on device data.
They'll release a hardware co-processor similar to Google's when they're ready to ship this.
iMessage, apple photos, Siri inputs? These all seem problematic beyond the standard open crawl type datasets. They don’t have a web search corpus.
It’s a very good question though. That could make all the difference, huh.
But they're not competing with cloud AI. Why would a person need to go to the cloud to give you a reminder or download an app? They're competing against the current local assistant, Siri.
Large models are great but they can't fit on 8 or 16gb of ram. And that's a very big deal.
You can have the basic stuff on-device with the "smarts" of an LLM that can have conversations with the user and have context to previous questions.
The other stuff can be fetched from the cloud (with the user's permission OFC) and optionally saved locally.
You can't make a cloud model work 100% locally for privacy or during connection issues.
Who pays (and how) if Apple starts doing this for you, on their equipment?
Would you pay $5/mo for it?