What we know about the Apple Neural Engine
github.com
github.com
Anyone disappointed, here be full details on everything.
What kind of things are taking advantage of this right now? It's gotta be more than just Face ID right?
What's my laptop likely to be doing with that?
I think neural engine is absolutely key to Apple's strategy. They want people to buy expensive devices and they don't want to process user data on their servers.
Users get privacy. Apple gets money. It's a pretty coherent business model.
> Users get privacy. Apple gets money.
Apple also gets users to subsidize the cost of compute indefinitely (by buying the expensive phone), rather than using their servers.
Subsidizing the smart speaker hardware and running everything on their servers hasn't worked out well for Alexa and Google Assistant.
> The Alexa division is part of the "Worldwide Digital" group along with Amazon Prime video, and Business Insider says that division lost $3 billion in just the first quarter of 2022, with "the vast majority" of the losses blamed on Alexa. That is apparently double the losses of any other division, and the report says the hardware team is on pace to lose $10 billion this year. It sounds like Amazon is tired of burning through all that cash.
Google expressed basically identical problems with the Google Assistant business model last month. There's an inability to monetize the simple voice commands most consumers actually want to make, and all of Google's attempts to monetize assistants with display ads and company partnerships haven't worked. With the product sucking up server time and being a big money loser, Google responded just like Amazon by cutting resources to the division.
https://arstechnica.com/gadgets/2022/11/amazon-alexa-is-a-co...
On the downside, we have to acknowledge that it is hugely inefficient for everyone to own expensive hardware that has to sit idle most of the time because it would otherwise drain the battery.
Where low latency is not an absolute necessity, the economic pull of the cloud will be tremendous, especially if mobile networks become ubiquitous and fast.
And it's not even particularly wasteful to produce. While there are lots of devices the amount of materials in each chip is incredibly small.
In terms of financial cost the material cost is so low that most chip vendors include features in low end chips that are disabled but shipped anyway.
The battery is a key part as well. Doing a lot of on-device processing requires a bigger battery and/or more frequent battery replacements.
Expensive hardware like millions of 5G/6G modems, stations and fiber optic cables nessesary to send vast quantity of photos and videos to the cloud for analysis?
ait is more expensive to build out bandwidth than to give everyone compute
Also average car is unused 95/% of the time, so by this principle everyone should take a bus
I'm not sure about that. Storing all data on all devices (not just user data but also the pretrained models) means higher data transfer volumes.
On the other hand, syncing data doesn't require low latency and a lot of the data can be transferred over cheaper landline connections rather than mobile towers.
Right now though the default appears to be to upload everything to the cloud right away, regardless of where the data is processed.
And there's also communication and media streaming, which consumes the vast majority of network resources.
Net net I'm not sure whether on-device vs cloud processing will make a big difference for networking costs one way or another.
>Also average car is unused 95/% of the time, so by this principle everyone should take a bus
Yes, and it is hugely inefficient. People who own a car (I don't) are doing it for the benefits it has, not because it's efficient.
That’s exactly right.
Back when dictation was done in the cloud, I could dictate all day on my iPhone no problem.
Now that it's on-device it kills my battery in a couple of hours.
The latency is absolutely improved, and continuous dictation (not stopping every 30s) is a godsend.
But it does absolutely destroy your battery life.
"Moore's Law is Dead" is a bit of a joke in the hardware space because a staggering number of smart people have made staggeringly wrong predictions of this nature (well, typically correct in a narrow sense and wrong in a broad sense). Jim Keller frequently talks about this and has a convincing theory as to why it happens: the industry is full of specialists who are all chasing one particular S-curve and fully understand to the point of conservatively believing in another S-curve or two. Inevitably, this gives them the impression that Moore's Law has just a few years of gas left -- however, it's actually a consequence of limits on human communication, curiosity, and cognition that determines the number of promising S-curves a typical engineer is thinking about. It's not a true evaluation of the supply of additional S-curves waiting in the wings. There's a limit somewhere out there but it's not really "in sight."
I'm not quite sure what people think as noone seems to be stating it explicitly, but getting the impression some commenters here have their own personal definition of Moore's Law. The actual law is narrowly defined: if you're looking to discuss things related to it "in a broad sense" cool, but that isn't what I was referring to above. I was referring to Moore's Law.
To be clear, Moore's Law is about processor manufacturing tolerances & states some pretty concrete predictions for rates of progress in the manufacturing process. It doesn't state anything related to compute (i.e what uses those processors can be put to), it's purely about the physical properties thereof.
Apple’s always-on Hey Siri wake word detection uses ANE for minimal battery life.
I’m not sure but I think android’s “Now Playing” feature (the always-on Shazam thing that shows the names of songs playing around you on your lock screen ) also uses Edge TPU for similar reasons.
Gesture recognition has been around since before the neural engine shipped and doesn't appear to be different on devices with or without it.
>Machine learning is used to help the iPad’s software distinguish between a user accidentally pressing their palm against the screen while drawing with the Apple Pencil, and an intentional press meant to provide an input.
https://yugalchoubisa.medium.com/how-machine-learning-and-ar...
- Biometrics (Face ID and Touch ID)
- Image analysis (face matching, aesthetic evaluations, etc)
- Text to speech and speech to text (smaller models on device, used for privacy/latency/reliability)
- Small ad-hoc models like Raise to Speak on Apple Watch, the Hey Siri detector (https://machinelearning.apple.com/research/hey-siri)
These things have been in phones for 5 years now and have been used from day one
I could only find a blurry YouTube video of the instruction manual for an old old heater in my house.
I paused the video on the bit I needed the guy had zoomed into and was able to copy and paste the text that I could barely read into a notes doc.
There’s no one splashy thing just lots of little quality of life improvements.
Your iPhone won’t analyze your photos unless it is plugged in. It’s power intensive primarily due to the amount of data that needs to move to classify faces in a photo. Image files are quite large these days, and even if you downscale them you still need to load them from disk and decompress them in the first place.
The scarce resource on the iPhone is RAM, not compute. The Camera’s image pipeline uses so much memory that the phone is effectively unable to multi-task when the camera is open.
Newer phones are running on faster process nodes, so that is always a boost, but the phones also have more RAM for bigger cameras.
Apple’s unified memory architecture is a key advantage, because they can ship less RAM if it can be shared across the CPU/GPU, but all mobile SoCs have unified memory for this reason, so it’s not Apple’s secret sauce, just part of it.
I think the primary constraint will be how many models can fit in memory at once, not how many can run at once.
They also don’t often all run at once: you only use the TTS model when Siri speaks, so Siri only loads it when it’s needed. But, in that case, then the latency of loading the model from disk is a concern…
In short, performance for ML is really a small subset of performance more generally, and Apple takes performance seriously because efficiency is their game. That is how good battery life pops out on the other side.
I find it a little frustrating we aren’t using the built in capabilities of iphones more in our company, i still kinda think apple tech is kind of a pariah in some circles, so we have to run with stuff that runs on cloud that costs us money over, heaven forbid something you could run on an iphone
I know the Swift version of Stable Diffusion from Hugging Face uses this too.
Or they are keeping it obscure for commercial reasons?
Or just not very competent/don't care?
Seems weird having these amazing chips and only blunt tools
Directly exposing the ANE wouldn't make much sense, as it's an IP block that changes between generations in incompatible ways.
You might not want the abstraction, but love it or hate it, that’s kind of the Apple way.
It will be very interesting to see what their next chips look like since we’re getting to the point where HW designs will reflect the rise of the, uh, transformers.
"> Can I program the ANE directly?
Unfortunately not. You can only use the Neural Engine through Core ML at the moment.
There currently is no public framework for programming the ANE. There are several private, undocumented frameworks but obviously we cannot use them as Apple rejects apps that use private frameworks.
(Perhaps in the future Apple will provide a public version of AppleNeuralEngine.framework.)"
The last part links to this bunch of headers:
https://github.com/nst/iOS-Runtime-Headers/tree/master/Priva...
So might it be more accurate to say you can program it directly, but won't end up with something that can be distributed on the app store?
I can’t find any documentation about it though just everyone working under that assumption.
Just sayin'.
Great guys.
[0] Persist with their own models running locally, how much to integrate with rest of the OS and maintain privacy moral ground, that sort of thing.