The hunt for the M1’s neural engine
eclecticlight.co
eclecticlight.co
I was waiting for the Mac Studio to drop to come to a conclusion, since it's plausible one ANE could've been disabled for power reasons on the laptops... but with both secondary ANEs in each die in the M1 Ultra off, and with no reports of anyone seeing the first ANE being disabled instead (which could mean it's a yield thing), I'm going to go ahead and say there was a silicon bug or other issue that made the second ANE inoperable or problematic, and they just decided to fuse it off and leave it as dead silicon on all M1 Max/Ultra systems.
Could it be overheating if the ANE's are not used the right way?
Sony pulled a similar trick on their A7 series cameras, and enabled more advanced AF features just with a firmware upgrade. It made the bodies "new" and pushed them at least half a generation forward. It's not the same thing, I know, but it feels similar enough for me.
A big design with many cores, such as CPU or GPU cores, may have manufacturing defects that makes one or more cores bad. Or it may be on one side or another of a tolerance range and not be able to work at higher power or higher frequency. These parts may get "binned" into a lower performance category, with some cores disabled (because a flaw prevents the core from working) or with reduced maximum performance states.
These are still "good" parts, and can be sold at a lower cost with lower performance, while the "better" and "best" parts will pass more tests and be able to have more or all portions of the chip enabled.
So it's not so much to work around a "bug" which might be a common flaw to all part designs, rather to work around manufacturing tolerance and allow more built parts to be useful rather than garbage.
Binning is irrelevant to hardware bugs.
Any two-ANE design would have a lot of control logic that has to be right, e.g., to manage which work gets sent to which ANE, which cache lines get loaded, etc. It's easy to imagine bugs in this logic which would only show up when both ANEs are enabled. So it's likely that there is a chicken bit that you could use to disable one of the ANEs and run in single-ANE mode.
I can take my old D70s today and take beautiful photos, even with it's 6MP sensor, however a newer body would be much more flexible.
> I can take my old D70s today and take beautiful photos, even with it's 6MP sensor
I suspect if you do a wedding shoot with a 6mp interchangeable lens camera, some customers are rightly going to ask questions when you hand over the work... Of course professional photographic equipment gets obsolete - even lens systems get deprecated every 20-30 years too. Newer sensors have vastly more dynamic range than the d70s among other image quality benefits.
I think you argument holds water much more strongly in the context of amateur users, where for sure you can keep getting nice images from old gear for a long time.
Unless you're printing A3 pages, getting gigantic pictures, or cropping aggressively, D70s can still hold up pretty well [0].
> even lens systems get deprecated every 20-30 years too.
Nikon F mount is being deprecated in favor of Z because of mirrorless geometries, not because the lenses or the designs are inferior (given the geometry constraints). Many people still use their old lenses, or nifty fifties are still produced with stellar sharpness levels. I'm not entering into "N" or "L" category of lenses of their respective mounts. Not all of them are post 2000 designs, or redesigns, and they produce extremely good images.
> Newer sensors have vastly more dynamic range than the d70s among other image quality benefits.
As a user of both D70s and A7III I can say that, if there's good enough light (e.g day), one can take pretty nice pictures with a D70s, even today. Yes, it dies pretty fast when light goes low, or it can't focus as fast, or can't take single shot (almost) HDR images (A7III can do that honestly, and that's insane [4]), but unless you're chasing something moving, older cameras are not that bad. [1][2][3]
> I think you argument holds water much more strongly in the context of amateur users, where for sure you can keep getting nice images from old gear for a long time.
Higher end, action oriented professional cameras are not actually built with resolution in mind, especially at the top end. All of the action DSLRs and mirrorless cameras up to a certain point are designed with speed and focus in mind. You can't see A7R or Fuji GFX series in weddings or in stadiums. You'll see A9s, Canon 1D or Nikon D1 series cameras. They're built to be fast. Not high res.
A wedding is more forgiving, but again a high MP camera is not preferred since it's more prone to vibration blurring.
[0]: https://www.youtube.com/watch?v=ku3lT8MjyFM
[1]: https://www.flickr.com/photos/zerocoder/41901384135/in/album...
[2]: https://www.flickr.com/photos/zerocoder/28459579257/in/album...
[3]: https://www.flickr.com/photos/zerocoder/39910477633/in/album...
[4]: https://www.flickr.com/photos/zerocoder/33984196648/in/album...
> Unless you're printing A3 pages, getting gigantic pictures, or cropping aggressively, D70s can still hold up pretty well
You even qualified that later with "if there's good enough light" and "unless you're chasing something moving". No, a D70 won't work well for wedding photography. Yes, people shot weddings with much slower film. They don't anymore because, like the D70, slow film is obsolete. People shot weddings with manual focus lenses too and the D70 is awful for MF lenses from the tiny viewfinder to the lack of support for non-CPU lenses. When the D70 was a current product some people did (make no mistake the D70 was never marketed as a pro body) simply because the D70 was on par with its contemporaries.
> Nikon F mount is being deprecated in favor of Z because of mirrorless geometries
Even within the scope of the F mount the D70 is obsolete — it's incompatible with new E and AF-P lenses.
> A wedding is more forgiving
Wedding photography is about the most technically challenging, least forgiving (low light, constant motion, spontaneous behavior) type of photography out there. The point you were responding to still stands – older digital photographic equipment is obsolete in a professional context while having some utility for hobbyists. Nobody's taking a D1 out to shoot sports these days. In fact most people didn't when it was new because Nikon's autofocus was so far behind Canon's.
No.
> You even qualified that later with "if there's good enough light" and "unless you're chasing something moving". No, a D70 won't work well for wedding photography.
I didn't intend to say "You can shoot weddings with a D70s". I just wanted to say, a D70s can take beautiful photos, even today, with today's standards, that's all. Even Nikon didn't position D70s for that kind of action when it was brand new.
> People shot weddings with manual focus lenses too and the D70 is awful for MF lenses from the tiny viewfinder to the lack of support for non-CPU lenses.
As a person who shot MF on both film and D70s, I tend to disagree, but that's not a hill I'd prefer to die on, at least within the borders of this comment box.
> Even within the scope of the F mount the D70 is obsolete — it's incompatible with new E and AF-P lenses.
I think being able to select between this much [0] of lenses is enough for most people.
> Wedding photography is about the most technically challenging, least forgiving (low light, constant motion, spontaneous behavior) type of photography out there.
Sorry, no. I shoot at tango nights. You have much more freedom in weddings. You can use flash, come close, people expect you, etc. A "low light" wedding situation is "what you can expect in a good tango night". You can't use flash, use lenses slower than f/2.2 (or more specifically ~t/2.5), because even a latest generation sensor will just choke, you can't use big lenses and be a distracting element, or come close for any reason. So, no. A wedding is not a piece of cake, but much easier from a technical point of view. Weddings have their problems like the distance/area you have the cover, the equipment you have to carry on you, duration, and storage and energy logistics, I agree, but it's not as challenging in terms of light or camera capabilities.
> ...older digital photographic equipment is obsolete in a professional context while having some utility for hobbyists.
I'd rather rephrase this. Newer photographic equipment is much more capable and makes professionals' life much, much easier. Obsolescence is something different in my eyes, and it's not the same as "not as useful today as of yesterday". Even an analog Pentax MF body is not obsolete in today's photographic world, even for professionals. It might not be their everyday body, but's it's neither useless, nor obsolete.
[0]: https://www.lensora.com/lensesfor.asp?camera=nikon-d70s
Take a look at what equipment qualifies you for professional support (NPS). Even the D3xxx series qualify. The D70 does not. The D70 does not because it is considered obsolete by Nikon.
> As a person who shot MF on both film and D70s, I tend to disagree, but that's not a hill I'd prefer to die on, at least within the borders of this comment box.
Sure, I shot manual lenses for quite a while with a D200 (with and without a split prism screen). DSLRs (autofocus bodies from any manufacturer really), especially lower end ones, are not well suited to manual focus lenses. Fast lenses like you would need for a wedding exacerbate this as you simply cannot resolve enough contrast to nail the focus with a big aperture.
There's no way around the fact that the D70 has a small viewfinder (95% coverage, sure, but only 0.75x magnification). The D200 had a 0.94x viewfinder and even that was well more challenging than the old ME Super I cut my teeth on.
> I think being able to select between this much [0] of lenses is enough for most people.
The D70 is still obsolete.
> Sorry, no. I shoot at tango nights. You have much more freedom in weddings.
Sorry, no. You can re-do a tango night. Try to redo a wedding shot and you'll be dealing with bridezilla at best. And, sure, you can use a flash (assuming the venue is okay with it) just like you can create unhappy customers. Even the brighter, outdoor weddings I've been to (not as a photographer thank god) have way more uneven lighting than any sort of indoor dance venue.
> I'd rather rephrase this. Newer photographic equipment is much more capable and makes professionals' life much, much easier. Obsolescence is something different in my eyes, and it's not the same as "not as useful today as of yesterday". Even an analog Pentax MF body is not obsolete in today's photographic world, even for professionals. It might not be their everyday body, but's it's neither useless, nor obsolete.
A D70 is still obsolete. Your old Pentax will shoot just fine with whatever K (or whatever depending on the age) lenses. Your D70 will function with a subset of F mount lenses and both older and newer lenses won't work. There are probably exceptions for those who are wedded to the novelty sensors Nikon used in some of their older cameras but no pro is going to be shooting with a D70.
I mean, look, you can use old manual focus (but not pre-AI) lenses with the D70. You'll have to bring a separate light meter though (or just chimp it) because the meter does not work at all with non-CPU lenses. Stop down metering? Nope. Nada. Much, much easier is a major understatement. What pro work are you going to do without an in-body light meter? Studio stuff? The D70 is obsolete.
Here's the list of DX bodies that Nikon considers pro gear: D500, D300S, D7500, D7200, D7100, D5600, D5500, D5300, D3500.
...maybe Apple disabled it for a reason, y'know?
It would be reasonable to assume that if both engines work, then the second is always the one to be disabled. Therefore to have the second enabled you'd need to find one where the first engine has failed, and has no other chip killing faults. Depending on the yield TSMC gets, these could be quite rare, so you'd have to have quite a large survey to find them.
Or as other people have noted, it could be an errata meaning the second core is broken, as this isn't the only possible reason.
The linked posting notes that they were able to get the ANE to draw 49 mW. Is this such a significant amount of power that its worth permanently disabling for laptop power draw? Or is there likely much more power being used elsewhere to support ANE in addition to the 49 mW that can be measured directly?
[0]: https://github.com/geohot/tinygrad/tree/master/accel/ane
https://www.youtube.com/watch?v=mwmke957ki4
https://www.youtube.com/watch?v=H6ZpMMDvB1M
I wonder why Apple didn't provide low-level API's to access the hardware? It may have various restrictions. I recall Apple also didn't provide proper API's to access OpenCL frameworks on iOS, but some people found workarounds to access that as well. Maybe they only integrate with a few limited but important use cases, TensorFlow, Adobe that they can control.
Could it be that using the ANE in the wrong way overheats the M1?
Using a high level API probably makes it easier to implement a software version for hardware that doesn't have the neural engine, like Intel Macs or older A-cores.
[1] Although this probably starts a long conversation about various GPU and ML core APIs and quite how low level they get.
I honestly don't know of a single company offering custom machine learning accelerators that let you do anything except use Tensorflow/PyTorch to interface with them, not a chance in hell any they actually will give you the underlying ISA specifics. Maybe the closest is, like, the Xilinx Versal devices or GPUs, but I don't quite put them in the same category as something like Habana, Groq, GraphCore, where the architecture is bespoke for exactly this use case, and the high level tools are there to insulate you from architectural changes.
If there are any actual productionized, in-use accelerators with low level details available that weren't RE'd from the source components, I'd be very interested in seeing it. But the trend here is very clear unless I'm missing something.
Oh, and they have an open-source UM software stack for those but it's really not usable. Doesn't allow access to the systolic arrays (MME), only using the TPCs is just _starting_ to enumerate what it doesn't have. (but, it made the Linux kernel maintainers happy so...):
https://github.com/HabanaAI/SynapseAI_Core#limitations (not to be confused with the closed-source SynapseAI)
Frankly I kind of expected the whole result of that kerfuffle to just be that Habana would let the driver get deleted from upstream and go on their merry way shipping drivers to customers, but I'm happy to be proven wrong!
One of the earliest lessons along this line was Itanium. Itanium exposing so much of the underlying architecture as a binary format and binary ABI made evolution of the design extremely difficult later on, even if you could have magically solved all the compiler problems back in 2000. Most machine learning accelerators are some combination of a VLIW and/or systolic array design. Most VLIW designers have learned that exposing the raw instruction pipeline to your users is a bad idea not because it's impossibly difficult to use (compilers do in fact keep getting better), but because it makes change impossible later on. This is also why we got rid of delay slots in scalar ISAs, by the way; yes they are annoying but they also expose too much of the implementation pipeline, which is the much bigger issue.
Many machine learning companies take similar approaches where you can only use high-level frameworks like Tensorflow to interact with the accelerator. This isn't something from Apple's playbook, it's common sense once you begin to design these things. In the case of Other Corporations, there's also the benefit that it helps keep competitors away from their design secrets, but mostly it's for the same reason: exposing too much of the implementation details makes evolution and support extremely difficult.
It sounds crass but my bet is that if Apple exposed the internal details of the ANE and later changed it (which they will, 100% it is not "done") the only "outcome" would be a bunch of rageposting on internet forums like this one. Something like: "DAE Apple mothershitting STUPID for breaking backwards compatibility? This choice has caused US TO SUFFER, all because of their BAD ENGINEERING! If I was responsible I would have already open sourced macOS and designed 10 completely open source ML accelerators and named them all 'Linus "Freakin Epic" Torvalds #1-10' where you could program them directly with 1s and 0s and have backwards compatibility for 500 years, but people are SHEEP and so apple doesn't LET US!" This will be posted by a bunch of people who compiled "Hello world" for it one time six months ago and then are mad it doesn't "work" anymore on a computer they do not yet own.
> Could it be that using the ANE in the wrong way overheats the M1?
No.
As for transparency in hardware, it probably will become more transparent once Apple feels that it is done and a finished science. They don't want to repeat Itanium.
Especially when Apple is involved. Hell there are still people who see them as beleaguered and about to go out of business at any moment :p
Or not. It's their hardware, they just won't be selling any Macs to me with that mindset. The only thing that irks me is when people take the bullet for Apple like a multi-trillion dollar corporation needs more people justifying their lack of interoperability.
In much the same way the App Store is an infuriating shh-don't-call-it-censorship bottleneck that gives Apple total and final control over what your (sorry, Apple's) devices can do, I wonder if political considerations represents a portion of Apple's motivation to keep things reasonably locked down. Obviously Apple can just kick apps it doesn't like out of the App Store, and binaries that would need to be downloaded and run directly on Macs is exceedingly unlikely to go viral to the same extent, so perhaps I'm overthinking things to the point of paranoia.
To be fair though, I mean. I'm mostly a bitchy nerd, too. And broadly speaking, taking the piss is just good fun sometimes. That's the truth, at least for me.
If it helps, simply close your eyes and imagine a very amped up YouTuber saying what I wrote above. But they're doing it while doing weird camera transitions, slow-mo shots of panning up the side of some Mac Mini or whatever. They are standing at a desk with 4 computers that are open-mobo with no case, and 14 GPUs on a shelf behind them. Also the video is like 18 minutes long for some reason. It's pretty funny then, if you ask me.
Maybe this works for some people. I can't knock someone for an opinionated implementation of a complicated system. At the same time though, we can't be surprised when other people have differing opinions, and in a perfect society we wouldn't try to crucify people for making those opinions clear. Apple notoriously lacks a dialogue with their community about this stuff, which is what starts all of this pointless infighting in the first place. Apple does what Apple does, and nerds will fight over it until the heat death of the universe. There really is nothing new under the sun. Mocking the ongoing discussion is almost as phyrric as claiming victory for either side.
What consumer apps actually use Neural Engines?
I think something like Photoshop, maybe. But wouldn't it just train a model and use it as regular code?
I'm interested in AI, but it's a joke to me about startups and jargon more often than not.
I feels weird to add this to all chips when I can't see that much usage.
Of course if you’re using third party libraries that don’t rely on macOS APIs this won’t be happening in your app.
https://www.theverge.com/2021/12/15/22837631/apple-csam-dete...
- Visual Lookup - Animoji - Face ID - recognizing "accidental" palm input while using Apple Pencil - monitoring users' usage habits to optimize device battery life and charging - app recommendations - Siri - curating photos into galleries, selecting "good" photos to show in the photos widget - identifying people's faces in photos - creating "good" photos with input from tiny camera lenses and sensors - portrait mode - language translation - on-device dictation - AR plane detection
Core ML API allows third party developers to use Neural Engine to run models
I know that C++ is out of vogue, but last time I wrote drivers in linux and OSX (10 years ago) I left with the distinct impression that it was oversold, at least compared to C. C clunks hard and C++ addresses the worst of it. I've never had to corral a bunch of overzealous junior C++ programmers, which I suspect is where the C++ reputation comes from, but Apple went down that path and wound up with something pretty decent.
IMO today it's a shrug and 25 years ago it was forward looking.
C++ when used as a better C, really is better.
Of course, the OSX driver code is going on 25 years old, so it's not evidence of anyone's opinion that C++ beats Rust for OS dev.
Basically it's used by anything involving machine learning:
- Speech recognition
- Face recognition
- Visual lookup (image recognition)
- Live Text (OCR)
It allows all these functions to be performed efficiently on-device rather than shipping data off to the cloud.
The second part of the chain is that you also tailor the operations the die area support to those operations necessary in a typical neural network, further optimizing the chip.
The end result is very power-efficient execution of neural networks, which allows you to beat a GPU or CPUs power curve, improving thermals of the core, and in the case of mobile devices, optimizes battery usage.
Those efficiency wins are almost certainly worth it even if third-party developers don't use it much.
Hmm, it seems like there's also a new API that can use the neural engine sometimes, "ML Compute". But only for inference? https://developer.apple.com/documentation/mlcompute/mlcdevic...
Reading this makes me wonder if it's not just a placeholder for some kind of intrusive system that will neural-hash everything you own, but I'm sure I'm just being paranoid.
That would seem like the logical conclusion. Perhaps there are hardware bugs/shortcomings that makes it very hard to use for the neural network API. Perhaps the software team is just behind and still building that.
It however turns out that a _lot_ of customer apps today don't use those accelerators at all.
(and about the attempt at using BNNS functions, that's not offloaded, it runs on the host CPU cores w/ the AMX tightly bound accelerator)
The Asahi project is putting a deliberate low priority on the ANE but I have seen some other small reverse engineering attempts.
I think some use of the ANE outside of Apple APIs will be possible soon.
So when running a neural net iOS / macOS needs to make a decision about where to run each net. Even if a net has layers in it that are a perfect match for the ANE, there’s still a trade off from having move the workload back and forth between a CPU core and the ANE (although the unified memory should eliminate a big chunk of this cost).
It might be that in the general case it’s not worth the latency hit from using mixed processors when running net that isn’t 100% ANE compatible, or could just be that Apple haven’t got round to implementing the logic needed to gracefully spilt workloads across the ANE and a CPU core. Which would make sense, because they’ve got the time and expertise to ensure all their nets fit within the ANE. Something that’s difficult for 3rd party devs to do, because they don’t have access to detailed ANE docs.
It’s also actively counter-productive: if they wanted to do this sort of tracking, they could just have done what all of their competitors do and send the data straight to their servers. This is hardware (and thus expenses) which is only necessary because of their stance on privacy, and avoiding off-device work.