[0]: https://github.com/geohot/tinygrad/tree/master/accel/ane
[0]: https://github.com/geohot/tinygrad/tree/master/accel/ane
https://www.youtube.com/watch?v=mwmke957ki4
https://www.youtube.com/watch?v=H6ZpMMDvB1M
I wonder why Apple didn't provide low-level API's to access the hardware? It may have various restrictions. I recall Apple also didn't provide proper API's to access OpenCL frameworks on iOS, but some people found workarounds to access that as well. Maybe they only integrate with a few limited but important use cases, TensorFlow, Adobe that they can control.
Could it be that using the ANE in the wrong way overheats the M1?
Using a high level API probably makes it easier to implement a software version for hardware that doesn't have the neural engine, like Intel Macs or older A-cores.
[1] Although this probably starts a long conversation about various GPU and ML core APIs and quite how low level they get.
I honestly don't know of a single company offering custom machine learning accelerators that let you do anything except use Tensorflow/PyTorch to interface with them, not a chance in hell any they actually will give you the underlying ISA specifics. Maybe the closest is, like, the Xilinx Versal devices or GPUs, but I don't quite put them in the same category as something like Habana, Groq, GraphCore, where the architecture is bespoke for exactly this use case, and the high level tools are there to insulate you from architectural changes.
If there are any actual productionized, in-use accelerators with low level details available that weren't RE'd from the source components, I'd be very interested in seeing it. But the trend here is very clear unless I'm missing something.
Oh, and they have an open-source UM software stack for those but it's really not usable. Doesn't allow access to the systolic arrays (MME), only using the TPCs is just _starting_ to enumerate what it doesn't have. (but, it made the Linux kernel maintainers happy so...):
https://github.com/HabanaAI/SynapseAI_Core#limitations (not to be confused with the closed-source SynapseAI)
Frankly I kind of expected the whole result of that kerfuffle to just be that Habana would let the driver get deleted from upstream and go on their merry way shipping drivers to customers, but I'm happy to be proven wrong!
One of the earliest lessons along this line was Itanium. Itanium exposing so much of the underlying architecture as a binary format and binary ABI made evolution of the design extremely difficult later on, even if you could have magically solved all the compiler problems back in 2000. Most machine learning accelerators are some combination of a VLIW and/or systolic array design. Most VLIW designers have learned that exposing the raw instruction pipeline to your users is a bad idea not because it's impossibly difficult to use (compilers do in fact keep getting better), but because it makes change impossible later on. This is also why we got rid of delay slots in scalar ISAs, by the way; yes they are annoying but they also expose too much of the implementation pipeline, which is the much bigger issue.
Many machine learning companies take similar approaches where you can only use high-level frameworks like Tensorflow to interact with the accelerator. This isn't something from Apple's playbook, it's common sense once you begin to design these things. In the case of Other Corporations, there's also the benefit that it helps keep competitors away from their design secrets, but mostly it's for the same reason: exposing too much of the implementation details makes evolution and support extremely difficult.
It sounds crass but my bet is that if Apple exposed the internal details of the ANE and later changed it (which they will, 100% it is not "done") the only "outcome" would be a bunch of rageposting on internet forums like this one. Something like: "DAE Apple mothershitting STUPID for breaking backwards compatibility? This choice has caused US TO SUFFER, all because of their BAD ENGINEERING! If I was responsible I would have already open sourced macOS and designed 10 completely open source ML accelerators and named them all 'Linus "Freakin Epic" Torvalds #1-10' where you could program them directly with 1s and 0s and have backwards compatibility for 500 years, but people are SHEEP and so apple doesn't LET US!" This will be posted by a bunch of people who compiled "Hello world" for it one time six months ago and then are mad it doesn't "work" anymore on a computer they do not yet own.
> Could it be that using the ANE in the wrong way overheats the M1?
No.
As for transparency in hardware, it probably will become more transparent once Apple feels that it is done and a finished science. They don't want to repeat Itanium.
Especially when Apple is involved. Hell there are still people who see them as beleaguered and about to go out of business at any moment :p
Or not. It's their hardware, they just won't be selling any Macs to me with that mindset. The only thing that irks me is when people take the bullet for Apple like a multi-trillion dollar corporation needs more people justifying their lack of interoperability.
In much the same way the App Store is an infuriating shh-don't-call-it-censorship bottleneck that gives Apple total and final control over what your (sorry, Apple's) devices can do, I wonder if political considerations represents a portion of Apple's motivation to keep things reasonably locked down. Obviously Apple can just kick apps it doesn't like out of the App Store, and binaries that would need to be downloaded and run directly on Macs is exceedingly unlikely to go viral to the same extent, so perhaps I'm overthinking things to the point of paranoia.
To be fair though, I mean. I'm mostly a bitchy nerd, too. And broadly speaking, taking the piss is just good fun sometimes. That's the truth, at least for me.
If it helps, simply close your eyes and imagine a very amped up YouTuber saying what I wrote above. But they're doing it while doing weird camera transitions, slow-mo shots of panning up the side of some Mac Mini or whatever. They are standing at a desk with 4 computers that are open-mobo with no case, and 14 GPUs on a shelf behind them. Also the video is like 18 minutes long for some reason. It's pretty funny then, if you ask me.
Maybe this works for some people. I can't knock someone for an opinionated implementation of a complicated system. At the same time though, we can't be surprised when other people have differing opinions, and in a perfect society we wouldn't try to crucify people for making those opinions clear. Apple notoriously lacks a dialogue with their community about this stuff, which is what starts all of this pointless infighting in the first place. Apple does what Apple does, and nerds will fight over it until the heat death of the universe. There really is nothing new under the sun. Mocking the ongoing discussion is almost as phyrric as claiming victory for either side.