The Secret Apple M1 Coprocessor
medium.com
medium.com
The cited announcement here was for ARMv8-M only, the microcontroller variant of ARMv8 for embedded systems.
From the linked EETimes article:
>> Arm is opening up its instruction set to customers’ customized instructions for Cortex M cores
>> For Cortex-A cores, Arm is still a long way from offering any customizable instructions
From https://www.arm.com/why-arm/technologies/custom-instructions:
> Arm Custom Instructions are a standard feature of the Armv8-M architecture
Apple M1 is not an ARMv8-M core, it's an ARMv8-A core.
I'd very much like to understand the contractual terms under which Apple put custom instructions into their ARMv8-A core, but I don't think it is as simple as "ARM is totally fine with all licensees doing this now." AFAIK, the latest revisions of ARMv8-A still do not permit custom instructions.
Other reasons I don't think AMX is in the category of the publicly announced "custom instructions" part of the ISA:
- arm custom instruction are prohibited from accessing memory, but AMXLD* and AMXST* exist.
- arm custom instructions are prohibited from having their own register states, but AMX has amx0, amx1, amx2 registers.
https://developer.arm.com/architectures/instruction-sets/cus...
Basically, Apple doesn't use core IP and "only" licenses the ISA (which is far more expensive if my memory is not lying) and so they can do whatever they want.
The broadest and rarest type of ARM license that is publicly documented is an ARM architecture license, which allows the licensee to implement one of the ARM ISAs with their own microarchitecture.
Apple and Samsung both have architecture licenses, though IIRC Samsung recently closed down their CPU design team and so is no longer making use of their architecture license.
The normal terms of architectural licenses do not permit the licensees to do whatever they want, they are typically bound by the ARM ISA specification. Normal architecture licenses don't allow the licensee to add custom instructions.
But even if you have full license, it doesn't necessarily mean you have time and resources to do full custom implementations all the time.
>Architectural licensees get a set of specs and a testing suite that they have to pass, the rest is up to them. If they want to make a processor that is faster, slower, more efficient, smaller, or anything else than the one ARM supplies, this is what they have to do.
https://semiaccurate.com/2013/08/07/a-long-look-at-how-arm-l...
However, neither Apple or Intrinsity had any IP rights over Samsung's core, as far as I ever saw.
Didn't know the part about Intrinsity, thanks for that :)
An architecture license is not usually a free pass to modify the architecture and fragment ARM's ecosystem at will.
It's easy to forget that Apple is/was one of the original computer companies, and had a founding hand in two ISA/architectures: PowerPC and ARM. They also used the Motorola 68000 and x86 (32+64), but in those cases they were relying on the manufacturer for the chips and architectural direction.
I think those patents would have expired by now.
This would be surprising to me.
Suppose Amazon acquires an architecture license, designs a custom core for Graviton3, and deploys it in AWS. Could they add custom instructions at will?
I would think ARM would be very unhappy when there are applications that run on amazon's arm64 CPU that don't run on ARM's arm64 CPUs, and I would imagine that that would be reflected in the license terms.
By the time the 386 was released, this method was so common that Intel had to preserve these quirks to maintain compatibility with existing software. Apple may find themselves in a similar situation if they're not careful.
I can remember when Microsoft decided to port some of the underpinnings of early versions of Windows to the Mac so they could run a unified code base for Word 6 on the Mac and on Windows, but they had to abandon that effort almost immediately as it was bloated, too slow, and didn't use the native UI.
>Mac Word 6.0 was a crappy product. And, we spent some time trying to figure out how not to do that again. In the process, we learned a few things, not the least of which was the meaning of the term “Mac-like.”
https://web.archive.org/web/20040514091238/http://blogs.msdn...
That's the only time I can remember them doing something unique that other vendors weren't also doing.
I don’t remember any Excel-specific code to keep Excel running, but there was a special error code from the “launch program” system call saying “program has special memory requirements”, but AFAIK, it actually meant “program is Excel version such and such” (changing the programs’s creator code made “launch program” succeed, but the program almost immediately crashed.
With the level of vertical integration they have now it is unlikely they will ever start to care.
Details are described at [1], tl;dr:
- store register values in RAM
- set magic value in CMOS
- cause triple fault to reset CPU
- BIOS sees magic value and instead of resetting rest of the system, it resumes from your stored state in real mode
So it's more like a BIOS quirk. A20 gate backward compatibility on the other hand ...
[1] - http://rcollins.org/articles/pmbasics/tspec_a1_doc.html
Start at 53:14: https://youtu.be/b13xnFp_Los
There is positively nothing "secret" about this.
RISC-V is fixing it, but it's great that Apple is also going into this direction.
Though it is unlikely that vector instructions will help very much for typical JS or compilation workloads.
To clarify, though, the changes introduced in newer SIMD extensions are rarely just a width increase. AVX-512, for example, increased the width from 256-bit to 512-bit, but also brought some substantial programmability improvements.
Linus was quite angry about it for a good reason.
https://community.cloudflare.com/t/archive-is-error-1001/182...
# These
# instructions have been reversed from Accelerate (vImage, libBLAS, libBNNS,
# libvDSP and libLAPACK all use them), and by experimenting with their
# behaviour on the M1.
The Accelerate framework is the only thing that has the Matrix co-processor support.
> You (and other commenters) are aware of NEON, but apparently not of AMX. AMX may not work for the sorts of JSON parsing weirdness for which you use AVX256 (that’ll have to wait for SVE/2, probably next year) but it does solve the problem of “I want to execute dense linear algebra fast”.
> You might want to run some comparisons of that for your M1 vs Intel MacBooks… The API’s to look at are in Accelerate()
The Core ML [4] docs have a nice diagram that demonstrates the relationship with the lower level Accelerate [5] and Accelerate BNNS [6]:
> Core ML is the foundation for domain-specific frameworks and functionality. Core ML supports Vision for analyzing images, Natural Language for processing text, Speech for converting audio to text, and Sound Analysis for identifying sounds in audio. Core ML itself builds on top of low-level primitives like Accelerate and BNNS, as well as Metal Performance Shaders.
Apple's TensorFlow implementation for the M1 is analogous to Core ML. It probably invokes the appropriate Accelerate BNNS and/or Metal library calls. Maynard Handley was exactly right, SIMD/ML/DL code that targets Apple silicon should link against Apple's Accelerate or Core ML libraries. These libraries in turn invoke the appropriate AMX, Neural Engine, and GPU instructions for each target platform.
[1] https://news.ycombinator.com/item?id=25408853
[2] https://news.ycombinator.com/item?id=25773109
[3] https://lemire.me/blog/2020/12/11/arm-macbook-vs-intel-macbo...
[4] https://developer.apple.com/documentation/coreml
[5] https://developer.apple.com/documentation/accelerate
[6] https://developer.apple.com/documentation/accelerate/bnns
That's what I wanted to know. Maybe I missed it.
Of course Apple doesn't care and has been shipping custom instructions for longer than that…
Plus you seem to imply that Apple has done this without Arm's consent which seems unlikely and even if so I would be astonished if that were in the public domain.
On the contrary, ARM has long worked closely with one or more licensees when developing new features.
>ARM picks 2-3 lead licensees for each market segment and works closely with them. This narrows down the pool of initial partners to a manageable number and also ensures that anyone who is picked is up to the task of successfully getting a new core to market correctly. The partners get about a year lead time to market, critical for some, far less important for others.
https://semiaccurate.com/2013/08/07/a-long-look-at-how-arm-l...
It's a bit like the way the GPL has requirements that apply if you distribute the software (you have to release source code), but if you use your modified code internally the requirement doesn't apply. In this case "distributing" is equivalent to selling the CPUs as components, in contrast to using them in your own products (even ones you sell to customers) which counts as internal use and so the requirement doesn't apply.
But you implied that Apple flouted what was permitted by Arm.
> Just off the top of my head
That's not a reference.
As others have commented its much, much more likely that Apple and Arm have a very close relationship and that Apple are just at the leading edge.
I have no idea what issues you had with my examples but if you describe them perhaps I can make them acceptable to you?
On the first point you said 'Apple doesn't care' which I read as Apple doesn't care about the licensing rules - if you meant something else then fine but I think most people would read it that way.
Yes.
Depends on whether these instructions can be executed with unprivileged user-land code. Once the genie is out of the bottle, backward compatibility means imitating the old instruction set. Since a recompile is often cheap enough (and esp. if coprocessor calls can be remapped at the bitcode compilation time), this should be a non-issue.
This is just like how (some) GPU vendors don't document their hardware's data structures and instruction set. You're living dangerously if you choose to rely on their exact behaviour. The vendors really only expect you to use their API (Vulkan/Metal/Direct3D/etc) because that abstraction layer allows them to radically and incompatibly change their hardware on a regular basis to improve performance or functionality.
GPUs are external processors, that require DMA transport, and can perform useful jobs in parallel with the CPU, and are produced by multiple vendors. This almost certainly requires a HAL, and there's little impact of making dynamic libraries calls compared to the benefit received.
Compare this to a co-processor, that basically can be called directly from CPU instructions, without DMA and a remote instruction queue. Something like this gains great benefits from a compiler intrinsic, as a function call, particularly as a dynamic library is quite expensive. This compiler intrinsic can just map to bitcode, which can then map to whatever at compile-time on the app store deployment side.
Why would people not want this option? Why should this only be available to Apple's internal products? Why should Photoshop/Lightroom/Blender/physics simulation/video encoding/AAA Games etc. run slower?
There's no excuse for Apple not providing a way to use these instructions, either via bitcode, or providing a mechanism to detect relevant CPU metadata. At a minimum, I expect hackers to take advantage of this to speed up their own applications, as well as open source forks.
Apple has always been merciless on this. If you use stuff they don't support and your software later breaks, that is your problem.
This is different from say Microsoft which has added support for quirks and bugs in software in their OS updates to keep old software running.
Apple doesn't do that, and I think it is the right decision. It makes developers used to following standards. But it is a tough call as it really pisses off people when their software breaks due to OS updates.
But why not support it at the bitcode level? Matrix operations have broad applicability, so they are just slowing down the software for customers using third party products that could use those instructions beyond what the library offers.
Since the instructions are completely unofficial and only available through what I assume is a dynamically linked system framework, there is 0 chance Apple will care.
For prior art, when Apple silently decided to change the ABI of the syscall underlying gettimeofday(2) in Sierra, projects who made raw syscalls despite that never being officially supported (cough cough golang) had to fix their stuff.
So?
> And if wager that Apple will be using those instructions under the hood of their own products.
Which they can, whether through the public library they offer or through one of their private frameworks.
This almost never comes up, because they typically don’t need to break things, but I can absolutely envision them performing microcode updates that make the CPU incompatible with previous releases.
They owe nothing to anyone regarding these instructions, and they will act accordingly. That’s what unsupported means here: not the PC or Android meanings where everyone has to kludge it in, but the Apple meaning where they will release without warning an OS point release version 15.2.7 that shatters an entire ecosystem built on unsupported APIs.
These instructions are just unsupported APIs in Silicon. Use them at your peril.
Somehow I doubt it. But hey. Maybe M1 will change their minds. Maybe something more powerful will. A 27” iMac that can play FPS games at 60fps/5K/HDR with virtual Atmos surround out of the box? Near-effortless console porting to a future Silicon AppleTV, with gameplay compatibility across both devices? They’ll realize what they’re missing out on eventually :) But they’re not going to code M1 assembler until they do, and it’s easier for them to use Apple’s frameworks than play games with the CPU.
(Yes, ffmpeg will probably do this. Good for them! And I’m sure Adobe will give it serious consideration - and then reject it, since it would get their apps rejected from the app stores and harm their relationship with Apple.)
> it’s easier for them to use Apple’s frameworks than play games with the CPU
The things I'm talking about are broader than what their accelerated libraries support. Matrix operations are ubiquitous across heavy-weight computation, including stress analysis, physics simulation, machine learning, video and sound processing, compression, ray-tracing.. why would Apple limit their market? There's a whole host of 3rd party applications that would be augmented with this co-processor. And it should be trivial to add a compiler intrinsic that compiles to some bitcode primitive that can ensure it works across all devices and OS's... or just provide a mechanism to detect the extension.
> And I’m sure Adobe will give it serious consideration - and then reject it, since it would get their apps rejected from the app stores and harm their relationship with Apple.
This is bad for Apple. And bad for the consumer. Final Cut Pro isn't the only tool that video professionals need to use, and you can bet FCP will be using the coprocessor directly under the hood.
Apple offers frameworks that accelerate the relevant operations, and has over time shown an ongoing willingness to limit their market, even when that’s seen as unpalatable or unprofitable or unhelpful to themselves or their users. I assume that the frameworks offering acceleration will expand their uses somewhat over time. I expect those frameworks to be the only compelling solution for having portable matrix acceleration across the complete spectrum of Apple Silicon hardware revisions over time, especially if Apple revises their undocumented operations and updates their frameworks annually (which they can easily do, since each year’s hardware release already demands the latest macOS to boot at all).
In general, the oldest macOS you can boot on any given Mac is "the version the Mac shipped from the factory with". So, for example, the 2017 iMac Pro can't boot a 2016 or earlier macOS.
This is still unsupported, mind you — if your 2017 iMac Pro is firmware updated to 2020 and you're booting a macOS from 2018 on it, you may encounter random bugs and/or crashes and/or mysterious issues due to assumptions made by the older macOS that are no longer entirely valid on 2020 firmware — but in general, it should work as long as the macOS is newer than the hardware.
(You probably won't be able to boot an M1 macOS external drive on an Intel Mac, and vice versa, for other reasons unrelated to this. In that scenario, Migration Assistant is probably your only choice anyways.)
I did this several times because I couldn't believe what I was seeing. I can't say for sure where the source of the problem lies—maybe something is weird with my installation media—but my theory is that a firmware and/or microcode update subtly broke something in Apple's old installers, possibly when they added APFS support. I looked into downgrading the firmware, but it seems to be impossible from what I can tell[1].
Mind you, outside of this one quirk the OS and installer run fine.
---
1: https://forums.macrumors.com/threads/guide-how-to-get-back-o...
With intel macs you can boot and install older macOS's up to a certain point. Like the previous poster says, it's about what firmware versions you have.
For example, if your mac came with Sierra, and through the regular usage and update, you patched it up to Mojave, and installed (or was forced to install by the way, there is no way to reject certain types of updates) firmware updates, you find that all of a sudden you can't re-install Sierra anymore. You'd need to install whatever version your firmware now supports.
This behaviour is undocumented and it's case by case - for each mac hardware there is a firmware for it - there are lots of people online who have figured out patch levels are for each revision of hardware, etc - but it's more trial and error to see what older version of macOS you can downgrade to.
I have never seen any record of this happening. I'm really quite sure you've always been able to install the earliest OS that your model of Mac shipped with.
The odd situation I mentioned above is the closest I've ever seen to a firmware update causing problems (if that's even the cause), and even then the OS runs fine.
There actually is a table of firmware versions, etc - once you upgrade your old Mac hardware past a certain firmware, you will no longer be able to run certain versions of macOS "before" the firware was supported.
This isn't just between "major" versions, also between "minor" versions of macOS
For example, between the current version of macOS Sierra, which is on 10.12.6 (6 being the minor version), if you try to install for example 10.12.x where x is below 6, etc - you might get the infamous ? or the circle with a slash through it on bootup.