I found out about this issue on Github and I learned that the M1 chip does not support AVX (which is indeed documented) and that Homebrew developers assumed that all CPUs supported by Big Sur were compatible with AVX instruction set (which is true except for the M1 under Rosetta 2).
I found it interesting to see how they rapidly corrected the situation by merging a partial revert. I also found their CI infrastructure interesting to watch while it was live-testing the commit.
cc -dM -E -xc /dev/null | less
Anything with a preprocessor macro is supported on all OS configurations that you are targeting. You can see, for example, that __SSE4_1__ is defined, so SSE 4.1 is ok to use without checking it at runtime with e.g. sysctlbyname("hw.optional.sse4_1", &enabled, ...)I have always thought that it is worth a bit of paranoia to try and do the check for optional features the “correct” way, whatever method the OS advertises (like sysctl for macOS) rather than doing something like cpuid. After all, every once in a while it’s possible to run into a configuration where the CPU does support some particular feature but the OS does not, and if you rely on cpuid you could be up shit creek without a paddle.
Within a few days problems became evident and Homebrew turned the compiler flag back off. Now it's just a question of the best way of remediating the already-compiled packages in the simplest way possible.
Really the fact that this is now a front-page-of-HN story is going to cause more confusion to people. It was just a short-lived bug.
So I personally find it interesting because I would not have expected Rosetta 2 to be missing AVX instructions if all the supported Big Sur hardware possessed them, and this is an interesting failure mode to watch and see how the Brew developers handle.
AVX can’t be emulated by Rosetta due to Intel patents.
On the other hand, disabling AVX for all Intel machines would make those programs significantly slower, so it's clear why there is reluctance to do that...
side bar: is there documentation for the instruction set or abi for that hardware?
I'll make a wild guess that getting data to the neural engine is still probably not quick because I assume it's some kind of statically scheduled type affair (exposed pipeline?). We literally seem to know almost nothing about it sadly.
https://jobs.apple.com/en-gb/details/200205070/neural-engine...
Even the job listing gives away next to nothing, other than poor English "Knowledge in compiler is a plus" ;)
No. Apple stans, food for thought.
I would hope there will be something soon although it's Apple so not much.
Edit: Still basically no, but https://github.com/geohot/tinygrad/tree/master/ane has got the instruction format apparently.
You can attack an instruction set blind (https://recon.cx/2012/schedule/events/236.en.html).
More edit: Bingo, patent: https://patents.google.com/patent/US20190340491A1/en?oq=2019...
Often you are happy to get a 1.25x speed up with AVX. Sometimes it actually goes slower.
If you were to emulate that code with a 1.25x speedup with AVX on the M1, you would end up with all the disadvantages of going to 8-wide, but with none of the speedup.
That 1.25x speedup is halved and the emulated AVX code actually runs at about 0.625x the speed of the emulated SSE code path.
The base requirements for x86-64 mandate SSE2 IIRC. Those patents expired this year, so Apple was now able to release an x86-64 “emulator” without negotiating patents.
(Do you have links or identifiers for the specific patents, by any chance?)
As for patents, I don’t, but someone else here linked in the one for FMA[0]. It’s a bit more complicated than just an API, but it seems broad enough that anything implementing that API would be covered. But IANAL.
Like it or not they are patented https://patents.google.com/patent/US7499962 (FMA, for example)
Yeah, and AVX is also relatively new, intro'd in 2008 and first shipped in a chip in 2011. AVX2 wasn't until 2013. So even with R&D and patents happening years beforehand it'll still be a good long while before they expire (that FMA example being a case in point, not until end of 2026).
Granted in Apple's specific case that's actually not a bad thing. Precisely because AVX is so new, many Macs supported up until the last version or two of macOS didn't have it. So AVX isn't at all a widely expected dependency for the kind of older software that may never get an ARM port and in turn most needs Rosetta 2.
Discussion of Intel patents from a few years ago
"Emulation is not a new technology, and Transmeta was notably the last company to claim to have produced a compatible x86 processor using emulation (“code morphing”) techniques. Intel enforced patents relating to SIMD instruction set enhancements against Transmeta’s x86 implementation even though it used emulation"
[edit: this seemed like quite an uncontroversial thing to say, please understand it was not intended to hurt anyone's feelings. I was a mac user for 12 years and simply don't click those links any more. Please don't downvote comments just because you're a Mac user]
The actual issue is that Rosetta identifies itself as a CPU without AVX so it’s clearly a homebrew bug. And in fact if you tell homebrew not to download a prebuilt binary it works fine.