And so it is in life
Realistically though, how likely would a GCC/clang be to emit these instructions when I'm working on some lookup tables, assuming I permit it to use them (e.g. via `-march=native` on a machine that supports the extension)? My gut feeling would be that unless I specifically make sure to structure my code to be as close to the semantics of the instructions as possible, these instructions would never ever be emitted. Or has the world of compiler optimizers advanced enough that rewriting that is commonplace now?
The "strange" instructions are actually not that niche, it's just that usage tends to be "indirect" and therefore people don't notice.
[^0] E.g. https://xoranth.net/memcmp-avx2
Also super great for emulation, and anyone else who does a lot of bulk bit-twiddling.
The whole discourse has become super weird (up to and including Linux himself ;) because of Intel 10nm delays. With the only AVX-512 products being 14nm-based intel server chips for 5 years, and then only coming to laptop for another couple years, and then only a single terrible generation of desktop parts that nobody bought, and with AMD launching super competitive (usually leading) products in those segments, obviously there wasn't a whole lot of real consistent adoption in software. And what adoption there was, was complicated by the fact that the largest adopter (server market) had to drop clocks massively and even pause processing to allow voltage to swing up enough, because they were 14nm products on a feature that was really aimed at 10nm and beyond. And then Intel yanked it out of all the desktop and laptop chips and seems poised to just ignore it for another 5 years.
Everyone just decided that because it wasn't getting adopted that it was inherently useless, up to and including Linus himself. But it wasn't getting adopted because it was a complete mess on the Intel side and AMD didn't even support it, so why bother?
The AVX-512 story is inextricably bound up in the 10nm delays and the organizational problems that have plagued Intel ever since. It's such a great thing that AMD didn't buy into the naysaying.
[1] https://en.wikipedia.org/wiki/Michael_Abrash [2] https://www.anandtech.com/show/2580/9
Most Intel ISA extensions come from either customers asking for specific instructions, or from Intel engineers (from the hardware side) proposing reasonable extensions to what already exists.
LRBni, which eventually morphed into AVX-512, was developed by a team mostly consisting of programmers without long ties to Intel hw side, as a greenfield project to make an entirely new vector ISA that should be good from the standpoint of a programmer. I strongly feel that they have succeeded, and AVX-512 is transformative when compared to all previous Intel vector extensions.
The downside is that as they had much less input and restraint from the hw side, it's kind of expensive to implement, especially in small cores. Which directly led to its current market position.
Sony Japan's documentation for how to use a mouse & keyboard on the PS2 was literally just the URL "https://www.usb.org/document-library/usb-20-specification". Eventually, they provided a binary-only keyboard library that everyone complained was buggy, but actually just had documentation that was so brief it was easily misunderstood. After black-box testing it for an hour it was clear it worked fine, just not how anyone would expect it to.
Many years ago I made a tiny stir online by writing a stream-of-consciousness report of the experience of dealing with stuff like this for a decade. https://venturebeat.com/games/what-is-making-games-like-for-...
Different constraints and challenges on both sides of the aisle give rise to compromises which end up with lowered performance or lowered ease of use. This is one area where great authority over the entire stack lends you lots of leeways, e.g. Apple designing Metal API and the HW for it.
I have a data structure library (in Rust) where I would love to have these. The problem is that AVX-512 just isn't common enough to rely on it yet, and I don't even have it on my workstation CPU (Radeon 6850, from just last year).
But in particular whether they had something in mind, I suspect Intel was thinking about video codecs and containers for a lot of these. If you read through the specs for them, you will find all sorts of places which call for things like this.
But yes, whether compiler developers can make good use of these. Questionable. They are really for specialized optimization workflows.
> They must have some programs in mind that could be run faster with them
Yeah, all new instructions are built with some workload in mind. This may or may not be specified in the architecture manual or you have to reverse-engineer it from the press releases.