Advanced Performance Extensions (APX)
intel.com
intel.com
* A new REX-like prefix that extends the number of addressable GPRs to 32 from 16. This only supports instructions that have one-byte opcodes or the 0f prefix, so recent GPR instructions like ADCX or BLSR aren't supported with this format, except.
* The EVEX prefix (used for AVX-512) is also extended to be usable for GPR instructions instead of just vector instructions. This allows three-address instructions to be defined.
* The EVEX prefix for GPR also has a dedicated bit for "do you want to set flags as a result of this instruction."
* New instructions that push/pop 2 GPRs at once
* New instructions that let you conditionally set flags (basically you can do OR/AND in the hardware flags, this sounds useful for compilers).
* New instructions for predicated loads.
* New 64-bit absolute jump instruction
* Also, implementation of the predicated stuff in AVX-512, but for 256-bit vectors. With this note:
> A “converged” version of Intel AVX10 with maximum vector lengths of 256 bits and 32-bit opmask registers will be supported across all Intel processors, while 512-bit vector registers and 64-bit opmasks will continue to be supported on some P-core processors.
I realized I’ve completely lost track of Intel’s architecture extensions reading this:
“They do not change the size and layout of the XSAVE area as they take up the space left behind by the deprecated Intel® MPX registers.”
Apparently MPX was Memory Protection Extensions and the design was so flawed, it was removed entirely soon after introduction.
SSE, AMD64, AVX, and AVX-512 say hi.
There are 8 64-bit MMX registers. To avoid having to add new registers, they were made to overlap with the FPU stack register. This means that the MMX instructions and the FPU instructions cannot be used simultaneously. MMX registers are addressed directly, and do not need to be accessed by pushing and popping in the same way as the FPU registers. [1]
[1] https://en.wikipedia.org/wiki/IAPX [2] I know you're thinking about the iAPX432 and forgot the rest
I have always been curious as to why the number of GPRs were limited for so long on X86 given that the instruction set is already variable length, and the CPU have typically a very large number of internal arch-register that could be cheaply addressed.
Having looked at the pain of developing a good register allocate in LLVM, and how critical memory access can me in hot/tight loops i would have loved to have even more register something closer to 64 or 128, and let the cpu manage the spilling internally.
Because it's hard-ish to wrench in the encoding bits. x86 registers are addressed via the ModR/M byte, which gives you 4 addressing modes × 2 register operands (if you have 8 registers). Note that non-register-register addressing modes include a second SIB byte which means you have up to three registers encoded in an instruction. x86-64 extended the number to 16 by adding a prefix that has 4 bits, which encoded a 32-bit/64-bit selector and 1 bit for each of the three possible registers (R, X, B). Note that to keep the prefix to only a 1-byte length, they had to reclaim 16 possible opcodes, whose space is pretty limited.
Encoding large register numbers gets pretty chonky in instruction space (3 register ids of 5 bits each is 15 bits of your instruction for operands, and that's before you start considering possible opcodes). It also increases the size of context switches (especially thread contexts), since you have to spill all of those registers even if they're not filled with useful data.
8 registers is definitely too few; it's not clear to me if the extra working set size afforded by 32 registers is worth it over the fatter instructions.
Even though it's only one extra bit to index 16 registers vs 8, it takes an entire extra byte in the encoding (you have to add the REX prefix with the R, X, and B bits set appropriately for the MSB of the registers you're modifying in the MODR/M byte). And you only get 4 bits to use out of that byte because the leading four are a way to disambiguate it (also iirc, this stuff is terribly documented).
And paying an extra byte per GRP instruction is actually kind of expensive. If you look at your compiler output you'll see a bunch of Exx instructions even in 64 bit long mode when optimizing for instruction decode/size.
It turns out a variable length encoding is really hard to extend without getting weird with it.
No, it has been extremely well documented since 2000 (three years before AMD had the first 64-bit chip ready).
(Shared libraries need to have a pointer to their code and data. This is done by creating a small table of entry code in front of the "real" code that sets a register to point to the library code/data memory before jumping to the real code. This code will look different for each process that has loaded the shared library so it gets its own little memory block. The rest of the code/data will be the same in all processes so it can be shared between them -- unless/until some of the data is written to, of course.)
So code in shared libs really only had 5 registers unless the frame pointer was fomitted. That's why it was necessary to tell the compiler whether to build normal code or shared library code.
The AMD64 not only had more registers, it also has an RIP-relative addressing mode so it doesn't need to steal that register for shared libraries.
I wonder why this long delays are still necessary. In the old days yes as there were so many parties to coordinate but nowadays, in theory, Intel could release hardware, the new ISA and compiler/OS patches and binaries on the same day.
Most new ISA support requires at least minimal support from the OS to use correctly (e.g., something like saving state on context switching, or reporting hardware support bits correctly). Releasing patches on the same day you release hardware means the new features are literally unusable for all of your customers, and it generally takes several months to go from a patch to a usable OS release. You could in theory release the patches on the same day as the new ISA documentation, but in practice, it's likely to take some time because the people writing the patches aren't the people writing the new ISA whitepapers and approving their publication and it takes time to go from "oh, I can talk about this now" to actually doing so.
Taking a little time with these sorts of things isn't bad. This will essentially be a new epic where AMD and Intel may end up being incompatible, at least for a while, there are tooling changes to make, OS support, all sorts of things. If they're serious about dropping 16bit and 32bit from the architecture, doing it during a change like this might make some sense. Do it wrong and they could Osborn themselves, but an Apple like migration strategy would be nice and I'd appreciate it.
After all why spend a shitload of money on developing a feature you'll want to market when there won't be any reason for customers to buy your new hardware?
The main reason to do a coordinated release would be competitive advantage, I guess, plus to ensure that there are actually compilers that support their new instructions available right from the start. Intel already release their own Linux distro optimized for their chips because other vendors were targeting the lowest common denominator
None of the products announced for 2024 (e.g. Sierra Forrest, Granite Rapids, Arrow Lake, Arrow Lake S, Lunar Lake) support this.
Making sure the thing actually works? At least I hope they're still doing that... and it's ironic to see this comment at the same time a rather horrible bug was discovered in AMD's CPUs:
https://news.ycombinator.com/item?id=36848680
If anything, Intel should probably slow down and take more care: https://news.ycombinator.com/item?id=16058920
> "In addition, legacy integer instructions now can also use EVEX to encode a dedicated destination register operand – turning them into three-operand instructions and reducing the need for extra register move instructions."
Overall, APX is providing 10% fewer instructions, 10% fewer loads and more than 20% fewer stores.
Also adding pop2/push2 instructions for moving state faster.
And adding more powerful conditional instructions (loads/stores/compares) and flag-suppression.
"The processor tracks these new instructions internally and fast-forwards register data between matching PUSH2 and POP2 instructions without going through memory."
I wonder if this implies that pushes don't have to commit to memory if they are popped soon enough? It has always bothered me that we have these huge physical register files but force all the spill and restore to go through memory because of silly anachronistic processor semantics. With a more flexible PUSH/POP semantics we could essentially get the register windows for free.
1. <https://www.researchgate.net/publication/224114307_Code_dens...>
2. <https://web.eece.maine.edu/~vweaver/papers/iccd09/ll_documen...>
x86 ISA is growing more RISC-like. Definitely saving on stack spilling is a Good Thing™
I was not familiar with the AVX vector instructions at this level of detail.
That's because it hides inside the BOUND instruction in a way that allows it to be used outside of 64-bit mode code. The BOUND instruction never existed in 64-bit mode, so Intel was free to do whatever they wanted with the 62h opcode, but they thought it would be enable EVEX in 16/32-bit code too.
The BOUND instruction must take a memory operand -- so the MOD bits can never be 11b (which would specify a register operand). The MOD bits are the upper two bits of the modrm byte (the byte right after the 62h opcode). So, if we don't allow the "upper" registers in 32-bit mode and only the "lower" 8 registers AND if we invert bit 3 of their register numbers and put those extra register specifier bits in bit 7/6/5 of the byte after the 62h opcode THEN we can fit EVEX into 32-bit mode, because those bits will always be 111b in 32-bit mode!
So outside of 64-bit mode, if we have a 62h opcode with a modrm byte with MOD!=11b: it's a BOUND instruction. If MOD=11b: it's an EVEX prefix.
- +16 registers (thus 32) and optionally separate destination (looking very RISC like now)
- PUSH2/POP2 with full forwarding
- Much expanded predication, including predicated loads and stores
This is pretty interesting. Especially the latter can make a big difference for highly unpredictable memory intensive code, like compression.
REX2 is D5 + a byte that is a lot like the lower nibble of a REX prefix, only twice as big.
It has two extra bits for the up to three registers that can be named in normal x86 instructions: either two registers or a register operand and a memory operand that can use two registers + a displacement for the memory address.
It also contains the W bit like REX does (64-bit).
It also has a bit called M0 that indicates whether the instruction is in opcode map 0 (primary opcode map) or opcode map 1 (0F opcode map).
That means that a REX2 instruction from opcode map 1 takes the same number of bytes as a REX instruction from opcode map 1. Instructions from opcode map 0 are one byte longer with REX2 than with REX.
Some of the normal ALU instructions also get an EVEX encoding (in map 4). That allows for a different data destination than before (separate from source the operand(s)). It also allows for ALU instructions that don't change the flags, which must be really nice for the out of order/data forwarding circuitry.
This subset AVX10/256, is the reason for this new specification. It is the Intel response to AMD Zen 4.
When their competitor supports AVX-512 on all products, Intel had to do something to remain competitive. Because they believe that supporting the full AVX-512 on their E-cores is too expensive, they have created a subset of AVX-512, including only the instructions with an operand size up to 256 bits.
Now, a different method has been defined for discovering through CPUID which AVX-512 a.k.a. AVX10 features are implemented, so only now it has become possible to implement an up to 256-bit subset.
Moreover when 128-bit vector and 256-bit vector instructions have been added to AVX-512, 2 bits from the EVEX prefix that were previously used for rounding control have been reused to encode the length of the vector operands.
Because of this, only the 512-bit vector instructions and the scalar instructions can specify the rounding control. So if the 512-bit vector instructions are deleted, there is no longer any way to specify the rounding control for vector instructions.
To solve this problem, in the first CPUs that will implement the 256-bit subset of AVX-512, new encodings will be used for the 256-bit instructions with rounding control.
Also the XSAVE and XRSTOR instructions had to be modified to save and restore correctly the new vector registers and mask registers.
So implementing the 256-bit subset of a AVX-512, a.k.a. AVX10/256, is not so straightforward as a microcode update, it requires changes in the instruction decoders and in the structure of the CPUID registers and other smaller changes.
I don’t entirely agree. See the second graphic here: https://www.phoronix.com/news/Intel-AVX10
Mentioned above the graphic:
> Part of making AVX10 suitable for both P and E cores is that the converged version has a maximum vector length of 256-bits and found with the E cores while P cores will have optional 512-bit vector use.
512-bit support will be optional, so maybe every P-core won’t have it… maybe it’ll be restricted to higher end processors? But it sounds like some will have it, or it wouldn’t be an option at all.
"A “converged” version of Intel AVX10 with maximum vector lengths of 256 bits and 32-bit opmask registers will be supported across all Intel processors, while 512-bit vector registers and 64-bit opmasks will continue to be supported on some P-core processors."
So all future Intel CPUs starting in 2025 will support a 256-bit subset of AVX-512, where AVX-512 is rebranded as AVX10.
Only some P-core processors will support the full 512-bit AVX-512 a.k.a. AVX10, which is to be understood that only those server CPUs that contain only P-cores, i.e. the successors of Granite Rapids and Granite Rapids D, will support 512-bit registers and instructions (and 64-bit mask registers instead of 32-bit mask registers).
P-cores in hybrid CPUs will be 256-bit (like E-cores).
P-cores in (server) CPUs that contain only P-cores will be 512-bit.
The question is when will AMD adopt it. Zen 5 is done and Zen 6 may be too late for these changes. Zen 6 is already looking at 2026. If they waited til Zen 7 it will be at least 2028.
Intel is still 35% behind Apple in terms of Pref / Clock on Geekbench.
In other words, their initial sloppiness was a Feature™: our initial mess was so bad that changes like these don't make it any worse!
AFAIK there was AVX512 in higher-end desktop SKUs, but not the latest E-core designs in Intel's client chips. The main problems are that:
- Operations at 512-bit register widths have significant power draw. Intel chips that support AVX512 have to downclock themselves on AVX512 workloads until their voltage regulators have boosted up to a higher voltage.
- The AVX512 register file is too big to physically fit in the E-core[0] footprint.
Incidentally I do remember Linus Torvalds specifically complaining that AVX512 was being used to implement memcpy in gcc, because it meant running certain programs would lower system performance. So these new architectures tend to be used a lot sooner than the time it takes for it to be safe to make them your minimum compile target.
[0] The BIOS on my Framework laptop refers to these as "Atom cores" - no clue if the current E-core design is derived from Atom or if this is a miscommunication or nickname AMI picked.
What has significant power draw is the use of double 512-bit floating-point multipliers, as implemented in the Intel server CPUs (though one core with such multipliers draws significantly less power than two cores having the same throughput).
AMD uses only double 256-bit floating-point multipliers and in general it uses exactly the same execution units for both 256-bit and 512-bit operations, so the AVX-512 operations do not increase the power draw even when they use 512-bit registers.
Also the AVX512 register file is not too big to physically fit in the E-core. Even if the AVX512 register file is 4 times greater than the AVX register file, the E-cores have much a much larger register file used to rename the architecturally visible registers.
Despite these facts, Intel still believes that implementing the full AVX-512 ISA in the E-cores is too expensive, so they have created this new specification of AVX10/256, which is just a subset of AVX-512 including the instructions with an operand size up to 256 bits, and which will be implemented in all future E-cores after some date, perhaps starting in 2025.
CachyOS and Clear Linux already do this.
Base Arch Linux and openSUSE are working on it. Maybe Fedora too, but I can't remeber
And it wouldn't be totally insane for Windows to do this either.
I'm honestly a little amazed that more people don't either use Gentoo or adopt its model, given the smorgasbord of mutually incompatible instruction set extensions. It seems very strange that people will buy a CPU that has 32 64-byte ZMM registers with 3 operand instructions, and then use that CPU to run code that operates on 8 16-byte XMM registers with 2 operand instructions.
This also means "riskier" methods (like LTO) have to be omitted by default.
Gentoo is great for libre software, security, manual patches, embedded computing and such. But for pure desktop performance, the Clear Linux way is best: aggressive compilation flags/libraries, tested by the package maintainers, shipped in 3-4 tiers. And as the Clear Linux devs said, most of the native instructions dont even matter, as the compilers can't use them.
https://gcc.gnu.org/onlinedocs/gcc-13.1.0/gcc/Function-Multi...
And there are different schemes for multiple architectures in the same program, like hwcaps.
https://developers.redhat.com/blog/2021/01/05/building-red-h...
There is discussion of a fifth level. Someone in the Intel Clear Linux IRC said a fifth level wasn't "worth it" for Sapphire Rapids because most of the new AVX512 extensions were not autovectorized by compilers, but that a new level would be needed in the future. Perhaps they were thinking of APX, but couldn't disclose it.
We can throw away any hope of v4 being a standard baseline.
Work out what it would cost to compile - say - a terabyte of C code at typical cloud spot prices.
A large VM with 128 cores can compile the 100 MB Linux kernel source tree in about 30 seconds. So… 200 MB/minute or 12 GB/hour. This would take 80 hours for a terabyte.
A 120 core AMD server is about 50c per hour on Azure (Linux spot pricing).
So… about $40 to compile an entire distro. Not exactly breaking the bank.
the hardware is the same design for cost saving purposes, but different features are unlocked for $$$ by a software license key.
You want AVX-512? pay up and unlock feature in your CPU and you can now use the feature. This could also enable pay-as-you-go license scheme for CPUs, creating recurring revenue for Intel
from the hardware perspective - the same silicon, but different features sold separately
Arguably it will be only JITd languages that benefit from this for quite a while. These sorts of fundamental changes are basically a new ISA and the infrastructure isn't really geared up to make doing that easy. Everyone would have to provide two versions of every app and shared library to get the most benefit, maybe even you get combinatorial complexity if people want to upgrade the inter-library calling conventions too. For native AOT compiled code it's going to just be a mess.
Or Xerox PARC, with their microcoded CPUs loading the desired interpreter on boot.
I guess, it is an idea that keeps being revalidated.
Is the goal here to increase the decode bandwidth of Intel CPUs?
Is the goal to reduce demands on load-store units by increasing the number of registers?
Are they hoping to make it easier to port or JIT armv8 asm to Intel CPUs?
Nevertheless, there are a few instructions inspired by Armv8, mainly PUSH2 and POP2, which correspond to the load register pair and store register pair of Aarch64.
On ARM that instruction is a pain in the ass when reading disassemblies, because the on-fail condition bits are just specified as a number from 0 to 15; the disassembler doesn’t bother to label which bits are specified, let alone what conditions they correspond to. Unfortunately it seems like Intel is doing the same thing in their assembly syntax, at least if I’m reading the document correctly.