This subset AVX10/256, is the reason for this new specification. It is the Intel response to AMD Zen 4.
When their competitor supports AVX-512 on all products, Intel had to do something to remain competitive. Because they believe that supporting the full AVX-512 on their E-cores is too expensive, they have created a subset of AVX-512, including only the instructions with an operand size up to 256 bits.
Now, a different method has been defined for discovering through CPUID which AVX-512 a.k.a. AVX10 features are implemented, so only now it has become possible to implement an up to 256-bit subset.
Moreover when 128-bit vector and 256-bit vector instructions have been added to AVX-512, 2 bits from the EVEX prefix that were previously used for rounding control have been reused to encode the length of the vector operands.
Because of this, only the 512-bit vector instructions and the scalar instructions can specify the rounding control. So if the 512-bit vector instructions are deleted, there is no longer any way to specify the rounding control for vector instructions.
To solve this problem, in the first CPUs that will implement the 256-bit subset of AVX-512, new encodings will be used for the 256-bit instructions with rounding control.
Also the XSAVE and XRSTOR instructions had to be modified to save and restore correctly the new vector registers and mask registers.
So implementing the 256-bit subset of a AVX-512, a.k.a. AVX10/256, is not so straightforward as a microcode update, it requires changes in the instruction decoders and in the structure of the CPUID registers and other smaller changes.
I don’t entirely agree. See the second graphic here: https://www.phoronix.com/news/Intel-AVX10
Mentioned above the graphic:
> Part of making AVX10 suitable for both P and E cores is that the converged version has a maximum vector length of 256-bits and found with the E cores while P cores will have optional 512-bit vector use.
512-bit support will be optional, so maybe every P-core won’t have it… maybe it’ll be restricted to higher end processors? But it sounds like some will have it, or it wouldn’t be an option at all.
"A “converged” version of Intel AVX10 with maximum vector lengths of 256 bits and 32-bit opmask registers will be supported across all Intel processors, while 512-bit vector registers and 64-bit opmasks will continue to be supported on some P-core processors."
So all future Intel CPUs starting in 2025 will support a 256-bit subset of AVX-512, where AVX-512 is rebranded as AVX10.
Only some P-core processors will support the full 512-bit AVX-512 a.k.a. AVX10, which is to be understood that only those server CPUs that contain only P-cores, i.e. the successors of Granite Rapids and Granite Rapids D, will support 512-bit registers and instructions (and 64-bit mask registers instead of 32-bit mask registers).
P-cores in hybrid CPUs will be 256-bit (like E-cores).
P-cores in (server) CPUs that contain only P-cores will be 512-bit.