Intel AVX10 Drops Optional 512-Bit: No AVX10 256-Bit Only E-Cores in the Future
phoronix.com
phoronix.com
But the whole approach with fixed-length instructions seems terrible to me. It takes Intel a decade to add another batch of instructions for another width, and the existing applications don't benefit from the new instructions, even if they already process wide batches of data.
The CUDA approach is so much more appealing: here's my data, and you can process it in how many small or large units as you want.
There exists also for Intel AVX/AVX-512 a compiler that implements the CUDA approach (Intel Implicit SPMD Program Compiler). Such compilers could be written for translating any programming language into AVX-512, while using the same concurrency model as CUDA.
Moreover, as a software model the "CUDA approach" is essentially the same as the OpenMP approach, except that the NVIDIA CUDA compilers are knowledgeable about the structure of the NVIDIA GPUs, so they are able to map automatically the concurrent threads specified by the programmer into GPU hardware cores, threads and SIMD lanes.
We routinely write code that works on 128-512 bit vectors. Some use cases are harder than others, e.g. transposing.
If that were true, it would be a good sign for his competence.
According to the steam hardware survey for Feb 2025, AVX512 support is already up to around ~16%. That may not sound like a lot, but considering that it's basically all coming from Zen4&5 based CPUs..
Once Intel's next design also supports it, it's probably worth targeting for consumers in relevant markets.
> Intel has dropped the 256-bit-only approach and going for 512-bit everywhere. Thus it would seem to indicate that Intel E cores of the future will properly support AVX 512-bit operation!
> It looks like AMD's widespread support for AVX-512 since Zen 4 and the rather confusing AVX10 implementations previously pursued by Intel are now over. With updated GCC compiler patches posted today, that 256-bit mess proposed for future AVX10 versions is being removed.
* https://en.wikipedia.org/wiki/X86-64#Microarchitecture_level...
Or that it'll support _both_ 256 and 512 instructions going forwards (and stop doing the nonsense where some cores support 512 and others don't?)
So none of the current instructions will be removed.
The previous plan of Intel was that in consumer CPUs the 512-bit instructions shall be removed, keeping only 256-bit instructions, 128-bit instructions and scalar instructions (FP32 & FP64).
Nevertheless, the most ancient versions of AVX-512 had only 512-bit instructions and scalar instructions.
The 256-bit instructions and 128-bit instructions have been added in Skylake Server, as a workaround for the bad power management of Intel at that time, which forced huge drops in clock frequency for long times when using wide instructions.
On modern CPUs there is no need to use 256-bit or 128-bit instructions. You gain nothing with them. AVX10 instructions have masks, so you can process any arbitrary length with a 512-bit instruction, in the case of loop prologues or epilogues.
The use of 512-bit instructions simplifies many optimized programs, because one instruction processes one cache line.
The default auto-vectorization tuning for current Intel server CPUs using 256-bit registers, which is perhaps another counterexample.
For atomic, I'm curious how you make use of that?
And then opens it up.
x86-64v4 was defined around 2017 and requires AVX-512. Few Intel CPUs comply with it even today.