I guess we will have to wait for at least one more generation.
[1] - According to Intel® Architecture Instruction Set Extensions Programming Reference: https://cdrdv2-public.intel.com/826290/architecture-instruct...
I guess we will have to wait for at least one more generation.
[1] - According to Intel® Architecture Instruction Set Extensions Programming Reference: https://cdrdv2-public.intel.com/826290/architecture-instruct...
Panther Lake will introduce FRED (Flexible Return and Event Delivery) a new manner of handling interrupts, exceptions and system calls.
FRED will bring tremendous changes to the operating system kernels, but it will have little influence on user programs, except that the computer will spend less time running OS kernel code than now.
For now it is expected that Intel will introduce AVX10 in its consumer CPUs only in Nova Lake, the Intel 2026/2027 CPU.
Meanwhile, AMD Zen 4 and Zen 5 are already happily supporting AVX10, except for implementing the CPUID AVX10 flags. AVX10.1 differs from AVX-512 only by adding a simpler method for identifying which instructions are supported. AVX10.2 will add only some instructions that are not needed on the CPUs that support the 512-bit AVX-512 instructions, like Zen 4 and Zen 5. AVX10.3 has not been defined yet and it is far in the future.
It's weird to see an Intel so... Broke? That they are seemingly forced to recycle old architectures endlessly
You can't emulate that via just two regular 256-bit uops, you need four (maybe more for blending the results together). And if you don't have the two-register table 256-bit variant (e.g. Tiger Lake doesn't, though for 512-bit of course; it splits it into three uops), that'd end up at a rather massive 12 uops.
The smaller size of the Intel E-cores is not only due to their different microarchitecture, but also because only their L1 cache memories are non-shared, while their L2 cache memories are shared within groups of 4 E-cores.
The shared L2 cache may not matter much for many general-purpose programs, but for other multi-threaded programs, which depend on having a great total throughput for the transfers with the L2 cache, the performance of each group of 4 E-cores becomes similar to that of a single core, instead of being 4 times greater.
The AMD compact cores have the same non-shared cache memories as the big cores. Only the shared L3 cache blocks that service a group of compact cores are smaller than for the same number of big cores.
The register file size makes sense, I didn't think they were that much of the die on those processors but I guess they had to be pretty aggressive to meet power goals?
https://i.imgur.com/WdMPX8S.jpeg
According to this, Zen4s FP register file is almost as big as its FP execution units. It's a pretty sizable chunk of silicon.
But looks more like they're giving up on people writing code for wide vectors, instead settling on trying to make the existing code faster.
Only the availability of AVX10/256 in Intel's consumer CPUs and in its server CPUs with E-cores is in the proposal phase (mainly because Intel has yet to design and launch, as the successor of Skymont that is being launched now, an E-core supporting AVX10/256; this is expected only in H2 2026).