how is AVX-512 innovative when it's essentially SSE with 4x bigger registers? am i missing something?
how is AVX-512 innovative when it's essentially SSE with 4x bigger registers? am i missing something?
It is much more: Just to give examples:
- The introduction of the opmask registers to mask many AVX-512 instructions
- AVX-512 Conflict Detection Instructions (CD) enables lots of loops to vectorize that could not vectorized before.
Also the fact that now 32 SIMD registered can be addressed does not seem very innovative from the outside, but implies a deep change: Before only 8 (32 bit mode; bite 5-3 of MOD-REG-R/M field) or 16 (64 bit mode; additionally use REX.R field) SIMD registers could be addresses for deep instructional encoding reasons. This also holds when using a VEX prefix. So being able to use 32 SIMD registers requires a completely new prefix scheme (EVEX). This new scheme contains lots of new capabilities (source: https://en.wikipedia.org/w/index.php?title=AVX-512&oldid=841...):
- Expanded register encoding allowing 32 512-bit registers.
- Support up to 4 operands.
- Adds 7 new opmask registers for masking most AVX-512 instructions.
- Adds a new scalar memory mode that automatically performs a broadcast.
- Adds room for explicit rounding control in each instruction.
- Adds a new compressed displacement memory addressing mode.
I'm really not convinced by this list.
But it's difficult to practically use AVX-512 when the turbo boost throttling prohibits breaking even with an equivalent AVX/AVX2 workload. I had a small project that would trivially scale up to AVX-512 register sizes. Despite doubling the vector width, it was actually much slower than the AVX/AVX2 version in practice -- just because it had such a low turbo boost ratio when running AVX-512 instructions.
This was also a problem with the first AVX implementations, so this is very typical: In the first generation, Intel makes such an instruction set extension available, so that one can write applications that make use of it (though they will usually not be faster, sometimes even slower). In the following generation this new feature is made fast, so that the newly written algorithms really get to profit from it.