As best I can tell, it would be reasonable to expect around 12 AVX2 cores operating on the same power budget as the current 8 AVX512 cores. Does the average user have enough AVX512 workloads to make that tradeoff worth it?
As best I can tell, it would be reasonable to expect around 12 AVX2 cores operating on the same power budget as the current 8 AVX512 cores. Does the average user have enough AVX512 workloads to make that tradeoff worth it?
Any media processing can generally gain significant performance with SIMD instructions. Even web browsing is a workload that is affected, as (among other things) libjpeg-turbo uses SIMD instructions to accelerate decoding. It doesn't yet use AVX512, but it does use AVX2. If there are gains with AVX2, why wouldn't there be with AVX512.
Even if AVX512 would remain 256 bits, it would be an improvement over AVX2 as new instructions enable acceleration of workload that were not possible before, plus there is an increase in productivity/ease of use.
However, AVX512 won't be used much for consumer applications until there is big enough market penetration of CPUs that support these instructions, as it has been the case with any new SIMD instructions when they were first introduced.
To a limit. JPEG (and many video codecs based on JPEG) have 8x8 macroblocks, which means the "easiest" SIMD-parallel is 64-way. And AVX512 taken 8-bits at a time is in fact, 64-way SIMD.
To get further parallel processing after that, you'll probably have to change the format. GPUs go up to 1024-way NVidia blocks (or AMD Thread groups), which are basically SIMD-units ganged together so that thread-barrier instructions can keep them in sync better. 1024-work items corresponds to a 32x32 pixel working area.
But that's no longer the format of JPEG. It'd have to be some future codec. Maybe modern codecs are seeing the writing on the wall and are increasing macroblock size for better parallel processing 10 years into the future (they are a surprisingly forward looking group in general).
We did indeed do this for JPEG XL - the future is now :) 256x256 pixel groups are independently decodable (multi-core), each with >= 64-item (float) SIMD.
Cache lines are probably 64 wide for the purpose of burst length 8 (64 bit burst length 8 is 64 bytes / 512 bits).
There was AMD's 3DNow! that saw limited software adoption, because Intel didn't support it. Newer instruction sets are getting adopted progressively slower as consumers are replacing computer less often, AMD is slower at adopting each new AVX instruction set and Intel is getting more aggressive with market segmentation. Because market penetration of new instruction sets is getting slower, SW adoption is also much slower.
Not even all gamer CPUs have SSE4 (https://store.steampowered.com/hwsurvey), so it seems that runtime dispatch is unavoidable.
Given that, if we can afford to generate code for new instruction sets and bundle it all into one slightly larger binary/library, the problem is solved, right?
Highway makes it much easier to do that - no need to rewrite code for each new instruction set. As to Intel's market segmentation, we target 'clusters' of features, e.g. Haswell-like (AVX2, BMI2); Skylake (AVX-512 F/BW/DQ/VL), and Icelake (VNNI, VBMI2, VAES etc) instead of all the possible combinations.
It's less than 2x, because the core downclocking for AVX-512 on older CPUs is higher than for AVX2: 60% vs 85% on Skylake, so only a ~1.4x speedup. Newer CPU architectures do not downclock, though.
Newer CPU architectures may not have enforced downclocks, but they will eventually respond to the higher thermal output.
[1] https://en.wikichip.org/wiki/intel/microarchitectures/sunny_...
[2] https://www.realworldtech.com/forum/?threadid=193291&curpost...
That analysis might be slightly underestimating the area penalty of AVX512, because the consumer Skylake cores that didn't have AVX512 execution units still reserved space for the AVX512 register file. (And given that fact, it's all the more surprising that while Intel was repeatedly refreshing 14nm Skylake for the consumer market, they never added the rest of the AVX512 bits or redid the layout of the consumer cores to reclaim the blank space of the register file.)
If you're going to hypothesize about a rearranged Skylake core that no longer reserves space for the full 512 bit vectors in the register file, then you probably should also try to estimate the other area savings that would result from truly removing AVX512 from the core design at a high level, rather than merely masking off a few regions of silicon that serve no purpose other than AVX512 support. And answering that question is a lot more subjective than simply tallying up the area occupied by the extra execution units hanging off the side of the AVX512-enabled server Skylake cores.
Or, to put it another way: there unquestionably is an area cost to AVX512 beyond that of the extra execution units tacked on to SKL-X cores. But that cost was already being paid by consumer CPUs years before SKL-X/SKL-SP shipped. So it wasn't really a contributing factor to the loss of two cores when Skylaked-derived Comet Lake was succeeded by Rocket Lake.