That's pretty disingenuous. The standard implementation is known to be terribly slow because of constraints. They should compare against current state of the art.
That's pretty disingenuous. The standard implementation is known to be terribly slow because of constraints. They should compare against current state of the art.
I’m sure an exotic watercooled setup will fare much better, but those aren’t generally what we run in production.
Even on a desktop this sort of thing is sometimes necessary, for example my CPU has different clock speeds depending on how many processors are running, so I have to lock it to the all-core clock if I want to see proper parallel speedups.
This might be annoying for day-to-day usage (although, CPUs really are insanely performant nowadays so maybe it will not be too bad).
By the way the performance penalty for using AVX-512 on multiple cores when the multiple cores were already active is zero. There is no penalty in most server scenarios.
That is a penalty due to licensing [0], not thermal throttling. As I wrote elsewhere, I’ve seen my clockspeed get cut in half across all cores on a physical die when running AVX-heavy operations for a sustained period of time, due to thermal throttling.
[0] https://travisdowns.github.io/blog/2020/08/19/icl-avx512-fre...