AVX-512: when and how to use these new instructions
lemire.me
lemire.me
If I'm understanding OP, this means that to use the AVX-512 instructions well, a compiler that has to think about instruction speed as a function of what other instructions are around it. It might be faster to write operation X with these instructions than without, but only if you don't also write operation Y with them, because then the CPU would get too hot.
That sounds so much harder! Hot damn! I know CPUs are complicated and 1 instruction = 1 cycle is wrong in many ways, but this just sounds especially difficult.
Doable, but someone will have to take the first jump.
The power throttling and voltage gating that goes on takes a long time--at least microseconds, up to a few milliseconds. The scheduling concerns that compilers deal with are worried about around tens to hundreds of clock cycles, a factor of well over a thousand.
This is why it is advisable for SIMD libraries or libraries that use SIMD to always offer some levers for this kind of thing.
This makes it a pain as you may have to write the same function 2 or 3 times in different SIMD instruction sets, or use a library to do that for you:
Compilers can't hope to realistically handle this. JITs at least have a chance, but adding in handling for this behaviour surely requires a lot more complexity than I'd imagine most runtime developers would want to add to their code.
The exception would be if your kernel can make some good use of other wide instructions such as memory access or shuffles.
...though, come to think about it, this would be pretty easy with a tracing JIT.
One of the co-authors talks more about the exact limits here: https://www.realworldtech.com/forum/?threadid=179654&curpost...
The quick answer is that on the W-2104 system he tested, you can sustain one FMA every 2 cycles while still remaining in the medium speed L1 state.
His avx-turbo tool (https://github.com/travisdowns/avx-turbo) can be used to check the situation for your particular processor.
I think given the current state of things it would be irresponsible for compilers to generate heavy instructions unless asked. Forget trying to be smart about it ... we already fail to be smart about things that are much simpler and more visible.
More interestingly, this may be what all CPU behavior looks like in 10 years, because if Intel has to resort to this kind of design now, why would hat change any time soon? Instead of worrying about primarily keeping the execution units full, people trying to write fast code may be primarily concerned with keeping them NOT full so that the chip doesn’t slow down. Which sounds crazy and hard to deal with.
So to use cutting edge instructions you generally have to hand code them (either in intrinsics like Lemire did there or in asm)
For integer operations, auto-vectorization is more prevalent since everything can be reordered more freely. clang especially auto-vectorizes a ton of stuff even at -O2.
So we have now being at the closest point to the CPU power wall in history ...
I remember joke opcodes back in the day. One was "Halt and catch fire". Is that seriously what this is doing?
Indeed. If you have some piece of code in a different security context that conditionally executes a heavy instruction based on a decision made over some sensitive data, doesn't this provide a way to obtain information about that data?
Yes. http://www.numberworld.org/blogs/2018_6_16_avx_spectre/
Amusing Zen is emulating AVX512 and AVX2 via micro-op splitting and it performs better under some workloads.
The real issue is path propagation delays of 512bits worth of electricity is extremely non-trivial, and costs a shit load of power. Just `mov`'ing to the AVX-512 instructions (initially when AVX-512 is not warm) can stall the CPU for 10,000+ cycles as it tries to power on all those registers.
Also, source on the power up stall? AVX(2) didn’t have that and I’m highly surprised AVX512 would. Agner at least claims the same reduced throughput during warmup, but I think he only has early silicon.
On to Skylake-SP, however, that chip is reported to have both reduced throughput and fully halted periods in [2].
Some have speculated it has to do whether chips have an integrated IVR: the models with integrated IVR having less capability of handling high dI/dt events. I don't know about that though (Skylake-SP still has external VR, right?).
[1] https://www.agner.org/optimize/blog/read.php?i=378#378 [2] https://software.intel.com/en-us/comment/1926876#comment-192...
Changing the clock speed based on thermal headroom is really hard to do well but Intel does it well and most other chip-makers are trying to duplicate Intel. This is really the opposite of those old chips that would sometimes burst into flame because a modern CPU which throttles based on temperature will never burst into flame even if you remove the heat sink and put it in a 150 degree over.
https://software.intel.com/en-us/forums/intel-isa-extensions...
Wow, let's give the word "license" semantics not related to software licensing, and cause confusion between L1 meaning "L1 cache" and "license 1".
Let's us begin
However, there are also deterministic frequency
reductions based specifically on which instructions you
use and on how many cores are active (downclocking).
As you are likely using AVX-512 in a cloud deployment you don't have access to any of the information, and are likely sharing that hardware who may not respect the same engineering rigor.Also, nobody setting there CPU affinity to ensure the down-clocking stays on a single core. You have to pray the scheduler doesn't shuffle your workload around. This requires a lot of platform specific C and most people are writing Java/Go/Ruby/Python where doing bit twidling NUMA management is impossible, furthermore the information you have access to in a cloud environment (which is where you'll be using advanced AVX-512 unless you work for Amazon, Google, Intel, or Cloudflare) may just lie about core count, and NUMA architecture.
Also this is just false [1]. Running AVX-512 adjusts the BASE clock for the package. There are throttling attempts to ensure every other core on the package throttles back. This is a package wide effect, not a per-core effect.
Light instructions include integer operations other than
multiplication, logical operations, data shuffling
(such as vpermw and vpermd) and so forth.
This is false according cloudflare [2] which you've linked. They test your "light" carry less adding, shifting, and xoring (these are the only operations in ChaCha20 [4]). It cost too much. We have chosen to only include two columns.
I'll include the whole thing [3]. Wow yeah the entire package's power curve is changing. The base clock, cores not effected, all their clocks are changing. Its almost like having 1 out of 24 cores still effects all 24 cores.... For example, the openssl project used heavy AVX-512
instructions to bring down the cost of a particular
hashing algorithm (poly1305) from 0.51 cycles per byte
(when using 256-bit AVX instructions) to 0.35 cycles
per byte, a 30% gain on a per-cycle basis. They have
since disabled this optimization.
The literal example to show AVX-512 is good at ends with statement that people using AVX-512 are now actively avoiding it.This is less then content, do you have an agenda or are you just an idiot?
[1] https://en.wikichip.org/wiki/intel/xeon_silver/4116
[2] https://blog.cloudflare.com/on-the-dangers-of-intels-frequen...
The explanatory page accompanying that table contains text that contradicts your claim: https://en.wikichip.org/wiki/intel/frequency_behavior
> The frequency of each core is determined independently based on the workload described above. That is, cores running Non-AVX workloads can enjoy the full regular turbo frequency, whereas cores executing AVX-512 or AVX2 will operate at their own designated turbo frequencies. [...]
> In Haswell, an AVX2 workload on one core meant all cores were capped at AVX2 Turbo frequency. This had the undesirable effect of reducing performance for non-AVX workloads on cores that were unrelated to the cores executing AVX2 workloads. This behavior was changed with Broadwell which grouped cores executing AVX2 workloads together and cores executing non-AVX workloads separately, allowing the former cores group to execute at the lower AVX2 turbo frequency while having the later cores group execute at full non-AVX2 turbo.
Also:
> The literal example to show AVX-512 is good at ends with statement that people using AVX-512 are now actively avoiding it.
It's an example to show that cycle-by-cycle speedups due to the use of AVX-512 are not always worthwhile, especially in library code. Which was one of the points the article was trying to make. It's fine if that was an obvious point from your perspective, but it doesn't contradict the article and it doesn't make it "less than content".
I'm not a fan of the delivery, but I absolutely agree with you.
It's reminiscent of the infamous "disemvoweling" strategy used on a few other forums, where the reader is forced to decide whether they want to painstakingly reconstruct offensive and abusive comments or blindly trust someone else to restrict what they see.
Life would be so much easier if they just displayed the comment score like most other moderated forums and let the reader decide the merits of the comments based on visible information.
This was true of server CPUs before Skylake-SP, but on Skylake-SP this effect is per core. This was widely touted by Intel since it was an improvement over the old behavior.
What the other cores are doing still matters since the count of "active cores" is used to look up the turbo frequency for each other core, depending on its license: but for this purpose it only matters if the others are running or halted, not _what_ they are running.
If you don't believe it, I've shared a benchmark you can try yourself if you have access to Skylake server [1]. Run with --spec avx512_fma_t/1,scalar_iadd/3, for example, to kick off 1 core of heavy FMA ops in parallel with 3 cores of scalar-only ops. You'll see only the FMA core drops down to the L2 license.
Light instructions include integer operations other than
multiplication, logical operations, data shuffling
(such as vpermw and vpermd) and so forth.
> This is false according cloudflare [2] which you've
> linked. They test your "light" carry less adding,
> shifting, and xoring (these are the only operations
> in ChaCha20 [4]). It cost too much.Again, you can test this yourself with avx-turbo [1], there are a variety of tests there that show everything except FP operations and integer multiplications (which execute on the FP unit) are treated as light. Of course, the tests aren't exhaustive, but they hit the main categories of instructions.
Note that even light AVX-512 instructions cause the chip to transition to a lower frequency (the so-called L1 license, which is usually about half way between the fasted L0 and slowest L2 speed). So even if ChaCha doesn't use any heavy instructions, any AVX-512 at all will slow down your frequency. They only reported a 5% to 7% reduction in performance, which could easily be consistent with a downclock to the L1 frequency.
> The literal example to show AVX-512 is good at ends > with statement that people using AVX-512 are now > actively avoiding it.
I think the example is intended to show that AVX-512 indeed significantly speeds up the parts of the code where it is applied: but if that makes up a small part of your code overall, it might not be worth it, because the rest of your code may suffer a frequency penalty.
Unlike some earlier ISA extensions, there is no simple answer like "use it" or "don't": there are complex tradeoffs. That's what you should get out of this article.