(Intel Hotchips slide: https://images.anandtech.com/doci/15984/202008171757161.jpg)
(Intel Hotchips slide: https://images.anandtech.com/doci/15984/202008171757161.jpg)
I don't think anyone ever said that.
What people are really saying, AVX-512; in its current form, limited by Clockspeed, Thermals, and most importantly market segmentation.
And given all these three things are only slightly improved in the current 3 years roadmap render the instruction useless. ( But latest HotChip information seems to suggest they are well aware of it. So Roadmap can change )
And even more importantly that is in reference to current stack of Intel issues.
Yeah, Torvalds said pretty much exactly that, in his typically boisterous form.
That it is "only useful in special cases", "only exists to make Intel look better on benchmarks", and that "the transistors would be better spent elsewhere".
https://www.phoronix.com/scan.php?page=news_item&px=Linus-To...
Are you disagreeing with the parent comment or not? Are you maybe saying that the instruction is useless in chips that are in computers right now but will be useful in chips that will be released in future?
I mostly agree, but for me the frequency scaling is almost irrelevant. Here’s a link, expand “Other settings” and you’ll see: https://store.steampowered.com/hwsurvey
If you implement AVX2, the code will run on 76% of PCs. When you want good performance, such optimization is a good use of limited resources.
AVX512 is only at 0.42%. That’s why it’s useless to implement AVX512 except for servers where you have full control over hardware. And with current state of Intel, not necessarily a good idea even for servers, it’s quite possible in a couple of years you’ll want to switch to AMD.
From what? Automatic vectorization is very limited, it’s only good for pure vertical algorithms.
> and dynamically choose depending on CPU capabilities
Only Intel compiler does that, the rest of them don’t.
Sometimes I ship multiple versions of binaries. Other times I do runtime dispatch myself, it’s only couple lines of code, __cpuid() then cache some function pointers, or abstract class pointers. It’s even possible to compile these implementations from the same C++ source, using templates, macros, and/or something else.
It would have been nice to see the vcore and thermal values graphed as part of the benchmark. Do they increase faster for AVX-512 vs the other instruction sets?
I've had problems in the past with Sandybridge, where AVX hit thermal throttles before SSE. I ended up having to disable them in my build because of it. Presumably, the same behaviour would be seen here now that the vector unit is wider and there are more densely packed transistors flipping.
You you use more power (and get hotter temps: these are exactly proportional, so you can mostly just talk about them as one) with wider vectors because you are doing more work. When you look at it on a per-element basis, you use less power per element with wider vectors. E.g., you might use 1 pJ per element for 256-bit FMA but only 0.8 for 512-bit FMA.
Of course, since you can do 2x as many total elements in 512-bits on a 2 FMA machine, you can be both more efficient but use more total power, so you can get TDP or thermal limits with 512-bit code that you wouldn't on 256-bit, but it should still per faster and more efficient per element.
All of this assumes you can usefully use the 2x more work with the larger vectors. Sometimes the scaling is worse: e.g., for lots of short arrays, when a lookup table is involved, when additional shuffling or transposition is required with larger vectors, etc. In that case you could end up less efficient with larger vectors.
If the only issue with AVX-512 is thermal downclocking because you end up using more power, it's almost definitely because you are getting more work done per time. A few AVX-512 instructions in a mostly scalar workload is not going to significantly increase power dissipation and therefore should not induce thermal downclocking, while a heavily utilized AVX-512 kernel will burn power, but should also be doing work twice as fast per instruction.
If I had a little background deamon that used 512 because they were cool or in the hotchips presentation, and that bonked my overall system performance that would be annoying.
It's also annoying because intel benchmarks with no mitigations. So what can happen is you think you should be seeing X performance, and then with mitigations applied and some 512 instructions hitting you are Y performance.
And it looks like they've reduced how often a single instruction will cause a lockup as the core shifts to a different power level. But until they've eliminated that issue, it's still scary to toss in a few AVX-512 instructions.
> There is a transition period (the rightmost of the two shaded regions, in orange14) of ~11 μs15 where the CPU is halted: no samples occur during this period16. For fun, I’ll call this a frequency transition.
It stops executing for 35 thousand cycles. I call that a "lockup" "as it shifts".
From later in the same post:
> Here, we have the worst case scenario of transitions packed as closely as possible, but we lose only ~20 μs (for 2 transitions) out of 760 μs, less than a 3% impact. The impact of running at the lower frequency is much higher: 2.8 vs 3.2 GHz: a 12.5% impact in the case that the lowered frequency was not useful (i.e., because the wide SIMD payload represents a vanishingly small part of the total work).
Interestingly enough, this is another feature that is supposed to have been improved on server Icelake. The frequency transition halt time is now pretty much negligible. The "core frequency transition block time" goes from ~12 us on CLX (similar to the number quoted above) to ~0 us on ICX.
(Slide with frequency transition info: https://images.anandtech.com/doci/15984/202008171754441.jpg)
Yes, if you insert a single 512-bit FMA that runs every so often in your code you will get a 15% performance hit from the lower frequency, but that's much less likely than the old case where people who were trying to use AVX-512 for memcpy and the like would slow down scalar code.
Now that the older and bigger case is fixed, this case remains the last sticking point. Because you still can't trust the CPU to do the right thing when there are a small number of heavy instructions. Even if they cut the halting time to 0, it's still bad for a single instruction to cause a prolonged downclock.
Some of that is from overclocking. Some is from old and/or defective power supplies. Some is from motherboard VRMs. And some, like the original Ryzen 1700X, is from bad SOC-internal power management.
At any rate, I have read forum posts reporting system failures caused by AVX. Either overcurrent or clockrate changes crashing it.
So it seems to have support. I wouldn't call that "misinformation."
There are plenty of people with both Intel and Ryzen systems that are straight-up broken and don't power on at all.
Defects happen. Improperly designed systems happen. Misconfigurations happen. While those situations are unfortunate for the small percentage of people experiencing them, they shouldn't be used to judged the platform's capabilities as a whole.
> So it seems to have support. I wouldn't call that "misinformation."
Using anonymous, largely unverifiable anecdotes posted on web forums as evidence for a population-wide problem is a textbook case of selection bias.
And if you're OK with that, let me throw in my own anecdote.
All of my systems, which include:
* Skylake, Coffee Lake, Haswell/Broadwell, Sandy Bridge/Ivy Bridge, and Westmere/Nehalem Intel processors,
* Zen 2, K10, and K8 AMD processors,
...have been able to reliably execute supported vector instructions (SSE, AVX, etc.) for extended periods of time without any problems.