> What benchmarks? Not a lot of real programs actually use AVX512, and a lot of the ones that did discovered that it made performance worse, so they stopped.
This paints a picture of what happened to Intel's fab process rather than AVX-512 itself.
AVX-512 was proposed back in 2013. It was NOT designed for desktops. The original designs were for their Phi chips (basically turning a bunch of x86 cores into a GPU). These Phi chips ran between 1 and 1.5GHz, so power consumption and clocks were always matched up.
Intel wanted to move these instructions into their HPC CPUs. The problem at hand was ultra-high turbo speeds. These speeds work because the heat is a bit spread out on the chip and a lot of pieces are disabled at any given time. With AVX-512, they had 30% of the core going wide open for the vector units (not to mention the added usage from saturating the load/store bandwidth).
They wanted 10nm and then 7nm to fix these issues. At their predicted schedule, 10nm would have launched in 2015 and 7nm in 2017.
Given the introduction of AVX-512 in 2013, they had plenty of time. In fact, the first official product was Knight's landing in 2016. Skylake with AVX-512 didn't launch until 2017 when they were supposed to be on 7nm.
Intel was forced to backport their designs to 14nm++++++++++++ which forced a deal with the devil. They had to downclock the CPU to keep within thermal limits, but this slowed down EVERYTHING. Maybe they could have created a separate power plane for AVX, but this would be a BIG change (and probably politically infeasible).
What happens with the downclocking? If you run dedicated AVX-512 loads, then maybe you should have been looking at Phi instead. If not, mixed loads suffered overall because of the lower clockspeeds.
Their second revision of 10nm superfin is still a generation larger (well, probably a half-generation) than what they anticipated. There might still be downclocking, but I'd guess that it's nowhere near what previous iterations required.
TL;DR -- AVX-512 was launched two nodes too early which screwed over performance. It should become acceptable either with 10nm Superfin or 7nm when it launches so the CPU doesn't have to downclock constantly.