Wide-datapath designs will generally have some tradeoffs like needing to ramp up power delivery, so there will be things like initial instruction latency for wide-datapath instructions, etc. That's kind of inevitable; I suspect the same will be true of the new fancy "variable length" vector ISAs if the underlying implementation and vector usage is wide enough, too.
Also: you don't need to use wide 512-bit vectors with AVX-512! You can use the instruction set with 128-or-256/bit vectors just fine.
https://travisdowns.github.io/blog/2020/08/19/icl-avx512-fre...
Actually supposedly Skylake-X/Xeon-W had much lower AVX downclocking than Skylake-SP too... InstLat64 made a tweet at one point showing this was 10-20% for workstation vs 30-40% for server iirc. Tweet has been removed unfortunately.
Intel really throttled it down on server chips, for whatever reason. Probably didn't want datacenter chips to run the Unlimited Voltage that was necessary for full-clock dual-unit AVX-512 on 14nm.
My 10980XE runs AVX2 at 4.2 GHz all core, and AVX512 at 4GHz.
[ General observation, not directed at parent comment: ]
Frequency throttling, even on the most affected Skylakes, has always been a non-issue if you run say 1ms worth of continuous SIMD instructions. How could a 10-40% drop negate speedups from 2x vector width plus double the registers and a much more capable instruction set?
It is time we buried this myth :)