From what I've seen, but haven't heard discussed much, the naive implementation vs AVX512 is a huge gain, but AVX2 vs AVX512 was not very impressive for the application I was looking at. The complexity this code added, and the cases where we needed it to run on AMD (for other reasons), basically made taking advantage of the feature undesirable for a single digit gain.
Things like VNNI or AMX are better wins, but they are only needed in very specific cases. VNNI in particular looked to be a 30% improvement in a BERT workload.