Mishandling aside, the issue I've seen is there really isn't consumer demand for this. Prior to AMD having AVX512, most of the comments were around wasting the silicon on SIMD, rather than improving other aspects of the CPU. I'm pretty sure there was good reason to think it was largely a dark area of the chip.
From what I've seen, but haven't heard discussed much, the naive implementation vs AVX512 is a huge gain, but AVX2 vs AVX512 was not very impressive for the application I was looking at. The complexity this code added, and the cases where we needed it to run on AMD (for other reasons), basically made taking advantage of the feature undesirable for a single digit gain.
Things like VNNI or AMX are better wins, but they are only needed in very specific cases. VNNI in particular looked to be a 30% improvement in a BERT workload.