The only exceptions are the loads/stores from/to the L1 cache, which have double throughput on Intel and the floating-point fused-multiply-add units, where the most expensive Xeon SKUs can do 2 FMAs per cycle, while Zen 4 can do only 1 FMA + 1 FADD per cycle.
Zen 4 implements the BF16 instruction set, which is likely to increase the speed many times for any AI/ML workload that uses BF16. It also implements the VNNI instruction set, which will accelerate any inference that uses INT8.
Even when these dedicated instructions are not used, AVX-512 is usually much faster on Zen 4, by eliminating bottlenecks caused by instruction fetch and decoding and by using the better designed AVX-512 instructions.
If you make simple 1-to-1 transition from AVX256 to AVX512, the speedup is usually <2x, unless there is a major bottleneck in instruction fetch and decoding. Also AVX512 in Zen is still double-pumped. Regarding FMA, again, if you compare FMA in AVX256 and AVX512, latencies and throughput are the same[1].
Comparing performance between different datatypes is probably fine, but it should be stated directly. Unless "Owners of CPUs like Zen4 can expect to see 10x faster prompt eval times" means comparison with Skylake.
[1] https://uops.info/table.html?search=vfmadd132ps&cb_lat=on&cb...
And as the sibling comment mentions, AVX512 is not just about the width but also about newer instructions available at 128b and 256b widths as well under AVX512VL.
Intel has only three 256-bit execution units, which are also ganged into two 512-bit execution units, by adding an extra 256-bit unit, which stays idle in 256-bit mode.
So the throughput for most register-register 512-bit operations is the same for Intel and AMD, except that AMD has a single FP64 multiplier vs. two FP64 multipliers on Intel and that the path to the L1 cache has double width on Intel (while AMD does one 512-bit load per cycle + one 512-bit store every other cycle, Intel can do two 512-bit loads from L1 + one store per cycle).