The only exceptions are the loads/stores from/to the L1 cache, which have double throughput on Intel and the floating-point fused-multiply-add units, where the most expensive Xeon SKUs can do 2 FMAs per cycle, while Zen 4 can do only 1 FMA + 1 FADD per cycle.
Zen 4 implements the BF16 instruction set, which is likely to increase the speed many times for any AI/ML workload that uses BF16. It also implements the VNNI instruction set, which will accelerate any inference that uses INT8.
Even when these dedicated instructions are not used, AVX-512 is usually much faster on Zen 4, by eliminating bottlenecks caused by instruction fetch and decoding and by using the better designed AVX-512 instructions.