"M1 core which is [...] much wider than both AMD and Intel" The core width you're referring to is the decoder width yeah? As another poster pointed out the M1 has a large reorder buffer as well. Combined with the ability to index a larger L1 cache what I'm getting here is that the M1 can be a lot better about scheduling instructions (and perhaps even running non-interfering instructions in parallel on a given core (is that a thing?)) because the frontend has more power and space to do so.
I guess that efficiency is then a big part of the puzzle on why the increased bandwidth to ram makes such an impact?
Even on the CPU side it can be a big win, in particular on SpecFPRate (a collection of heavy floating point real world codes, not microbenchmarks) Anand has this to say: The fp2017 suite has more workloads that are more memory-bound, and it’s here where the M1 Max is absolutely absurd. The workloads that put the most memory pressure and stress the DRAM the most, such as 503.bwaves, 519.lbm, 549.fotonik3d and 554.roms, have all multiple factors of performance advantages compared to the best Intel and AMD have to offer.
To drive this home compare the Spec2017 FP Rate, the M1 Max gets 81.07, the Ryzen 5950x (high end desktop with twice as many fast cores and a 105 watt TDP) gets 62.27.
So a low power M1 Max with half as many cores and much lower power is 30% faster than AMD's highest end desktop chip. Instead of a desktop size/volume/power, you can get it in a laptop that's 2/3rd of an inch thick.