Even when they were new, they competed with AMD's high end desktop chips. Many years later, they're still excellent in the laptop power range - but not in the desktop power range, where chips with a lot of cache match it in single core performance and obliterate it in multicore.
https://www.cpu-monkey.com/en/compare_cpu-apple_m4-vs-amd_ry...
And in laptop form compared with a m4 max: https://www.cpu-monkey.com/en/compare_cpu-apple_m4_max_14_cp...
Why does it matter how they achieved their thunderous performance? Why must it be diminished to just a boatload of cache? Does it matter from which implementation detail you got the best single-core performance in the world? If it's just way more cache, why isn't Intel just cranking up the cache?
It's worth noting that Intel is not a stranger to building CPUs with lots of cache - they just segmented it into their server chips and not their consumer ones.
It matters because it is useful to understand why a given chip is faster or slower than its competitors. Apple didn't achieve this with their architecture/ISA or with some snazzy new hardware (with some notable exceptions like their x86 memory emulator), they did it by noticing how important cache was becoming to consumer workloads.