If I have a core with four 128-bit neon vector units, I have the same throughput as an x86 with two AVX2 units or one AVX512 unit. However that 4x128-bit core is actually more flexible than the other two as I can do 4 different things at once, or 4 scalar operations per cycle. (Of course the downside is you spend more frontend resources on decode).
Given that most code isn't vector code, the multiple short vector length approach is actually superior on many common real-world workloads that aren't machine-learning (and CPU is unit-of-last-resort for large ML workloads anyway).
The M1 and later big cores can dispatch four NEON FMA instructions per clock, so 512 bits worth of vector math, which compares OK with most Intel or AMD chips (Zen 4 can do two 256-bit MUL and two ADD, and Intel "client" bigcores since Sunny Cove typically do three 256-bit FMA).
spoiler: it isn't