So how many hardware systems does Apple silicon have for doing matrix multiplies now?
1. CPU, via SIMD/NEON instructions (just dot products)
2. CPU, via AMX coprocessor (entire matrix multiplies, M1-M3)
3. CPU, via SME (M4)
4. GPU, via Metal (compute shaders + simdgroup-matrix + mps matrix kernels)
4. Neural Engine via CoreML (advisory)
Apple also appears to be adding a “Neural Accelerator” to each core on the M5?