Can anyone comment on where efficiency gains come from these days at the arch level? I.e. not process-node improvements.
Are there a few big things, many small things...? I'm curious what fruit are left hanging for fast SIMD matrix multiplication.
Are there a few big things, many small things...? I'm curious what fruit are left hanging for fast SIMD matrix multiplication.
NB: Hobbyist, take all with a grain of salt
It seems to me that floating point math (matrix multiplication) will over time mostly disappear from ML chips, as Boolean operations are much faster both in training an inference. But currently they are still optimized for FP rather than Boolean operations.
https://semiengineering.com/speeding-down-memory-lane-with-c...