But their point is that
>every hand-optimized vectorized x86 routine
is approximately 0 percent of code.
>every hand-optimized vectorized x86 routine
is approximately 0 percent of code.
But very often approximately 0 percent of a code is 90+% of a runtime
I know I've seen pretty huge speedups in my own code for "free" just from switching to an AVX version of BLAS. You can just think about how many different programs use BLAS (which is itself highly arcane internally), and AVX is definitely in a ton of other low level libraries out there.