2x/4x/8x is still thinking in terms of autovec.
I'm getting 50x faster code with manual ASM. That's the difference between audio code that runs in realtime and code that does not.
The competing implementations use SIMD and native code, autovec works nicely there. I symbolically invert the LinAlg system at compile time.