Lesson learned, sometimes the easiest performance gains are found by not being naive about memory access. The extra instructions were inconsequential.
Impressive, but definitely beatable given sufficient free time.
There's no chance you could achieve same by using scalar instructions. SIMD can access memory a lot faster than scalar.
Random access is another matter. The trick is of course to avoid non-sequential access patterns.
IOW: Are the scalar instructions slower than memory bandwidth?
I work with image processing, >1GiB/s. We optimize for each platform for hand, many embedded platforms but also x86_64. Most of our work is detectors and loss-less compression and we could not do these things in realtime if it weren't for SIMD/MIMD.
If your bottleneck isn't memory, then yeah, SIMD is a real boon.
Especially since high performance software is glad to have even a 1% boost in speed.