v210_planar_pack_8_ssse3: 402.5
v210_planar_pack_8_avx: 413.0
v210_planar_pack_8_avx2: 206.0
v210_planar_pack_8_avx512: 193.0
v210_planar_pack_8_avx512icl: 100.0
23x speedup. The compiler isn't going to come up with some of the trickery to make this function 23x faster.
800% is nothing.