> we combine the vector masks using the packs instruction and readily extract it using movemask just once
Blend instructions are generally slightly faster than pack instructions, _mm256_blend_epi16 in that case. The order is not the same so the permute() function needs different shuffle vector, but otherwise it should work the same way.
> B-tree from Abseil to the comparison, which is the only widely-used B-tree implementation I know of
Another one: https://github.com/tlx/tlx/tree/master/tlx/container It’s not using any SIMD but the code quality is OK, I didn’t have much issues vectorizing the performance-critical pieces in my projects.