Speaking about pre-shuffling and blending, you can try doing twice as much work with vectors. Pre-shuffle slices of 32 keys into [ 0, 4, 8, 12 .. ], [ 1, 5, 9, 13.. ] etc, then when searching combine 0-2 and 1-3 pairs of vectors with _mm256_blend_epi16(a, b, 0b10101010), then merge with _mm256_blendv_epi8 using _mm256_set1_epi16((short)0xFF00) for the blending mask.
This way _mm256_movemask_epi8 will get you a bitmap of 32 bits for 32 integer keys you can scan with _mm_tzcnt_32. Especially good on AMD where vpblendvb completes in a single cycle, Zen 3 can even run 2 of them per clock.