However, I have some comment on the benchmark, the iterations have no loop dependency on the previous result, so the CPU is allowed to pipeline some loads. This is rarely representative of real workloads, where the key being searched comes from other computations, so this is more a measure of highest possible reciprocal throughput than actual latency. Would be interesting to see what happens if they make the key being searched a function of the result of the previous lookup. (EDIT: as gfd below points out, the article actually discusses this!)
Also some of the tricks only work for 32-bit keys, which allow the vectorized search in nodes. For larger keys, different layouts may still be preferable.
[1] Ordered search is when you need to look up keys that are not in the dataset, and find predecessor/successor. If you only do exact lookups, a hashtable will almost certainly perform better, though it would require a bit more space.