Thanks! You're right, I missed it because it was in the section "Comparison with std::lower_bound" and I thought there would be just some gloating :)
Most published data structure benchmarks have this pitfall, and don't discuss it, then when trying to measure the improvements after plugging into a real workload it turns out that the microbenchmark was not predictive. When I see latencies that top out at 50ns when they clearly have to make more than one causal load I immediately get skeptical.
This is one of the very few articles I've seen that clearly explains this, which makes it even more excellent. I only wish it didn't label the other graphs "Latency" too, or at least put an asterisk there.