If you search for "reciprocal throughput" in the article, the author did do another benchmark with dependencies to compare "real latency".
Most published data structure benchmarks have this pitfall, and don't discuss it, then when trying to measure the improvements after plugging into a real workload it turns out that the microbenchmark was not predictive. When I see latencies that top out at 50ns when they clearly have to make more than one causal load I immediately get skeptical.
This is one of the very few articles I've seen that clearly explains this, which makes it even more excellent. I only wish it didn't label the other graphs "Latency" too, or at least put an asterisk there.