Software performance is counterintuitive
lemire.me
lemire.me
That is never the problem. The problem is that doing this creates tight coupling between the two operations, making it marginally (or sometimes significantly) more expensive when you only need one of them. The coupling also makes your code harder to understand (in all but trivial cases). When the operations are naturally inseparable, this approach makes sense. It also makes sense if you are trying to optimise out a specific problem after profiling.
As a general design, it's usually an anti-pattern and one of the reasons premature optimisation gets such a bad rap with Knuth et al.
Also this case is pretty trivial. There are many cases when trying to couple together multiple operation results in an explosion of edge-cases with unexpected side effects that end up making performance worse, unless you are super careful. This is when "just pick a better algorithm" is usually applicable, except if you tightly couple two algorithms together, it's much harder to swap one out.
The sad thing is that blog posts like that are easily found on Google by junior-level developers who don't know any better and take all the numbers on faith as representative of the language, even when the methodology is almost certainly complete nonsense.
Just because the additional instruction did not take more time (or cycles), doesn't mean it was free. Say there are several instructions any of which can be subbed in for `__builtin_popcountl` and still fit into the same cycle. Several used together may add up and cost a cycle on each round of the loop.
So although the added instruction doesn't cost any additional time. It isn't free.
One caveat to be aware of in the article's example: hyper-threading works by letting a second thread run on the CPU's execution ports that are unused by the first thread. Consequently, if you design your code to do an excellent job of saturating execution ports, it will often mean that hyper-threading is slower than using a single thread per core (all of the additional cache pressure, few free execution ports that allow the code to run). However, the speed improvements you can see on a single thread with excellent execution port saturation will often be higher than the speed improvements you will see from hyper-threading code that does not saturate the execution ports due to the aforementioned cache overhead of running a second thread concurrently.
Asymptotic complexity is the best place to start. I thank academia for this 'naive' model.