https://news.ycombinator.com/item?id=6842338
It seems that the solution is not the top voted comment, but due to mgraczyk.
I just confirmed that the results are the same on Skylake, but I don't have an explanation yet, although I don't think it's branch prediction related.
You can prove/disprove that by taking the branch out of the question and just generating N calls to increment. Or N increments interspersed with a call to NOP.
That said, micro-architecture abuse at this level is rarely applicable to non-synthetic workloads in my experience.
The person who could look at this and just whip out an amazingly brilliant and nuanced explanation of what was going on is Ian Taylor over at Google. I am always in awe of some of the amazing ways that he and the gcc team there could increase performance in something already highly optimized.
https://news.ycombinator.com/item?id=6842872
mgraczyk:
"I think It's because the branch target is a memory access, so the dcache causes the execution pipe to stall on the load. In the tight loop with a call, the branch target is a call and there is time enough to pull from dcache before the load data is needed. I suspect that the i7 can't pull data from dcache immediately the jne instruction, so you get a hiccup. Try adding a second noop in to tightloop as the target for jne."
and
https://news.ycombinator.com/item?id=6844264
pbsd:
"The CPU is using store forwarding to cache the 'memory' accesses in the store buffer, which means most accesses are not even accessing L1 cache (if this were the case, we would not have such a low count of cycles per iteration)"
Again the reason I suspect the branch predictor in this sort of case is that when the loops are essentially 100% inside the cache, practically the only thing that varies the actual execution rate is whether or not a branch is not predicted. That said, the flow through these sorts of pipelined execution units is anything but clear.