Here is a version of the code that runs even faster than the asm("call foo") one by inserting nops between the load and the use of the loaded value. However it became even faster by inserting a few nops at the beginning of the loop too.
The fastest version that I could find is to do a prefetch of counter at the end of the loop .... speaking of which isn't that nopw that GCC uses for alignment a prefetch instruction too? Is the CPU tricked by the fake address used there?
void prefetch() {
unsigned j;
for (j = 0; j < N; ++j) {
counter += j;
__builtin_prefetch(&counter, 0, 3);
}
}
void nop_wait() {
unsigned j;
for (j = 0; j < N; ++j) {
unsigned x = counter;
/* 4 or more nops seem right, 3 nops are slower */
__asm__("nop");
__asm__("nop");
__asm__("nop");
__asm__("nop");
counter = x + j;
}
}
tightloop:
3,000,314,154 cycles # 0.000 GHz ( +- 0.01% )
2,400,971,094 instructions # 0.80 insns per cycle ( +- 0.00% )
0.788256283 seconds time elapsed ( +- 0.02% )
loop_with_call:
3,000,314,154 cycles # 0.000 GHz ( +- 0.01% )
2,400,971,094 instructions # 0.80 insns per cycle ( +- 0.00% )
0.788256283 seconds time elapsed ( +- 0.02% )
nopwait:
2,679,341,970 cycles # 0.000 GHz ( +- 0.37% )
4,000,906,874 instructions # 1.49 insns per cycle ( +- 0.00% )
0.704209471 seconds time elapsed
vs prefetch:
2,586,821,497 cycles # 0.000 GHz ( +- 0.54% )
2,800,888,975 instructions # 1.08 insns per cycle ( +- 0.00% )
0.679998564 seconds time elapsed
This is on a Intel(R) Core(TM) i7-2600 CPU @ 3.40GHz