Here is a version of the code that runs even faster than the asm("call foo") one by inserting nops between the load and the use of the loaded value. However it became even faster by inserting a few nops at the beginning of the loop too.
The fastest version that I could find is to do a prefetch of counter at the end of the loop .... speaking of which isn't that nopw that GCC uses for alignment a prefetch instruction too? Is the CPU tricked by the fake address used there?
void prefetch() {
unsigned j;
for (j = 0; j < N; ++j) {
counter += j;
__builtin_prefetch(&counter, 0, 3);
}
}
void nop_wait() {
unsigned j;
for (j = 0; j < N; ++j) {
unsigned x = counter;
/* 4 or more nops seem right, 3 nops are slower */
__asm__("nop");
__asm__("nop");
__asm__("nop");
__asm__("nop");
counter = x + j;
}
}
tightloop:
3,000,314,154 cycles # 0.000 GHz ( +- 0.01% )
2,400,971,094 instructions # 0.80 insns per cycle ( +- 0.00% )
0.788256283 seconds time elapsed ( +- 0.02% )
loop_with_call:
3,000,314,154 cycles # 0.000 GHz ( +- 0.01% )
2,400,971,094 instructions # 0.80 insns per cycle ( +- 0.00% )
0.788256283 seconds time elapsed ( +- 0.02% )
nopwait:
2,679,341,970 cycles # 0.000 GHz ( +- 0.37% )
4,000,906,874 instructions # 1.49 insns per cycle ( +- 0.00% )
0.704209471 seconds time elapsed
vs prefetch:
2,586,821,497 cycles # 0.000 GHz ( +- 0.54% )
2,800,888,975 instructions # 1.08 insns per cycle ( +- 0.00% )
0.679998564 seconds time elapsed
This is on a Intel(R) Core(TM) i7-2600 CPU @ 3.40GHzIn fact replacing the nopw that is used for alignment by a prefetch instruction gives me something slightly even faster:
00000000004004e0 <prefetchit>:
4004e0: 31 c0 xor %eax,%eax
4004e2: 0f 18 0d 87 04 20 00 prefetcht0 0x200487(%rip) # 600970 <counter>
4004e9: 48 8b 15 80 04 20 00 mov 0x200480(%rip),%rdx # 600970 <counter>
4004f0: 48 01 c2 add %rax,%rdx
4004f3: 48 83 c0 01 add $0x1,%rax
4004f7: 48 3d 00 84 d7 17 cmp $0x17d78400,%rax
4004fd: 48 89 15 6c 04 20 00 mov %rdx,0x20046c(%rip) # 600970 <counter>
400504: 75 dc jne 4004e2 <prefetchit+0x2>
400506: f3 c3 repz retq
400508: 0f 1f 84 00 00 00 00 nopl 0x0(%rax,%rax,1)
40050f: 00
2,445,505,116 cycles # 0.000 GHz ( +- 0.52% )
2,800,857,971 instructions # 1.15 insns per cycle ( +- 0.00% )
0.643010038 seconds time elapsed volatile unsigned dummy = 0;
void loop_dummy_read() {
IACA_START;
unsigned j;
unsigned dummy_read;
for (j = 0; j < N; ++j) {
dummy_read = dummy;
counter += j;
}
IACA_END;
}There are remarkably few memory accesses actually being performed in that loop. The CPU is using store forwarding to cache the 'memory' accesses in the store buffer, which means most accesses are not even accessing L1 cache (if this were the case, we would not have such a low count of cycles per iteration).
The best result I've got is by inserting a 6-byte nopw right before the store:
.L3:
movq counter(%rip), %rdx
addq %rax, %rdx
addq $1, %rax
cmpq $400000000, %rax
nopw 0x1(%rax,%rax,1)
movq %rdx, counter(%rip)
jne .L3That's a P4 thing, not a Core thing http://en.wikipedia.org/wiki/CPU_cache#Trace_cache
Edit: ok, Core has a uop cache http://en.wikipedia.org/wiki/Micro-operation_cache#Micro-ope...
You don't need a trace cache to do what you suggest, it just predicts the branch and goes from there