Novel Optimization Technique
swapped.tumblr.com
swapped.tumblr.com
Is there any code before or after that has been omitted that might be contributing to the 5% speedup?
I have written my share of hand-optimized assembly, but this is just an optimizer peculiarity. I am not after trying to understand why it happened, and I was making a joke calling it a "technique". That said, here's an assembly dump of both versions, just in case :) -
Slower version - https://gist.github.com/4454234
Faster version - https://gist.github.com/4454238
Still, it would be very interesting to be able to actually try to analyze what's giving the speed-up.
p = malloc(); free(p);
in a tight loop for 5 seconds, generating about 60,000,000 traversals of the code in question. The slower version yields a loop count of around 67,... and the faster version bumps it to over 70,... Not quite 5%, but in thereabouts.... also, have you considered pointer arithmetic? your style of for loop is nothing i would ever write to begin with. perhaps this is why the compiler has a hard time with it - it is quite unusual - the extra loop counter seems pointless and can be removed by a single pre-calculation
Good point. Tried it and makes no difference. Looking at the assembly, it appears that the difference in performance lies in the prolog code for the loop rather than in the loop itself.