My conviction inspired me to repeat Richard's test to see how the numbers compared when run on a current Haswell i7-4770 processor. Here's what I got for the versions he compared when compiled with GCC 4.8.1 and using 'perf stat speedtest1 --size 5'[1] (sizes are for the sqlite.o lib only):
Cachegrind per Richard:
3.8.7a -Os: 953,861,485 cycles
3.7.17 -Os: 1,432,835,574 cycles
Haswell:
3.8.7 -O3: 469,956,519 cycles (870,767,673 instructions, 1,069,032 bytes)
3.8.7 -Os: 542,228,740 cycles (884,541,860 instructions, 648,360 bytes)
3.17.7 -O3: 651,680,545 cycles (1,320,867,830 instructions, 1,002,192 bytes)
3.17.7 -Os: 771,304,471 cycles (1,406,527,795 instructions, 605,976 bytes)
One can make of the numbers what you will, but here are some conclusions I drew:While the improvement is about the same magnitude that Richard sees, the actual number of cycles is off by about a factor of 2.
Despite having a larger binary, and thus theoretically worse instruction cache behavior, -O3 beats -Os by about 20% in speed in both cases, at a cost of 40% (400K/1MB) in size.
'perf record -F100000 speedtest1 --size 5; perf report' is a really slick and easy way to figure out where the program is spending it's time. If I were optimizing this, I remain fairly certain it would be more a more effective approach than using Cachegrind.
[1] I had to comment out the "sqlite3_rtree_geometry_callback" line in speedtest1.c to get it to compile, which might affect my numbers slightly, but from the source comments I don't think it was actually being used.