The author assumes an average of (slightly over) 2 instructions per iteration, which means the theoretical limit is a half iteration per clock, assuming no cache misses ever. How can you reduce that to 1 processing cycle per clock in a single thread, is there a load-add-store instruction, or a way to get multiple instructions per clock through a single thread?