> Because the point of this experimentation was to iterate over a list by accessing the memory and so to measure the time for each iteration, I had to be sure every executed operation at each iteration was the most stripped-down set of instructions. Hard to be shorter than one cpu instruction for the given loop:
I’ll wager that the author didn’t look at the generated assembly as if that loop has only one CPU instruction, that instruction would be an increment, making the analysis of cache effects moot.