I created a huge array of 32 bit ints - a few gigabytes in size. Then timed this code:
uint32_t checksum = 0;
for (int i = 0; i < 1000000; i++) {
unsigned offset = rand32() & (TABLE_SIZE - 1);
checksum += table_ints[offset];
}
I was surprised to see it only took about 10 ns per iteration, which is 5 to 10 times faster than a DRAM access. Because the table was large and the access pattern was random, each iteration has to do a DRAM access (there is low probability of getting a cache hit).The processor was able to execute 5 to 10 iterations of that loop in parallel in a single thread. Quite amazing.
The random number generator looked like this:
uint32_t g_seed = 12345;
uint32_t rand32() {
g_seed = 214013 * g_seed + 2531011;
return g_seed;
}