* Late edit up front to add: you are of course absolutely correct that the ram perf has room, this is demonstrated by the article’s speedups. If prefetching works, it means the baseline wasn’t bottlenecked on peak ram bandwidth.
The article was assuming that the peak theoretical throughput is instruction limited, not ram limited. Then it demonstrated empirically that ram was somehow slowing it down further from the theoretical peak. Providing a counter example might be better than arguing over various assumptions?
We could elaborate with some specifics on the ram speed you’re thinking of? Which kind are we talking about? What is your calculation of the bandwidth of this problem? I assume it’s 8 bytes per iteration (4/read, 4/write), so for a typical processor at, say 3.5Ghz, and one instruction per clock, that would require a ram bandwidth of 28GB/s. That’s higher than most single channel DDR4, right? I think most high end desktops are dual channel and higher ram bandwidth, but I don’t know what the bandwidth is in practice when you plow through memory linearly with minimal compute and no prefetching, the perf in the article doesn’t strike me as being strange.
> The cache can do a read and a write per clock, so should be good for 1 iteration per clock.
It seems like the data in the article mostly supports that statement, with the caveat that an occasional cache miss will bring the average down a little, and the author claims that switching between read and write every single word is slower than reading large blocks.
> I think what may have happened is that the compiler did exactly what it was instructed to do, read the value that it just wrote to memory.
I’m not sure I understand what you mean. The code doesn’t read the value after it writes, it writes over the last value read, and then moves to the next address, right? No need to speculate about the compiler, the author included x86 assembly, right? Are you suggesting the hardware might be treating the value as volatile and skipping the cache, and stalling a read? I think that would be a lot slower. Or do you mean something else?