The constraining factor in this case is CPU throughput.
When you looked at it with VTune, did you find clear evidence of this? That is, were you consistently executing 4 instructions per cycle, or were you maxed out at 100 percent utilization for a particular execution port? While it's possible to be at these limits (say if you were hashing each input) in most cases I think it's rare to get to that point when reading from RAM.
Similarly, by accessing data sequentially we make it possible for the CPU to predict what data will be accessed in the future, and have it loaded into the cache before we even request it.
One important caveat to this (at least on Intel) is that the hardware prefetcher does not cross 4KB page boundaries. So if you are doing an 8B read every cycle, and if you are able to do your other processing in the other 3 µops available that cycle, it's taking you ~500 cycles to process each page. And then you hit a ~200 cycle penalty when you finish each page, which means you are taking a ~25% performance hit because of memory latency, and probably have a CPI closer to 3 rather than 4! So if you (or your compiler) haven't already done so, you might see a useful performance boost by adding in explicit prefetch statements (one per cacheline) for the next page ahead.
Anyway, thanks for your response. I like your approach, and I wish you luck in making it even faster!