Narrators voice: It was always the ultimate spec, it is now becoming more obvious to more people.
One of the interesting things for me is how memory "operations" (which is to say transactions of the memory controller) can completely obliterate "bandwidth."
In the past (not sure how true this is on current microarchitectures) the controller between cache and DRAM worked by opening a "page" of dynamic memory. That was an artifact of how DRAMs use both a column address select and a row address select and then multiplex the address bits. So to "read" memory you needed to select the column, then select the row, and then read the memory. The good news was that if you needed the next word of memory in order, you could just read again. And again. Periodically if you were reading the chip would have to ask you to wait while it refreshed its contents.
Anyway, this memory operation of opening a new page was a lot slower than reading the next word in memory. So if you're requests were bouncing all around memory your effective bandwidth was limited by how fast your memory controller could open new pages. That could be one tenth the nominal serial access bandwidth. Controllers had multiple "page" registers so they could hold the state of two (or more) different DIMMS and try to interleave their access across DIMMs to hide the latency aspect of page mechanics.
Generally though, when you get to the point where your measuring memops and trying to layout your physical memory to minimize them you're in a different realm of system optimization.