So, now it has jumped from (disk -> network -> memory -> ...) to (network -> disk -> memory -> ...), which is a big change.
I've definitely noticed that when optimizing a query and deciding between a high number of seeks vs. a table scan, older versions of MSSQL will tend to be pessimistic of drive latencies and just go with the full scan (potentially incorrectly / prematurely). In an uncached scenario on an SSD, this is probably sub-optimal. My guess would be that instead of looking at actual seek latency, the optimizer was using reasonable guesses for spinning disks. I'm guessing newer versions are more SSD aware though.
It's more and more like network/disk -> L3 cache -> L2 cache. DRAM is pretty slow.
Because PCIe controller is anyways on the same chip as L3 cache, there's no reason to send the data on a long trip to DRAM and back. Until, of course, when the cache line gets evicted for reason or another.