Why do CPUs have multiple cache levels? (2016)
fgiesen.wordpress.com
fgiesen.wordpress.com
It's also why instruction and data caches are usually split at the L1 level. It's a bit more complex to keep them coherent but it almost doubles the size. Also these days it allows icache to store pre-decoded instructions, but that's secondary.
1) limiting pollution, so memcpy doesn't purge L1 of instructions 2) simplifies interface between CPU and icache - only need to support a handful of instruction fetch block sizes CPU actually uses
TL;DR: Time complexity of memory access scales O(√N), not O(1), where N is your data set size. This applies even without hitting theoretical information density limits, whatever your best process size and latency. For optimal performance on any particular task you'll always need to consider locality of your working data set, i.e. memory hierarchy, caching, etc. And to be clear, the implication is that temporal and spatial locality are fundamentally related.
Yes, some cache levels have lower intrinsic latency than others, but at scale (even at-home, DIY, toy project scales these days) it's simpler to think of that as a consequence of space-time physics, not a property of manufacturing processes, packaging techniques, or algorithms. This is liberating in the sense that you can predict and design against both today's and tomorrow's systems without fixating on specific technology or benchmarks.
No, it doesn't. Due to the Bekenstein bound, the amount of matter — and hence information — that can be stored in a sphere is ultimately proportional to surface area, not to volume. This is covered in part 2 of the article: https://www.ilikebigbits.com/2014_04_28_myth_of_ram_2.html
I guess that was covered in the comment.
> This applies even without hitting theoretical information density limits, whatever your best process size and latency.
But now I guess I don't understand why... :(. I feel like we can almost trivially show this isn't true by starting with a downright primitive "best process size"--maybe a system where every bit is so large I can see it and access it using a pair of tweezers--but maybe I am fooling myself?
I'm having a hard time mentally come up with a way a larger L1$ could be faster. Have more ways, sure. Or more read ports. Or more bandwidth. And I'm given to understand that Intel tags (or used to tag) instruction boundaries in their L1I$. But how do you reduce latency? Physically larger transistors can more quickly overcome the capacitance of the lines their destination and the capacitance of their destinations but a large cache makes the distance and line capacitance correspondingly larger. You can have speed differences with 6T versus 8T SRAM cell designs but as I understand it Intel went to use 8T everywhere in Nahelem for energy efficiency reasons. I guess maybe changes in transistor technology could have made them revisit that, but 8 transistors isn't that much physically larger than 6.
But in general there are a lot of complicated things about how the memory subsystem of a CPU work that are important to performance, add add a floor to how low a first level cache's latency can be, but they don't really contradict anything that was said in the article.
It’s faster because you’re comparing a two completely different physical implementations. SRAM does not use capacitors. DRAM does. You trade speed for density.
I always thought cache layers was because of locality but that is my imagination :) The article talks about different access patterns of the cache layers which makes sense in my mind.
It also mentions density briefly:
> Only what misses L1 needs to be passed on to higher cache levels, which don’t need to be nearly as fast, nor have as much bandwidth. They can worry more about power efficiency and density instead.
The space constraints are also caused by money. The reason we don't just add more L1 cache is that it would take up a lot of space, necessitating larger dies, which lowers yields and significantly increases the cost per chip.
Space constraints are caused by power and latency limits even with infinite money.
> added levels of muxing necessarily mean that there's a limit to how low the latency of a cache can be
L1 cache avoids muxing as much as possible, which is why it takes up so much die space in the first place.
> L1 cache avoids muxing as much as possible, which is why it takes up so much die space in the first place.
Every time you double the size of a cache, you need to add a single extra mux on the access path. Simply to be able to select from which half of the cache you want the result. You also increase the distance that a signal needs to propagate, but I believe for L1 the muxes dominate.
Cost of bits at L2 and L3 are more or less the same. They are typically implemented using the same kind of transistors.
The reason is latency, plain and simple. Making a pool of memory larger always makes it slower to access, so to have a very fast pool you need it to be very small. The optimal solution is a hierarchical stack of pools of increasing size and latency. And if you had infinite free transistors, it would still be the same.
| Memory _| Size __| Latency| Bandwidth
| L1 cache | 32 KB _| 1 ns _| 1 TB/s
| L2 cache | 256 KB | 4 ns _| 1 TB/s
| L3 cache | 8 MB+ | 10x L2 | >400 GB/s
| MCDRAM | ______| 2x L3 | 400 GB/s
| DDR DIMMs| 4 GB-1 TB| 2x L3 | 100 GB/s
[1] https://www.intel.com/content/www/us/en/developer/articles/t...
0. You're not going to have 32 GiB of SRAM anytime soon.
1. Amount of L3 translates into area which means cost. 1152 MiB of L3 on an 9684X ain't fucking free.
And cost plays no role in those decisions. Making the L1 or L2 larger would make them too slow. Only the size of L3 is limited by cost.
To mask this, we write back to cache and rely on cache coherency algorithms and multiway, multilevel caches to make sure main memory is written back to and read when cache tags are invalidated.
tl;dr - Current process technologies make SRAM very much faster than DRAM and multiple levels of multiway caches create a time based interface to maximise memory throughput to the CPU regsisters while maintaining coherent memory write backs.
It’s worth noting that Apple Silicon is fast because their DRAM bandwidth is much closer to the same machine cycle latency as the APU cores’caches and registers.