I'm having a hard time mentally come up with a way a larger L1$ could be faster. Have more ways, sure. Or more read ports. Or more bandwidth. And I'm given to understand that Intel tags (or used to tag) instruction boundaries in their L1I$. But how do you reduce latency? Physically larger transistors can more quickly overcome the capacitance of the lines their destination and the capacitance of their destinations but a large cache makes the distance and line capacitance correspondingly larger. You can have speed differences with 6T versus 8T SRAM cell designs but as I understand it Intel went to use 8T everywhere in Nahelem for energy efficiency reasons. I guess maybe changes in transistor technology could have made them revisit that, but 8 transistors isn't that much physically larger than 6.
But in general there are a lot of complicated things about how the memory subsystem of a CPU work that are important to performance, add add a floor to how low a first level cache's latency can be, but they don't really contradict anything that was said in the article.
It’s faster because you’re comparing a two completely different physical implementations. SRAM does not use capacitors. DRAM does. You trade speed for density.
I always thought cache layers was because of locality but that is my imagination :) The article talks about different access patterns of the cache layers which makes sense in my mind.
It also mentions density briefly:
> Only what misses L1 needs to be passed on to higher cache levels, which don’t need to be nearly as fast, nor have as much bandwidth. They can worry more about power efficiency and density instead.
The space constraints are also caused by money. The reason we don't just add more L1 cache is that it would take up a lot of space, necessitating larger dies, which lowers yields and significantly increases the cost per chip.
Space constraints are caused by power and latency limits even with infinite money.
> added levels of muxing necessarily mean that there's a limit to how low the latency of a cache can be
L1 cache avoids muxing as much as possible, which is why it takes up so much die space in the first place.
> L1 cache avoids muxing as much as possible, which is why it takes up so much die space in the first place.
Every time you double the size of a cache, you need to add a single extra mux on the access path. Simply to be able to select from which half of the cache you want the result. You also increase the distance that a signal needs to propagate, but I believe for L1 the muxes dominate.