As a result we will probably see the pendulum swing back from cache-heavy designs. Going ham on cache was an obviously advantageous strategy at 7nm, you can basically look at it as N7 having been two full node families ahead of the curve on sram density (Samsung is only catching up at 3nm/2nm). But since SRAM hasn’t scaled at all in the last 2 nodes it’s becoming comparatively more effective to spend your transistors on logic instead. You still want big caches of course (and cache can be easily stacked ala AMDs v-cache) but it’s more worthwhile to spend more heavily on logic than it was on 7nm.
N7 basically was the era of "let's throw cache on everything, even products that traditionally haven't had caches". GPUs never had L3 cache before, for example, but that became advantageous, even in GPUs, which are focused around logic/computation rather than deep cache structures. And now you are seeing that train grind to a halt - RDNA3 did not expand the cache further, although they did increase the bandwidth of the cache (2.7x higher, although bear in mind that RDNA3's memory subsystem is 50% wider which means cache bandwidth is effectively 1.8x higher in a relative sense).
Similarly Intel went completely nuts once they finally got to a 7nm-tier node (Intel 7 aka 10ESF). Raptor Lake in particular is just caches all the way down...
Since cache no longer shrinks at 5nm and 3nm, but logic does, it makes sense to do some logic-intensive things rather than just throwing all your area at cache like 7nm.
On the flip side though, since cache can be pulled out to a separate cache die fairly effectively (AMD v-cache), you can continue to scale cache there. RDNA4 and NVIDIA Blackwell are both rumored to be coming fairly quickly which suggests a potential respin. And both AMD and NVIDIA have things they need to work on, NVIDIA doesn't have DP2.0 and AMD seems to have screwed up RDNA3 fairly badly, so a "similar but improved" quick refresh makes sense. Rumors mentioned "[NVIDIA] Ada Lovelace with a stacked cache die" at once point and imo that is very plausible for the next-gen Blackwell chips as a potential quick-fix improvement to keep scaling Ada.
But I think we're going to see an overall trend towards "the logic die is for logic and L3 cache gets stacked on top" for now. That seems to be a formula that works without too much trouble. L1 and L2 on the logic die is unavoidable, there is too much incentive for proximity/latency improvements, but big stacked L3s seem very effective and doesn't cause MCM-style problems.
Going forward there will also be non-cache (and non-memory!) things bonded as well.