N7 basically was the era of "let's throw cache on everything, even products that traditionally haven't had caches". GPUs never had L3 cache before, for example, but that became advantageous, even in GPUs, which are focused around logic/computation rather than deep cache structures. And now you are seeing that train grind to a halt - RDNA3 did not expand the cache further, although they did increase the bandwidth of the cache (2.7x higher, although bear in mind that RDNA3's memory subsystem is 50% wider which means cache bandwidth is effectively 1.8x higher in a relative sense).
Similarly Intel went completely nuts once they finally got to a 7nm-tier node (Intel 7 aka 10ESF). Raptor Lake in particular is just caches all the way down...
Since cache no longer shrinks at 5nm and 3nm, but logic does, it makes sense to do some logic-intensive things rather than just throwing all your area at cache like 7nm.
On the flip side though, since cache can be pulled out to a separate cache die fairly effectively (AMD v-cache), you can continue to scale cache there. RDNA4 and NVIDIA Blackwell are both rumored to be coming fairly quickly which suggests a potential respin. And both AMD and NVIDIA have things they need to work on, NVIDIA doesn't have DP2.0 and AMD seems to have screwed up RDNA3 fairly badly, so a "similar but improved" quick refresh makes sense. Rumors mentioned "[NVIDIA] Ada Lovelace with a stacked cache die" at once point and imo that is very plausible for the next-gen Blackwell chips as a potential quick-fix improvement to keep scaling Ada.
But I think we're going to see an overall trend towards "the logic die is for logic and L3 cache gets stacked on top" for now. That seems to be a formula that works without too much trouble. L1 and L2 on the logic die is unavoidable, there is too much incentive for proximity/latency improvements, but big stacked L3s seem very effective and doesn't cause MCM-style problems.
Going forward there will also be non-cache (and non-memory!) things bonded as well.