We've kind of half-assed it with DDR memory banks, but it mostly introduces mysterious slowdowns that are difficult to reason about and I think we would be better served I think by making a formal thing. Instead of introducing an L4 cache we could do this instead, and reduce the size of the L1-L3 caches, which shortens lookup time and thus latency.
For legacy apps, you could provide facilities for the OS to 'page' blocks in from main memory, but the speed would come from managing the workload imperatively, starting loads in the background before the data is actually needed, and dumps after it is last touched.
You could also take a half-step in that direction with the DEC Alpha processor's extremely relaxed memory consistency model, and see how well things program for that. When I had a chance to poke at them, Alphas were about 2-3x faster than Intels, with much less ecosystem optimization effort burned -- and most software worked just fine.
You're only restricted by the fragmentation of the system memory which is an issue yes, but it's dealt with in other ways.