So the only interleaved memory is very small caches local to the compute cores.
(And I'm not even getting into power and thermals involved)
Latest generation high end GPU's use HBM2(E?) memory, which is a very wide and fast pipe compared to DDR4 used for "normal" CPU main memory.
As for systolic arrays, to some extent the matrix-matrix units in Google TPU's and NVIDIA Tensor Cores are systolic arrays. I suspect we'll see designs go further down that path in the future.
Although memory access patterns of typical GPU workloads are very different than those of CPU workloads, there are GPU-local, hierarchical caches similar to CPUs for different purposes that are exclusive to shader cores/computation units, which can be even partitioned dynamically. The struggle on the GPU side is keeping all the cores busy while servicing memory requests efficiently. Sharing memory across small group of threads run in lockstep, for example, is a great way of doing it.
I didn't say the memory necessarily has to be moved into the GPU. The other way around would also be a possibility.