Why aren't there e.g. 16 GB of L1D$ inside our CPUs? It'd be blazingly fast! :)
Although memory access patterns of typical GPU workloads are very different than those of CPU workloads, there are GPU-local, hierarchical caches similar to CPUs for different purposes that are exclusive to shader cores/computation units, which can be even partitioned dynamically. The struggle on the GPU side is keeping all the cores busy while servicing memory requests efficiently. Sharing memory across small group of threads run in lockstep, for example, is a great way of doing it.