But they are starved of memory bandwidth! And the lower latency memory CPUs prefer is not the same as the high bandwidth memory that GPUs like.
Also there is different kind of caching. Modern APUs often have two ways to access memory. Over there own cache, or over the CPUs cache. So shared memory for CPU<->GPU gets the full advantage of a cache, but still it is a trade off.
If you want some work done by the GPU part of an APU, by sharing a pointer, you can do that today. But from the point of view of the CPU there is no prediction beyond the "GPU do X" commands. And a very high latency until the job is done. So you need a minimum GPU job size for it to make sense.
That is not the only bottleneck involved.
Historically, GPUs have used GDDR ram as opposed to general purpose DDR memory. One of the key differences between GDDR and DDR is the bus width, which can be as large as 1024 bits, compared to conventional ram with a 64 bit bus width (although dual channel is effectively 128 bits). This much wider bus results in much higher memory bandwidth which is generally necessary to feed the truly enormous number of functional units in a GPU.
I suppose you could ask: why doesn't everyone just standardize on GDDR?
1. This would dramatically increase cache line size. I don't have data, but I assume this would generally be bad.
2. My recollection (but I don't have a source for this) is that DDR has lower latency than GDDR ram, so for branchy code (which CPUs often have to deal with, but GPUs typically never have to deal with), DDR could actually be faster.
3. DDR is cheaper to manufacture. Aside from being higher volume, a lower bus width just makes is simpler to manufacture.
Why would it change cache line size? GPUs also use cache lines in the range of 32-128 byte?! I think that is independent of the bus system/width.
I just assumed that bus width = cache line size. I guess I was wrong.
Sorry.
Does this mean there are over a thousand traces between the GPU chip and the memory chips? It would be pretty clear why regular motherboards don't use it if that's the case, the sockets for the chips would be enormous! You're talking about roughly doubling the pincount vs. a 64 bit memory bus on a modern LGA socket.
The types of workloads run on GPUs typically like very high memory bandwidth and are usually willing to live with higher memory latency to get it. Onboard GPU memory is usually built with this in mind (trade off capacity and latency for increases in bandwidth). This is generally speaking the opposite of what you want in a CPU where people often want very high memory capacity and lower latencies, but may not limited by memory bandwidth, so simply sticking a GPU on die and giving it access to a memory subsystem that was not designed to feed a GPU is not going to make anything better.
Memory bandwidth is typically the bottleneck in GPUs. Meanwhile, the access patterns are typically very predictable. So they're able to prefetch data, so latency is generally not a problem. So GPU memory is designed to have very high bandwidth, even if it means completely tanking the latency.
On the other hand, CPUs typically always need better memory latency, and most workloads do not saturate the memory bandwidth. Unlike in a GPU, many memory access patterns are unpredictable. The bottleneck of operations that are pointer heavy tends to be limited by latency. All tree operations and linked list iteration tend to be latency limited. Languages like C#, Java, Python and Javascript where all data lives behind a pointer tend to benefit significantly from improving latency. So while improvement memory bandwidth up to a point is important, there's much more attention given to latency.
GDDR RAM wants to be accessed in bulk - large rows at once. Things are easy when each thread wants a subsequent byte, but if not things become much slower. Caching and other techniques can help mitigate this, and they're (imo) the place for a lot of creativity in architecture design. Having more potential FLOPS just means more ALUs.
Thank you for explaining it so clearly.
Some (not all) GPU memory types are not cache coherent with the CPU. Some of the cache-coherent cases have poorer performance relative to the non-coherent memory from the perspective of the GPU's memory bandwidth and access latency.
https://fuse.wikichip.org/news/1634/hot-chips-30-intel-kaby-...