> I've heard that memory is a bottleneck for the GPU, specifically the time it takes to move stuff from RAM (main memory) to the GPU's RAM/memory. If it is such a big bottleneck, then why don't we (yet) see powerful GPU sharing the same die as the CPU and accessing/using/sharing RAM with the CPU (like integrated GPUs do)? Then there'd be a zero bottleneck. You would just give load whatever into RAM, and just give the (super-powerful integrated) GPU a pointer/address. Bam, done. Why hasn't this happened yet? / What am I missing/misunderstanding here?
That is not the only bottleneck involved.
Historically, GPUs have used GDDR ram as opposed to general purpose DDR memory. One of the key differences between GDDR and DDR is the bus width, which can be as large as 1024 bits, compared to conventional ram with a 64 bit bus width (although dual channel is effectively 128 bits). This much wider bus results in much higher memory bandwidth which is generally necessary to feed the truly enormous number of functional units in a GPU.
I suppose you could ask: why doesn't everyone just standardize on GDDR?
1. This would dramatically increase cache line size. I don't have data, but I assume this would generally be bad.
2. My recollection (but I don't have a source for this) is that DDR has lower latency than GDDR ram, so for branchy code (which CPUs often have to deal with, but GPUs typically never have to deal with), DDR could actually be faster.
3. DDR is cheaper to manufacture. Aside from being higher volume, a lower bus width just makes is simpler to manufacture.