AMD's high-bandwidth memory explained
techreport.com
techreport.com
That said, I'd expect this to become the memory for laptops in the not to distant future. A Core i7 with 16GB of DDR5 stacked on top of the CPU and all that freed up space for more battery. Look for it in a Macbook near you :-)
Also with the lower speeds used, heat will be less and may be the limit currently to avoid warping of individual wafers in the stack.
As for desktops etc, one avenue this does open up would be a more standard socket perhaps as the CPU can change and sits on another socket in effect, allowing changes to be done at that level to maybe stretch out sockets some CPU's would normally never reach.
Does this count as a 3DIC?
Also, obligatory: http://i.imgur.com/v7loIkF.jpg
I've seen GPUs advertised as having 128-bit buses, but I didn't know GDDR5 was only 32bit. For this it feels a bit like the RAMBUS vs DDR battle all over again; high-clock combined with a narrow bus caused higher latencies than a slower clock with a wider bus, which has the added benefit of being cheaper overall.
This is a multi-chip module or MCM (http://en.wikipedia.org/wiki/Multi-chip_module). Intel also uses them to put the North bridge on the same package as the CPU for its mobile CPU's.
There are several innovative aspects about this MCM, but MCM's have been around for a long time.
A single GDDR5 chip is 32bit. GPU's use many chips in parallel to achieve wider buses.
The questions are whether G/CPUs have:
- multiple or single SKUs for on-package RAM capacities
- addon RAM ability via sticks, sockets and/or surface-mount
It would come as no surprise if addon RAM became a feature of the higher end CPUs (e.g. i7, Xeon, Phenom, whatever)
The biggest problem after this is how much the GPU and CPU have to fight over main system memory, which really brings us to the end game of GPUs altogether. Sooner or later there won't be room enough for both in the picture, your single heterogeneous core or MCM will have both a CPU and GPU on it (and probably a half a dozen or more application-specific accelerators).
If you need a point of reference take a look at L2 cache. It was originally chips on the motherboard, then with the PII it was put on a Processor Card, then with the PIII it was eventually integrated into the Die.
The exact same thing happened with L3 cache.
We are going to have AMD HBM [1], NVIDIA stacked DRAM [2] and Hybrid Memory Cube [3]. But why do we need all three, when the latter is supposed to be a standard? Or are some of these actually duplicates?
[1] https://en.wikipedia.org/wiki/High_Bandwidth_Memory
A more interesting configuration would be attaching 12 1GB HBM chips to the gpu , and achieving a memory bandwidth of 128GB/s * 12 = 1.5Tbyte/sec (would increase power only by 30W over current model).
Maybe their gpu is too weak to support such massive memory bandwidth, and it would be quite hard to do so ?
One problem with HBM is especially an issue for large GPUs. High-end graphics chips have, in the past, pushed the boundaries of possible chip sizes right up to the edges of the reticle used in photolithography. Since HBM requires an interposer chip that's larger than the GPU alone, it could impose a size limitation on graphics processors. When asked about this issue, Macri noted that the fabrication of larger-than-reticle interposers might be possible using multiple exposures, but he acknowledged that doing so could become cost-prohibitive.
Fiji will more than likely sit on a single-exposure-sized interposer, and it will probably pack a rich complement of GPU logic given the die size savings HBM offers. Still, with HBM, the size limits are not what they once were.