Intel’s Long Awaited Return to the Memory Business
realworldtech.com
realworldtech.com
This article is about the same thing, presumably:
http://www.anandtech.com/show/6911/intels-return-to-the-dram...
They must be serving this from a Raspberry Pi.
What they can do though is achieve more aggregate bandwidth without increasing package pin count. Perhaps this explains why Intel was willing to allow Ivy Bridge to have only half the memory bandwidth of the previous Sandy Bridge: just not that many (non-server) systems were ending up with the four DIMMs required for full bandwidth.
So now they're integrating it into the package itself.
This is not some fundamental law; it's entirely dependent on bitline and wordline capacitance (array size). For external DRAM chips, huge arrays make sense, and you end up with tens of ns of latency. In an eDRAM cache scenario, you use much smaller arrays, and in the case of the POWER7 L3, latency is more like 6ns (for row hits, of course).
From early 1980's to early 2000's, DRAM latency went from 150 to 50 ns and has stayed around there since. That's 1-2 doublings of performance in that parameter in the last 30 years, compared to who-knows how many in transistor size and speed.
Still, Intel knows what they're doing and I can see a place for (yet another) spot in the memory hierarchy before the CPU resort to off-package DRAM.
There's a huge tradeoff space between latency, bandwidth, capacity, and cost, and market forces have forced a convergence around 2 design points, colloquially DDR and GDDR. For yield reasons, die area (cost) has been mostly fixed for a long time. Successive generations of DDR spend the dividends of Moore's Law primarily on extra capacity, then on a bit of additional bandwidth when possible. Successive of generations of GDDR prioritize bandwidth, primarily by dedicating tons of area to high-speed single-ended I/Os.
These two design points make sense for their most common use cases. In traditional disk-based systems, avoiding hitting the disk is more important than absolute DRAM latency, so increasing capacity is your best bet. On GPUs, you need enough bandwidth to feed a quickly growing number of functional units on the chip, and at least for graphics, the access pattern can be made to be extremely predictable, so latency is not as important there.
The advent of faster-than-HDD persistent storage (SSDs) and the desire to run more general purpose workloads on highly parallel machines like GPUs points to a need for a third DRAM design point.
Just think of this move as part of the development of those technologies. And the $1 per 256 MByte doesn't apply to on chip integrated memory with high bandwidth.
Maybe 3D isn't ready yet, but at least 2.5D integration is already in the market in FPGA's , and there's some claims from some companies[1] that it's ready for market for low cost applications.
[1]http://www.eetimes.com/electronics-products/electronic-produ...
I'm actually more excited that the CPU has access to this cache too.