The Future of Memory
semiengineering.com
semiengineering.com
We need more software engineering progress on paging & persistent storage systems.
So as you increase bandwidth you reduce program speed.
Higher frequency results in more heat.
We are fast approaching the need for a Wii like Broadway architecture where the program is running in "fast" SRAM and the data is on "slow" DDR.
Also most programs have cache-misses.
It's like the existing APIs for pining things in memory so they can't get paged out. They have very specific uses and normal programs generally don't use them and shouldn't.
That’s how GPUs are doing that. Each thread group can use a limited amount of SRAM, programs declare in advance how many bytes they need. Then in runtime the scheduler who dispatches tasks to cores enforces that limit by never dispatching too many thread groups on each core.
This can give better electrical characteristics of the bus, as the buffer chip to the DIMM connector can have simplified routing and higher power signaling without putting more load on the DRAM chips, and the buffer chip design being focused on this interface signaling rather than compromising between that and the actual DRAM cells.
It's a bit more expensive, being an extra chip on each DIMM, and has a latency penalty, as the buffer chip means everything on the DDR bus is effectively 1 clock behind what the DRAM chips themselves provide. But it's often necessary if you have a large number of DIMMs on a single channel or very long traces required for packing lots of DIMMs around a CPU, as that increases the electrical capacitance and noise of each path, which many DRAM chips can struggle to drive, especially at higher speeds.
As dram chip density increases you can get higher capacities without the longer bus traces and more DIMMs per channel that might require registered ram, there's nothing "fundamental" about 64gb needing registered ram, and you are already seeing 48gb DDR5 DIMMs that can work on consumer platforms, which often have no issues running 4 DIMMs without registered ram.
Can you get more specific then? Can you give any details or any overview? There must have been some information that led you to post this originally, can you link it?
Some of them include:
- GPU - CPU shared memory (already possible just not yet the standard)
- Higher DRAM bandwidth (already possible, just not yet a priority)
- system on chip FPGAs (always possible just very expensive to fit “AI models”)
- SOC NVM. Ideally even NVM on the same wafer as the GPU and CPU (possible today but requires a lot of work on the yield. NVM would take up a lot of real estate that could ruin yield).
- analog circuits
- new semi-conductors / photonics
- memristors