The Memory Wall: Past, Present, and Future of DRAM
semianalysis.com
semianalysis.com
These days you can do thousands of calculations waiting for a few bytes of memory. And not only is the speed difference getting worse, but memory sizes aren't keeping up.
Guess we're not far away from compressing stuff before putting it in memory is something you'd want to do most of the time. LZ4 decompression[1] is already just a factor of a few away from memcpy speed.
The RAM was so terrible that essentially you try to keep the processors running in cache for as long as possible. RAM access is painful.
There is a performance profiling tool built into F3DEX3 that now shows that approximately 70% of the time the system is idle while running Zelda OoT. It is just waiting for memory transfers. The folks at SGI/RAMBUS cut corners a little too hard building that system.
But turns out this kind of performance profile is just prep for were we are heading apparently.
Especially when you throw multiprocessing in. We need better benchmarking tools that load up competing workloads in the background so you can tell how your optimization really works in production instead of in your little toy universe in the benchmark.
DRAM isn't getting (much) cheaper. And it will continue to be the case in next 10 years. Considering there is nothing on the roadmap to suggest any breakthrough.
People old enough may remember in the late 00s and even up to mid 10s, there are words like we will get 16GB computer as baseline, or I could work with 64GB soon.
Reality is that DRAM price has fluctuate lately within the same range in the past ~13 years. It was only in 2023 the price floor broke through the $2/GB barrier.
Unless we somehow found a way that could magically shrink the capacitor by a substantial amount. Or we change the way we do programming.
It is too early for me to imagine what capabilities are unlocked if I personally had say 1TB of memory, but large RAM capacities seem only required for mammoth database servers.
Given a regular ROM structure on the silicon wafer, it would be possible to take advantage of parallel beams and design an electron beam lithography machine specifically for ROM to reduce the cost of programming. The final frontier would be to build a "wafer scale engine"-esque single wafer chip and we would have essentially reached the limits of what is achievable through clever design of silicon based semiconductors.
https://news.skhynix.com/sk-hynix-develops-pim-next-generati...
It sort of makes sense because SRAM cannot have the same transient nature of a calculations transistor. It has to hold the state for longer than one 3 billionth of a second. So it has to be a little more robust.
This is my intuitive take so it could be completely wrong.
(Of course, this is still very irrelevant for comparison to DRAM; the price difference is huge.)