DDR5 Memory Specification Released
anandtech.com
anandtech.com
DRAM runs on a separate process which is dominated by the difficulty of building the capacitors. These are roughly the shape of a pencil (long narrow hexagons) where the central structure which holds the capacitor needs to be etched to perfection in a process that can take days. The transistors underneath are, at that scale, about as large as the chad from a paper hole punch. The capacitors are just about as narrow as material science (limit to voltage arcing through the insulation layers) can make them so there is glacially slow progress in shrinking DRAM further. Meanwhile the transistors are at extreme limits of resolution for liquid immersion processing, as also are the lines needed to join the rows and columns. Getting those perfect requires very specialized and competent processing.
They are not easy, second rate circuits. They are a completely separate branch of the silicon world. Unfortunately since they don't scale much any more, current design methods were mature 8 years ago, the only way you get more of them is to build new factories. That means it is a seller's market in a game where building another fab costs $10B and will only succeed if staffed by really expert people. So, it is generally profitable. The 3 vendors cannot easily undercut each other since they all have roughly the same limits, and any attempt to flood the market takes 4 years to build and everyone can see it coming.
So there you are. DRAM is the pivotal technology of the current computer era. Fixing that will most likely require breakthroughs in fundamental memory technology - or a reason for demand to collapse.
So on 8 Channel 16 DIMM per socket you could fit a theoretical 32TB of memory. This is insane amount of memory and great for In-Memory Database. ( How is Intel Optane going to compete? )
This makes me wonder, what makes DRAM so expensive? It is still hovering at a median price or around $3/GB compared to NAND which is less than $0.1/GB.
Greed does. DRAM makers were antitrust busted at least 7 times on my memory in Taiwan, Korea, and USA.
They are not necessarily price fixing illegally. It's just that they all keep their production and capacity expansions closely in check to not ever let prices down.
That and in rare cases when prices are down due to unexpected decline in shipments they're all very swift to shift wafers to produce something else. Feel free to dig DRAMeXchange reports for for details.
I'm not sure how they have survived bad press, but Samsung is not a good company.
Sure, you have to sacrifice more transistors for the same capacity, but newer processes can fit more on the chip, right? I recall from computer architecture classes that the benefit of DRAM is the ability to use fewer transistors, but if transistors are cheap...
(I'm sure I'm missing something here. Power consumption / heat generation? I also never really understood why SRAM continues to be so expensive, when it seems like it would obviously benefit from smaller processes.)
So even if you were to manufacture SRAM on Intel's ultra-expensive 10 nm logic process, you'd need a massive amount of silicon for the same capacity.
[1] https://fuse.wikichip.org/wp-content/uploads/2017/12/isscc-2... [2] https://d3i71xaburhd42.cloudfront.net/f20203949a744276e338d6... [3] https://www.anandtech.com/show/13999/sk-hynix-details-its-dd...
But if you have issues scaling DRAM, and different scaling limits on transistor count / SRAM, it makes sense (to me at least) to start considering SRAM as an option (e.g. for lower latency, faster speeds, higher bandwidth transfers, etc). Just because you can't achieve the same capacity today doesn't mean there's no merit to it -- HDDs vs SSDs from a decade ago feels like the obvious comparison.
Supposedly [1] TSMC's 5nm process yields 256Mb on a 5.376mm² die, at roughly ~50Mb/mm², which would translate to a 3.5Gb die of the same size as the SK Hynix chip. Sure, that's no 16Gb die, but you could easily make 32GB sticks (assuming that you could just combine these chips in the same way as in DDR4).
I guess there's also a barrier to entry in that you'd also either need new hardware to deal with "SRAM sticks", or some sort of compatibility layer (a controller that implements the DDRx signaling logic, perhaps).
[1] https://www.anandtech.com/show/15219/early-tsmc-5nm-test-chi...
Of course that all depends on the generation of tech and only applies in an apples to apples scenario.
It's really not, especially in the 3D NAND flash era where only one manufacturer is still using a floating gate cell. It's so thoroughly not transistor based that the Chinese upstart's claim to fame is that they fabricate the transistors on an entirely different wafer from the memory cells, and glue them together later.
It's best to think of NAND, DRAM and logic as three separate categories that each require a very different mix of tools in the fab, especially on the back-end. (But you won't be finding EUV or quad-patterning in the front-end of a NAND fab, either.)
Seems like RAM manufacturing is like building a casino: expensive to do, but basically a sure bet.
You could theoretically spend several times more than that to try to get ahead over the course of two or three generations, but for that kind of money you could just as easily secure some very preferential pricing from one or more of the incumbents, thereby ensuring that everyone else trying to put a lot of RAM into a PC has to pay more.
The most viable path to establishing a new leading-edge competitor in this space is for a government to throw lots of money at the problem, knowing that it'll be years at best before it produces anything competitive, but having the advantage of being able to more or less ignore IP issues and having a potential demand far higher than any one memory customer can produce on its own. China is doing this for the NAND market, too.
[0] https://www.finanznachrichten.de/nachrichten-2010-05/1694023...
that being said, I do admit my post was sarcastic, but it also did mention what the problem here is (besides a much broader discussion about captialism) which is that they can rationally expect to get away with this. State-sponsored ones because their interests are the national interest (to an extent) and privately held ones because they're usually able to effectively capture their regulators.
Also worth considering that 32TB of DRAM would draw over 12kW, just sitting there.
That's the virtual address space. A page table entry has enough bits to have a 64-bit physical address space, it just wouldn't be able to have it all mapped at once in the same virtual address space. Although CPUs don't have 64 physical address lines yet, there's nothing fundamental in the x86 architecture preventing them from doing so.
Intel have already implemented 5-level paging, which would give a 2^57 bit virtual address.
Also, this wouldn't be the first time that the x86 has supported more physical memory than virtual. PAE allowed for 64GB of RAM on a system with a 32-bit virtual address space.
While we're here I hope anyone can explain to me why Cooper Lake Xeon parts are listed as supporting 4.5TB of DDR4.
Answering self: the 4.5TB support is mostly Optane memory, not DDR4.
So the number of bits in an n-level system is 9n + 12.
9x5 + 12 = 57
On 32-bit systems the number of PT entries was 2^10, which is why they could have a 2-level system. 10x2 + 12 = 32.
IBM Power System E980, max ram: 64TB.
Expanding from 48 toward 64 bits isn't difficult.
$ echo '(32 * 2^40) / (8 * 2^30) * 3' | bc
12288Here is the datasheet for a 128GB dimm from 2017 [1], which shows 3.4A IDD0 (normal operation) on the 1.2V rail at the highest speed of DDR4-2666, and 0.2A on the 2.5V precharge rail for a total of just over 6W. Also worth noting is that is a LRDIMM, which draws more power from the DC rails due to the additional buffering. A normal RDIMM draws a bit less static power.
Compare to a manual for a similar vintage 32GB stick [2], which consumes 2A on the 1.2V rail and 0.1A on the precharge rail for a total of a bit under 3w. One quarter the capacity, but still half of the power draw.
[1] https://www.samsung.com/semiconductor/global.semi/file/resou... [2]https://static6.arrow.com/aropdfconversion/d3b3ce1d78b0ad7d3...
Thanks for doing the math on the power story. I didn't realize about the scaling.
I thought RISC-V could with the RV128 ISA.
You should see how fast Optane will copy a project folder full of node_modules, bin, etc.. Blows the doors off an Evo Plus.
That said, probably not worth the premium for the vast majority of uses.
Not realistically; that would just introduce more problems than what I would avoid.
I have better things to do with my time than try to avoid npm.
JS projects are only a small fraction of my work anyway.
128gb dimms are more than double the price of 64gb dimms, so it not always economically viable unless you need max memory density.
Price ;=)
Also Optan drives (over PCI) are probably still interesting for some applications due to low latencies as far as I remember.
Look at base ram in MacBooks, the growth over time is pretty slow.
While both benefit greatly from economy of scale; the manufacturing tolerances, equipment, etc etc used influences pricing; but I'm not an engineer... so maybe someone here can chime in on that :)
The difference between 10µs and 60ns is merely a factor of 167 not 1 million.
The only take away I have from your comment is that you somehow confused DRAM with the L1 cache and SSDs with HDDs. That's the only way one could possibly arrive at your numbers.
I have used rounded numbers for illustrative purposes. They might be off by 30% or more but they are within the right order of magnitude.
This means we don't have to worry about ECC support by CPU/motherboard anymore, right?
> So on-die ECC is a bit of a mixed-blessing. To answer the big question in the gallery, on-die ECC is not a replacement for DIMM-wide ECC.
> On-die ECC is to improve the reliability of individual chips. Between the number of bits per chip getting quite high, and newer nodes getting successively harder to develop, the odds of a single-bit error is getting uncomfortably high. So on-die ECC is meant to counter that, by transparently dealing with single-bit errors.
> It's similar in concept to error correction on SSDs (NAND): the error rate is high enough that a modern TLC SSD without error correction would be unusable without it. Otherwise if your chips had to be perfect, these ultra-fine processes would never yield well enough to be usable.
> Consequently, DIMM-wide ECC will still be a thing. Which is why in the JEDEC diagram it shows an LRDIMM with 20 memory packages. That's 10 chips (2 ranks) per channel, with 5 chips per rank. The 5th chip is to provide ECC. Since the channel is narrower, you now need an extra memory chip for every 4 chips rather than every 8 like DDR4.
I'm looking forward to our ARM future. :-)
After all, DDR4 has higher latency than DDR3 running at the same clock speed.
If there are two 32-bit data busses rather than one 64-bit bus, arithmetic suggests they shouldn't need to find extra pins from somewhere.
So maybe the rationale for shrinking the CA busses (to 7 rather than 12) is something different?
DDR4 appears to have had 40 and 32 bit data buses, while this one has 40/40.
if so it seems like OS'es could track and publically tell on executables
Sure, any particular exploit may be less likely to work, but once you can hammer memory 3/4 of the code running on the system turns into a potential exploit vector. :)
How will this affect driver complexity and cache-misses?
As long as multiple cores are accessing memory or prefetching is on (it's almost always on), both channels will be utilized so software won't notice.
[1] When you do a read operation on DRAM you get a multi-cycle burst of data, not just one word. This amortizes command/address overhead and presumably matches the slow-but-wide internal DRAM array with the fast-but-narrow channel. See https://people.freebsd.org/~lstewart/articles/cpumemory.pdf sec. 2.2.
I am not a HW engineer, but:
With DDR, the difference of all traces in the same channel (data & clock) has very tight tolerances (on the order of 1/8 or 1/16 of a clock). Having fewer traces per channel may make it easier to route for higher clock speeds.
> How will this affect driver complexity and cache-misses?
I'm not sure what you mean? The memory controller should abstract almost all of the differences away. There are per-channel configuration settings that are usually configured by the SPD rom, so there will be twice as many to set, but multichannel memory controllers are already a thing, and going from N to 2N of something doesn't really affect software complexity once N is greater than one.
So you go from
REQUEST1---------RESPONSE1-REQUEST2---------RESPONSE2
to REQUEST1---------RESPONSE1RESPONSE1
REQUEST2---------RESPONSE2RESPONSE2
Each request individually is slightly slower, but the total bandwidth is greatly increased. REQUEST1-REQUEST2---RESPONSE1-RESPONSE2
can't it?Increasing the clock just makes it a lot harder for motherboard and CPU manufacturers to support those speeds, but on its own it doesn't really gain you a lot of speed.
Thanks for letting me know of your experience!
The article is actually explicit of this not being the case:
DDR5 DIMMs: Still 288 Pins, But Changed PinoutsThe real purpose of DDR is not actually to double the data rate, but to halve your clock speed and allow you to use the same frequency for your clock as your data. This mostly benefits signal integrity.
QDR is better understood as memory with two ports, one for reading and one for writing, which can be used at the same time. This is a lot more expensive and really doesn't have huge benefits for PCs compared to just adding more channels (as DDR5 does).
That said DDR5 is solidly on two transfers per clock, so I don't understand the suggestion to call it or use "QDR2" (what's QDR1 for contrast?).
CPUs have become so fast that relative to their "internal" speeds, RAM is the new hard disk. Databases are becoming in-memory, and going out to fixed storage, even SSD, is an anathema.
New applications are not designed to work on data sets bigger than physical memory. Disk-to-disk streaming algorithms are practically unheard of outside of a few niche scenarios. Like I said, even database vendors are moving to in-memory!
I love machines with huge amounts of memory. My laptop has 64 GB, and it's great! I can run entire fleets of servers in a local hypervisor. I can load huge blobs of CSV or JSON data into the shell and not have to worry about the 2-5x overhead of the in-memory representation. It'll fit just fine. I can run every "bloated" app at once and still have 50 GB free for "whatever". I've reindexed a database on my laptop in minutes that would have taken days(!) on a production server because it didn't have enough RAM and was thrashing the storage like crazy.
Another way to look at it is the "GB per CPU core". With existing AMD EPYC 2 CPUs having 64 cores and 128 threads, the typical 512 GB memory configuration is "only" 8 GB per core, or 4 GB per thread! With a dual-socket server, halve those numbers again. Similarly, mainstream desktop Ryzen CPUs have up to 16 cores, and that's not even talking about the not-so-mainstream Threadripper line. For 4GB per core, you'd need 64 GB.
It's likely that AMD will release 24 or 32 core mainstream CPUs in the near future, maybe as soon as 2 years from now when their 5nm products start shipping. I fully expect server CPUs to hit 96-128 cores per socket around the same time frame, or up to 512 hardware threads in a standard two-socket server. Terabytes of memory is going to become "standard" very soon now.
However, that does not address the sheer wastefulness of our technological trends to require more resources to do things slower, but displayed with smaller and more colorful pixels. Should everyone have a 64-core 512GB memory computer to view web pages, play minecraft or whatever? Will that be too small to write a text document in 20 years time? Will every person on the planet be expected to get a bigger computer because they can't run the (electron-in-ethereum-on-browser-in-container)^n pancomputer?
My servers could probably benefit from, idk, petabytes? If I could keep the entirety of my server's hard drive in RAM I'd be very happy.