Micron Introduces 128 GB DDR5-8000 RDIMMs with Monolithic 32 GB Die
anandtech.com
anandtech.com
Or even better "128 GByte ... 32 Gbit"
I know LLMs are undergoing rapid change, but what if someone (say Apple) wanted to commit a large model to an actual masked ROM? What kind of density can be achieved on a similar process as DRAM, and what kind of interface would such a thing use?
But you also have input dependent activations. You need RAM for these.
Then there's the other die, as in die-cast. I think these are typically referred to as dies in the plural. I'm not sure if this terminology is still in use, it seems to have been replaced with mold.
Then the other other die, referring to the fragment of silicon inside an integrated circuit. I've seen dice used more commonly, but both forms are in active use.
It's one of those frustrating inconsistencies of the English language.
I think die is still used when referring to metal stamping equipment and similar, but casting equipment is usually referred to with mold.
Don't they sound the same?
dies daɪz
dice daɪs
I never asked, but I want to believe their internal representation of “1D6” is one-dice-six.
Sorta like reading “1 x 6” as “one _times_ six” instead of “one _time_ six”
Words are fun.
https://journals.plos.org/plosone/article?id=10.1371/journal...
I can't see any reason to think that native speakers are changing their patterns much because of what non-native speakers do or understand.
Production was in Manassas, Virginia as well as numerous locations in Asia.
This could be a great product but the chance of even one bit randomly flipping is far too high.
Also if you look at the picture, there's five columns of packages per channel.
There's UDIMMs (typical RAM from desktops), and then RDIMMs (Registered), and LRDIMMs (Load-Reduced DIMM).
Because RDIMMs and LRDIMMs are only used on servers, they're all ECC / Error Correction just implicitly. At least, I've never seen an RDIMM or LRDIMM missing ECC.
The parent said the chances of a single bit flipping are too high without it, they are, which is why it’s in the spec :) These densities wouldn’t be possible without it.
ECC memory modules are better because they can fix and report errors, the base DDR5 DIMM will hand you a piece of memory with multi-bit corruptions that it couldn't fix. The ECC DIMM will fix larger errors and tell you about each detected corruption.
But there are ECC DIMMs with even better ECC. DDR5 moved the DIMMs from one 64-bit channel to two 32-bit channels. ECC DDR4 you'd protect 8 memory dies with 1 memory die holding the ECC, your computer will identify it as 72-bit memory. ECC DDR5 you protect the 4 memory dies in the half channel with a dedicated 5th die on each channel. Some ECC DDR5 will have 80-bit ECC, some will have 72-bit ECC. You want the 80-bit ECC because it stores 8 more bits of ECC for your memory controller to correct errors.
Micron isn't known for making 80-bit DIMMs.
It's really simple: just demand SECDED.
That's all you need to know, one acronym, SECDED.
https://cr.yp.to/hardware/ecc.html
Single Error Correction, Double Error Detection.
BTW, this sort of ECC fraud has been going on since the 1990's at least. It isn't going to just go away. Learn what SECDED is and demand it.
On Linux:
# cat /sys/devices/system/edac/mc/mc*/*/edac_mode
SECDED
SECDED
SECDED
SECDED
SECDED
SECDED
SECDED
SECDEDWhere a non-ECC DIMM is usually 8 DRAM chips (64 bits) and in DDR4 or earlier and an ECC DIMM would be 9 DRAM chips (72 bits), now with DDR5 can have either 9 DRAM chips or 10 DRAM chips (80 bits).
The question of whether ECC protection is strong enough to provide SECDED also is insufficient to address the complexities of having ECC protection only on the link, or end-to-end via sideband or in-band ECC, or using a combination of separately implemented link ECC and on-die ECC.
> the complexities of having ECC protection only on the link
That's not SECDED, since it can't correct a bitflip that occurs somewhere other than on the link, nor can it detect a double bitflip in those situations.
SECDED. Just ask for SECDED.
Even the Linux kernel knows about this:
# cat /sys/devices/system/edac/mc/mc1/csrow3/edac_mode
SECDED
SECDED. Just ask for SECDED.You're just making up arbitrary rules. The term SECDED does not incorporate any guarantees about which part of a system it applies to. It's just a mathematical statement about the strength of the ECC in use plus the implication that detectable but uncorrectable errors are reported.
What you're looking for requires more words to fully specify, probably including the term "end to end".
What I’ve read is that DDR5 has “on-die ECC” which means that the module itself handles the error correction, but then it can not handle errors that happen in transit to the CPU as full ECC can.
There are DDR5 full-ECC modules that would be real equivalents to DDR4 ECC.
> When the latest version of DDR memory – DDR5 – was introduced in 2020, marketing campaigns got one important fact wrong. DDR5 UDIMM, the regular desktop RAM we’re all familiar with, was touted as having “built-in” Error Correcting Code (ECC). Not true. What it has is built-in data checking, which is a very different feature that shouldn’t be confused with traditional ECC. For that, your customers will need DDR5 ECC UDIMM memory.
— Intel, “Error Correcting Code (ECC) vs. Built-in Data Checking”: https://www.intel.com/content/www/us/en/content-details/7549...