Samsung Unveils Industry's First 32Gbit DDR5 Memory Die: 1TB Modules Incoming
anandtech.com
anandtech.com
I have had defective RAM, and I got quite a bit of corruption before the first crashes, it is hardly noticeable when it is just a pixel changing color in a picture, but it is still something you don't want. ECC would have prevented that.
I know there is software resistant on random bitflips, like for satellites exposed to cosmic rays, but it is a highly specialized field. It is also a field where they use special chips, typically with coarser (and therefore less efficient) dies that are more resistant to radiation. You leave a lot on the table for that.
If it’s a random, once in several billion reads/writes issue, it can just stop/identify the bad data from further propagating. Sometimes. That data is still lost.
ECC does forward error correction, which is extremely rare for the type of data protection you’re talking about. and if the data is corrupted in RAM (say when initially loaded/read) before the software can apply FEC, there is nothing the software can do.
ECC modules just have more chips to store the extra parity information. In the high capacity RDIMM server market there are plenty of ECC options.
You can get ECC support on Intel 12th and 13th generation parts by buying a motherboard with a W680 chipset.
You can get ECC support on modern AMD CPUs by picking a motherboard that lists ECC support listed on the product page. It’s not that hard.
I know some people won’t be happy until every laptop has ECC RAM and is super cheap, but the reality is that the demand for ECC RAM is very low. The majority of users would choose the extra battery life and lower price if given the option.
Nice circular reasoning. But nothing will change till we're not vocal enough about ECC benefits and shady pricing. I assure you though, it's not about my happiness :)
AMD otoh has brought ECC to the table in Ryzens without the same shenanigans
For the higher priced models you cant even order them with non-ECC memory.
Registered (meaning ECC and buffered) RAM is common in the workstation market, so it is not limited to noisy servers.
Check out HP Z series and Dell Precision workstations. They are available used / refurbished at low prices.
Having your system’s available memory fluctuate up and down based on how many segments are currently set to ECC also doesn’t sound fun.
Having developers manually turn ECC off for regions where it’s unimportant sounds like a lot of complexity for a relatively rare use case.
There is in-band ECC in some newer Intel designs, but it’s all or nothing. Adding extremely complexity to memory management to selectively disable it sounds like a lot to ask.
But the parent comment was suggesting that it be on or off depending on the memory segment, which is a completely different problem.
I think your reading depends on thinking "application" means "process", while another reading would be that an application is a particular deployed system, where this setting can be altered e.g. at the BIOS level.
On AMD ECC support is pretty much standard on every chip they make, and always has been. Even my shitty 4-core phenom from over ten years ago on an el-cheapo motherboard supported swapping it's regular DIMMs for ECC ones. You're never going to get ECC "for free" but it would be totally possible for everyone to pay the cost once and just move to ECC-only for everything from now on.
Except intel, the company that brought software-locked hardware features to x86, love to price-differentiate.
I assume it is going to take another decade to fully unwind Intel's ECC market segmentation. Trying to get a sense on if I should pay the ECC tax for my next build. Of course noting that as a consumer, I will probably never notice a flipped bit.
Which is often background noise for home users, but no less problematic.
Often heat/load dependent too.
On-die ECC allowed DDR5 to be competitive with DDR4. Is it really protecting your data at rest if the DDR5 die is running at such tolerances that it's correcting single bit errors from internal signalling issues every transaction? It's only single bit ECC, if something else outside of the die(Cosmic Ray, sudden voltage change, sudden temperature change) induces a bit to flip while the internal circuitry causes a different bit to flip your data is now corrupt.
https://www.atpinc.com/tw/blog/ddr5-what-is-on-die-ecc-how-i...
When ECC starts tripping on a device outside of completely random times is when you should look into what's going wrong. You may have overheating or failing hardware.
So I'm not sure how this works, because I'm not sure if "true" ECC is better/worse/same as on-die ECC. A casual googling shows on-die to have more advantages.
Note that link ECC + inline ECC don't give you end-to-end protection, since the controller in the memory module can still flip bits. DDR5 is moving to on-die ECC (which, unlike DDR <= 4's side-band ECC) also isn't end-to-end.
I'd like to see side-band ECC continue to exist, but I think it is going to be phased out entirely.
This article defines all the terms, but is very vague about what things are mandatory, or how reliable the error correction schemes are. For instance, it carefully doesn't say that SECDED schemes detect all two bit errors, instead it says they detect at least some:
https://www.synopsys.com/designware-ip/technical-bulletin/er...
I doubt it will be phased out for servers. I haven't seen anyone reporting that on-die ECC in DDR5 has a reporting mechanism, and reporting on ram errors is important for server reliability.
Also: "what Andy [Grove of Intel] giveth, Bill [Gates] taketh away." with regards to speed.
Only tangentially related… So nice that the chip shortages seem to have mostly worked their way out of the system.
Used server prices in particular have gotten almost ridiculously cheap lately. Just bought a 4 node EPYC server with 128 cores, 512GB DDR4, 16x NVMe slots with 15.3TB of P5500 storage populated for $2500.
Seems like almost insanity to get so much compute for that price, and yet in 5 years, as always, that machine will be considered slow and inefficient compared to the latest iteration.
It seems like progress continues unhindered by the death of Moore’s law. The hidden toiling of billions in R&D that make this possible is truly a modern marvel of capitalism.
Keep in mind it's only that cheap because someone already considers it slow and inefficient.
The higher the energy cost, the sooner that upgrading to more efficient compute pays off. But it’s still amazing to me that just the CPUs for this system were selling for $10,800 in 2020.
AI is the biggest driver for power at the endpoint since when 3D games came out. (I remember when Quake drove an entire generation to upgrade their PCs.)
Or you could just use ClosedAI APIs and all that will lead to…
If you don’t see specifically what you’re looking for, check out the listings and look for other items being sold by the same sellers.
After spending some time watching it you’ll even get a feel for what’s coming “off lease”.
Companies decommission servers all the time. Look for the companies that buy them. Data protection laws, however, had made storage devices trickier to get.
That is a 32 Core EPYC CPU, 128GB DDR4, 4TB NVMe SSD per node for $625. I am assume you mean physical core and threads. But I am surprised you could get an 32 Core EPYC for $625, unless it is not even a Gen 3 EPYC but much older. The 128GB and 4TB together would have costed at least $200. Meaning you need to get the CPU for less than $350 to leave some cost with PSU, fans and Server Case.
Either it is heavily discounted or this is a 2nd hand deal.
>Seems like almost insanity to get so much compute for that price, and yet in 5 years, as always, that machine will be considered slow and inefficient compared to the latest iteration.
Likely not.
That is because DRAM and NAND price has fallen to close or below BOM cost. Your 512GB DDR4 and 15.3TB SSD is less than half if not close to a quarter of its price 18 months prior. The EPYC CPU also has stepper discount due to slower HyperScaler expansion. Basically unless we have two or more node cycle to bring down cost, the current purchase price wont last.
I can get the same SSD I bought in February for literally half the cost today.
You can get larger drives. 4TB has pretty high availability. If you want more, you can go server grade U.2 and get 8-32TB SSD - you will pay more though.
Enterprise space can be reasonable to look at.
7.68TB $400, similar to the QVO but much better quality flash and interface.
https://serverpartdeals.com/collections/solid-state-drives/p...
It is U.2, so you'll need an adapter.
OTOH, a 16TB per socket server would be quite a beast.
It's always annoyed me that the memory allocation for VMs is so static. If I allocate 16GB to a VM, it will fill these 16GB with its own filesystem cache, even when the host could make a better use of that memory (like using it for other VMs). There's virtio-balloon, but it has to be adjusted manually.
When not working on these tools, it was all just buffer cache for Linux while the actual processes were easily living in 8 GB or so. Even today, I'd comfortably live in 16 GB on a laptop, except once in a while miss the option to just do sloppy large allocations for a single task instead of worrying about chunking, IO, and careful access order.
Another reason was to enable me to hack with performance optimizations like mounting freqent read/write files into memory (tmpfs). Putting Chrome's cache there for example can be a noticeable increase in performance.
Another reason was just for the sheer badassness of having so much RAM
I slap $400 worth of memory in my workstation and I can load large datasets or piles of VMs without worrying that some mistake on my part is going to generate a huge bill on my part.
It's pretty common for me to run 60-80GB worth of VMs during the working day, so not using could is a pretty massive savings for me.
https://www.sqlite.org/appfileformat.html
> Conclusion
> SQLite is not the perfect application file format for every situation. But in many cases, SQLite is a far better choice than either a custom file format, a pile-of-files, or a wrapped pile-of-files. SQLite is a high-level, stable, reliable, cross-platform, widely-deployed, extensible, performant, accessible, concurrent file format. It deserves your consideration as the standard file format on your next application design.
This is why my 128Gb of ram is ECC. Long running programs are most at risk of memory errors.
With 64GB, I mount /tmp and ~/.cache as tmpfs, which speeds up web browsing and code compilation, but with more capacity I would mount /var or even / into RAM.
Granted, NVMe drives are so fast these days that the improvement of using RAM for storage is likely not noticeable. Plus there's the problem of persistence. If you're OK with volatile storage for your use case, then this won't be an issue. Otherwise, you need some way to ensure data is persisted.
Perhaps these RAM advancements will make hybrid DRAM/NAND drives cheaper and more performant.
On environments with log shipping, having /var/log as a tmpfs makes a lot of sense.
the neat thing is that these programs sync back periodically so the contents are preserved
this gives you both the speed of ram, eventual persistence, and reduced wear down on ssds