AMD and Jedec Are Collaborating on DDR5 MRDIMMs at 17,600 MT/S
hothardware.com
hothardware.com
If you have 96 CPU cores per socket, and only 6 channels per socket, that's a "mere" 625K IOPS per core. Up that to 128 cores (coming soon), then it's just 470K IOPS per core! That's worse in some sense than a laptop SSD for a single-threaded program. Not directly comparable, of course, but you can see why AMD would want to bump this number up.
For comparison, the equivalent IOPS for L2 cache is a blistering 260 million IOPS, and L1 cache is 850 million.
I see now why there's the phrase "memory is the new disk" is starting to get popular.
Not to discount what you've said but it's slightly better :p
I think you're mixing up latency and throughput. IOPS is throughput, and any single CPU core can have several RAM accesses in flight, not just one. So with DDR4-3200 you have more than 3 billion "IOPS" per channel, not 10 million as in your calculation.
On DDR5, channels are 32 bits wide, but there are two in a DIMM, so a normal desktop system is actually "quad-channel", with each channel doing up to that 400 million IOPS if the ram speed is DDR5-6400.
In practice, it's of course rare to completely saturate channels, just because of bank conflicts and refreshes, etc.
You can have 1M IOPS at 5MB/sec if your blocks of data are small enough, like 8k... vs 100k IOPS/sec at 1GB/sec throughput with 1MB request sizes...
In this case 100k IOPS are getting you more throughput than your 1M IOPS.
Yes, it is. You're thinking of "bandwidth"; the terms "throughput" and "bandwidth" are not interchangeable. IOPS and bandwidth are two different forms of throughput: the former has number of operations in the numerator, the latter has bytes.
That is not how any of this works. Neither RAM nor SSD can get anywhere near their peak throughput if accesses are serialized. If your laptop has DDR5, it can likely do > 1B iops. 100ns is the time to get a single random access to a closed row, but while you are waiting for that access to happen, you can issue new operations every 3ns to different banks and bank groups on the same channel.
> [can't] get anywhere near their peak throughput if accesses are serialized
But first of all, doesn't DRAM have an optimisation for serial accesses, which is (I hope I get this the right way round) to hold the RAS steady then do CAS/CAS/CAS/CAS instead of RAS/CAS + RAS/CAS + RAS/CAS + RAS/CAS?
In addition, if you do access RAM serially then there is some interleaving which is automatically inserted so that sequential accesses at some level are sent off to different banks (but I've never been clear if there interleaved at a 64 bytes, to fit a cache line, or at a 4K page size to suit the TLB, or some other size, and I'd really like to know – anyone please?)
What I mean is that if you are doing an infinite pointer hop (that is, load a value, then use the that value as a pointer to load another value, use that as a a pointer... etc), you only get a tiny fraction of the IOPS you get than if you launch 100 different loads to different addresses that you already know. And the inverse of your load to use latency is basically how many IOPS you get if you do that infinite pointer chase.
> In addition, if you do access RAM serially then there is some interleaving which is automatically inserted so that sequential accesses at some level are sent off to different banks (but I've never been clear if there interleaved at a 64 bytes, to fit a cache line, or at a 4K page size to suit the TLB, or some other size, and I'd really like to know – anyone please?)
They are interleaved at the size of the DRAM row, which is usually 2^16 bits, or 8192 bytes. The reasoning here is that if you are launching a ton of linear accesses, that is, doing CAS/CAS/CAS/CAS... to the same row, you can get full throughput from it. This is because when you have opened a row, you are not reading from the DRAM anymore, the entire row is in a SRAM buffer, and can read full interface throughput from it. Then you only have to open a new row once the current one runs out, which is always in a different bank so can be done in parallel with the read operations from the current row so long as your memory prefetchers are smart enough to start doing it early enough.
RAM also does overlapping requests and has a queue depth, so that's a reasonable comparison. The headline of this article says they are standardising RAM with a throughput of 17600 million IOPS.
If you have 96 CPU cores per socket, and 6 channels per socket going to separate MRDIMMs, that's 1100 million IOPS per core. At 128 cores, 825 million IOPS per core.
So, this new RAM standard is 825-1100 times faster throughput than your laptop SSD, when counted in IOPS (without regard to the data size).
If you want to compare latency instead of throughpt, the IOPS measurement for that occurs when the sender waits for each request to complete before issuing the next one. In that case you're right that it's will be much lower IOPS per CPU core, but this also makes it lower IOPS for the SSD so that comparison still favours RAM.
Each CPU core is independent, so each core would likely wait for its own previous request before issuing the next one (this actually happens when a CPU core is following a linked list), but the cores are doing this independently of each other, in parallel.
So with 96 CPU cores per socket, each core waiting for its own request to complete before issuing the next one, and assuming ~100ns latency, that's 960 million IOPS total, and 10 million IOPS per core. With 128 CPU cores it's still 10 million IOPS per core.
It is just a weird way to express throughput. An SSD that loads 4KB at a time has 64 times less IOPS than RAM that loads 64 bytes at a time assuming both have the same bandwidth.
From a practical perspective, the IOPS of RAM is critical for running programs. And if you're doing anything other than big file transfers, an SSD plugged into USB 2.0 at 35MB/s will beat a hard drive on USB 3.0 at 200MB/s.
Even better is IOPS at queue depth 1, IOPS at queue depth 16 or 32, and max bandwidth.
We’re very much in a situation where the software is catching up to the hardware. For example, look at the proposed NEST Scheduler for Linux where they were seeing 10-100% (yes, as much as double!) performance improvement with a 10-15% reduction in power consumption, mostly by focusing on keeping, “hot cores hot.”
IO scheduler and process scheduler improvements will offer truly material gains in coming years, even if you still use the same high-end hardware you have today.
A simple back-of-the-napkin maths is that for a system with 30GB/s memory bandwidth and 32-byte cache lines works out to about a billion I/O operations per second. This would be a typical laptop system, or the like.
However, 100ns per read is still a valid scenario for a single-threaded program "chasing pointers" as in a linked-list.
I think it's still fair to say that an un-optimised single-threaded program accessing memory randomly is only 10x faster than a parallel program efficiently utilising a modern NVMe SSD.
The corollary is that there is (still) no scenario where an SSD would "beat" the performance of RAM since the worst-case for RAM still beats the best-case for SSD.
All the old-school HDD-optimization techniques from traditional database programming, etc. are now fully relevant again for memory management.
For example HDD latency was very sequential with the heads stacked on top of each other. Meanwhile modern memory has has latency for a single request but you can make several requests at the same time and latency doesn’t stack linearly.
Modern RAM allows paralleled access where HDD effectively only had a single read/wrote head. People also on average do a lot more computation on any given bit of memory.
Wat? No it's not. It's closer to 5x-6x than it is to 5 orders of magnitude if you're looking at instructions the CPU normally executes.
> All the old-school HDD-optimization techniques from traditional database programming, etc. are now fully relevant again for memory management.
Those techniques were mainly managing the fact that it took 100s of milliseconds to random seek (like, human perceptible times to do one seek) so the algorithms tried hard to minimize that. It doesn't have much bearing on RAM with something like 100ns random access.
1993 was when the original 60-66Mhz Pentium first shipped; 5 orders of magnitude may be an exaggeration, but every step since then has had of increases beyond just clock speed, and clock speed alone has increased nearly two orders of magnitude.
https://en.wikipedia.org/wiki/CAS_latency#Memory_timing_exam...
The widening gap between memory latency vs speed of processing in the processor (and how that arguably makes access optinizations that used to make sense between processor and HDD sensible between processor and main memory) is the issue that was being discussed, so memory latency can't be counted as part of processor speed in that context.
The geometric midpoint (the only one that makes sense here) between the high end of 5-6x and the low end of 5-6 orders of magnitude is about 774×.
The best info I can find (unsurprisingly, direct comparisons of performance of long-separated-in-time processors in aggregate is hard to find) is that between 1995-2011, integer performance increased somewhere well over 128 times but less than 256 times and floating point significantly more than 256x, and was trending in the last several years of the period to increase at a pace of 21% per year for both. [0] Seems likely to be close to or significantly past the 774× level between 1993-2023.
[0] https://preshing.com/20120208/a-look-back-at-single-threaded...
Kind of. With memory, at least, we don't need to do elevator seeks.
But yeah, random access to memory is bad, you can only get to the max performance if each thread requests for somewhat contiguous memory blocks
Wrong how? The internal structure doesn't matter very much. Channels per socket and cores per socket, to calculate channels per core, is the correct math.
Any _particular_ RAS + CAS to read a 64-burst will be ~100ns latency. But you can perform 16 of them in parallel on DDR4 (probably 32?? in parallel on DDR5?).
That is: Bank-group#0-Bank#0 RAS -> Bank-group#1-Bank#0 RAS -> ... Bank-group#4-Bank#0 RAS -> Bank-group#0-Bank#0 RAS ... Bank-group#4-Bank#4 RAS -> Bank-group#0-Bank#0 CAS (finally finish the 1st read/write)... etc. etc.
DDR4 sticks are designed to only be fully utilized when you have 16+ concurrent RAS/CAS operations going all at once.
-------------
If you have two sticks on a line of DDR4, you have a "rank" as well. So Rank#0-Bankgroup#0-Bank#0 -> Rank#1-Bankgroup#0-Bank#0... (etc. etc.) 32x in parallel (alternating between Stick#0 and Stick#1).
DDR4 sticks already can be multi-rank. So this MRDIMM seems to be another layer of parallelism (and DDR5 upgrade was probably another layer of parallelism again, because 100ns latency doesn't seem to be improved no matter how many years go by).
Thats basically what Intel Falcon Shores is, and I am sure AMD is cooking up something similar behind the MI300.
https://www.amazon.com/ATI-Rage-Video-Card-109-43200-10/dp/B...
We already have the ability to put far more memory in GPUs than they currently have, but we don't because it hurts throughput. I don't think technical advances will change the existence of a throughput/capacity tradeoff, but I do think it's an interesting idea that maybe DL models benefit from a different tradeoff than graphics applications. Personally I'd guess that's a larger factor here than incremental changes in capabilities, because the two types of memory are both improving at the same time. But who knows! Maybe the technologies will converge enough that it will no longer make sense to have two different variants, though I wouldn't bet on it.
Actual ML GPUs transitioned to HBM, but the capacity is relatively limited. I think Nvidia thought 96GB per GPU would be plenty when they were planning things out years ago, but the explosion of model sizes seems to have taken all the hardware manufacturers by surprise.
However, if your usecase is sequentially reading massive chunks of memory, this will be useful for you. But dear god am I afraid of seeing the latencies on these beasts.
That OEM motherboard in your big box tower.
>All consumer AMD CPUs should be able to clock up to 6000 MT/s and Raptor Lake should do 6600 MT/s stable at the very least with Hynix M-dies. 5600 MT/s is not even an overclock on Intel CPUs.
DDR5 speeds currently crash like a rock if there's more than one DIMM per channel, even moreso if those DIMMs have two ranks rather than one.
I've got four sticks of 16GB DDR5 5600MHz in my Intel i7 12700K, Z690 system. The best I can get the RAM going is 4800MHz, which is still an overclock because Intel's memory controller goes down to 4000MHz rated clock speed with this setup (two DIMMs per channel, 1 rank per DIMM).
You can: https://www.newegg.com/corsair-64gb/p/N82E16820236933
Maybe once prices come down more I'll consider going to two sticks of 32GB instead of four sticks of 16GB, who knows.
https://www.corsair.com/ww/en/Categories/Products/Memory/VEN...
4 X 48GB DDR5 5200 40CL latency. G-Skill just came out with some new kits that are even faster.
How would you determine that with dmidecode?
CL22 is also pretty damn bad, most DDR4-3200 kits now have CL16, which makes them 10ns. But even then, your DDR4-3200CL22 has 13.75ns, your DDR5-5600CL46 has 16.5ns. and that is assuming your PC will handle a 5600, which isn't the JEDEC standard, so most OEMs can say goodbye to it.
* CPU not supporting such a high rate
* Motherboard not supporting it
* CPU lost the silicon lottery and has a bad memory controller that struggles above basic rates
* vendor put in DDR5-4800 in anyways to save up on costs
* BIOS not up to date on your motherboard
> JESD79-5A expands the timing definition and transfer speed of DDR5 up to 6400 MT/s for DRAM core timings and 5600 MT/s for IO AC timings to enable the industry to build an ecosystem up to 5600 MT/s.
JEDEC _default_, not standard. Satisfied ? As in, the default value is 4800, and now, with JESD79-5A, JEDEC also recognizes that DDR5-6400 can exist. Not "is the standard".
A bit like what Apple did with M1/M2, putting ram/flash/cpu/gpu on the same package with very short, wide and fast interconnects.
Leaving only IO, specialized chips and power outside of the main package.
If you need an absolute ton of memory, having it all directly on the die becomes impractical at some point.
Epyc supports 6tb per socket - try getting that much memory on-package any time soon.
M1/M2 ram is just 128bit like any other normal mobile processor. M1/M2 Pro comes with 256bit bus (2 custom LPDDR5 chips) and 200 GB/s, M1/M2 Max with 512bit.
POP ram makes sure nobody will be able to upgrade, hardware locked recurring revenue.
When the arguments for DIMMs are "field upgrades" which is a problem you can avoid by buying more memory to begin with (Also in practice few users actually upgrade memory), and "price segmentation" which isn't really a technical consideration, I see the writing on the wall. At least in the consumer space.
AMD 3D cache is also going in that direction integrating ram/cache on top of cpu with high density of interconnect.
This does not seem to be correct. Modules with an 128-bit data bus would obviously require a lot more pins. Instead what this actually seems to be about is placing a buffer between the two ranks (essentially two completely independent modules) and the data bus (presumably it buffers the command bus as well). The "128 bit" data bus on the host side is achieved by running QDR over the 64 bit wide bus.
Basically an evolution of LRDIMMs with a pinch of FBDIMM.
AKA a Mux - or Multiplexer
https://en.wikipedia.org/wiki/Multiplexer
> In electronics, a multiplexer (or mux; spelled sometimes as multiplexor), also known as a data selector, is a device that selects between several analog or digital input signals and forwards the selected input to a single output line.
An example of a common digital multiplexer would be the 74HC157 but if you think they're really not multiplexers I'm sure NXP and Texas Instruments would love to hear from you
Multiplexing is a concept anyway, the medium doesn't matter - you can multiplex radio signals, light, sound etc.
Digital the values “round-to-nearest”, so if supply is 1 V, then 0.6 V, 0.9 V, and 1 V all mean logical 1. Whereas for analog those different values mean different things.
Also, for digital you want to push the signals to the extremes so they are easier to read, while for analog that would be destroying information, you want to preserve the value.
(In practice there are fundamental limits to how much useful information you get out of analog signals due to noise.)
Clearly the output signal has to be 0 or 1, which must mean the signal is boosted or reduced to an appropriate level (assuming we not trying to encode multiple values onto a signal; I'm just talking binary). So I guess the signal has to be captured (latched?) then modified. I don't think a clock is absolutely necessary as there have been asynchronous digital CPUs, but the presence of a clock clock is very typical as well.
A clock is not necessary in this definition. For example a simple logic AND gate is digital and has no clock. The input is not captured or latched before being modified; nevertheless it is a digital circuit.
You can use an analog mux on digital signals, but it probably won’t be as good as a digital one. However, you can’t use a digital mux to switch analog signals, it won’t work.
Edit: the one exception to this is for bidirectional digital signals you may want to use an analog mux, since you don’t know which side may be driving the lines. This would apply to DRAM since the data lines are bidirectional, so probably using an analog mux.
Do you want hardware vendor RMAs to increase or something?
The yield is maybe a factor, but that is solved with experience. Why is no company even trying?
Is it an IP issue? I seriously don’t understand (especially for phones) why it is still optimal to keep RAM separate.
Could someone help me understand this?
and I’d really appreciate leaving “upgradability” out of this conversation. That is a separate topic that can be had but not interesting to me right now.
But the chiplet approach works, and nVidia and Apple are both shipping it. I dunno what's going on at Intel. Hopefully they'll pull themselves together.
Huh? Sustained all-core load is why I have a workstation in the first place, it's for anything that would be annoying to run on my laptop. I'm sure different people have different workloads, but I don't think I'm an outlier here.
At least mine spends most of the time waiting for my keystrokes. It does do background tasks, of course, and some heavy lifting sometimes (when I have to wait for them instead of the other way around) but it doesn't spend the day doing the heavy lifting because that'd make the machine much less responsive (and much noisier).
OTOH, if I had a 56-core 112-thread monster, it'd probably be perfectly fine to do the heavy lifting AND the interactive work at the same time. Worse case scenario would need to set core affinity so that the heavy lifting doesn't step over the cores running my GUI.
- Heat generation (harder to cool)
- Trade offs don't justify the pain
At the highest level isn't it all EUV? Why can't the laser print both patterns onto the wafer?
Sapphire Rapids has models with 64GB of "High Bandwidth Memory".
It's not upgradability any more. That means you buy the machine, run out of ram for your working set 2 years later, add more ram.
Now Apple refuses to sell you machines with enough ram from the start.
Would you like your preferred server vendor to do the same?
The PCIe link isn't going to be competitive with the DDR5 link, since both are measuring DRAM access speeds in this case.
Yeah, if you start getting higher speeds, it becomes more a question of if the memory controller on the cpu, the wiring on the motherboard, and the ram itself will work together at that speed, which you can't guess from just the number on the ram packaging.
It would be nice if JEDEC offered more ratings, but given that they don't, I'll take the XMP numbers.
Give us both the rated JEDEC clocks and timings and the manufacturer overclock clocks and timings.
To use the DDR4-3200 example, it is far more difficult than it should be to figure out whether that's JEDEC 3200MHz or actually JEDEC 2166MHz with 3200MHz XMP overclocking.
You've been lucky (or you're using latest and greatest mobos).
But I'm just getting DDR4-3600, not DDR4-5000 or something; and I think all my DDR4 setups are 1 dimm per channel, which helps. I didn't ever get DDR3 XMP to be happy though.