In the future even your RAM will have firmware; and the subject of POWER10 blobs
devever.net
devever.net
https://en.wikipedia.org/wiki/Fully_Buffered_DIMM
That SERDES chip unavoidably increases latency, roughly doubling it. This isn't likely to catch on as the only kind of DRAM in the system; latency is too important.
Maybe we'll see successful systems with two levels of DRAM: small low-latency parallel DRAM and a large heap of higher-latency serial DRAMs. That's plausible.
Yeah, you would think so. But it really hasn't improved all that much over the years. (Of course, that's partially because it's hard to do!).
Roughly speaking, when I designed with DRAM in 1983, the access time from RAS was 150 ns for the most popular parts. Now, with DDR4, you're talking CAS latency of about 10 ns, an improvement of 15x.
In the same timeframe CPU clock rate increased from about 10 MHz to about 3000 MHz, an improvement of 300x.
I grok your comment about perhaps having two levels of DRAM, but the whole memory hierarchy thing is already crazy. Most computers already have 3 levels of cache ahead of the DRAM.
IIRC, Rambus RDRAM was also serial, or at least less parallel than conventional DRAM.
It does! It's not in the mainline source tree, but is tagged. The most recent version appears to be https://github.com/open-power/ocmb-explorer-fw/tree/ea3fae3c...
Then in rapid succession starting in the early 2000s, we saw USB and SATA/SAS replace most IDE and SCSI use cases, and PCI Express replace PCI and AGP, along with Rambus taking a swing at conventional SDRAM.
Was there some fundamental breakthrough that made it possible to do high-speed serial busses, or was it that the costs of managing them finally undercut the hassles of signaling difficulties on a parallel bus?
The other thing I wonder about is what happened to SRAM? I know it always used a price concern-- "64K of DRAM is $10, and 64K of SRAM is $50", but given that DRAM is so cheap now, I'm surprised there's no SRAM product for enthusiasts/cost-no-object performance seekers. They'll already pay 5-10x the price of commodity DDR4 for the highest-binned overclockable stuff.
It's also $3000 per TB, triple the price of DDR4, and not much faster.
The real benefit vs. NAND is that it doesn't care about read/write mix and has max IOPS with any mix with consistently very, very low latency.
As I understand it, the switch to serial and point to point are because of three things.
a) pin/wire count reduction --- you can't have a small chip that uses a pc parallel port, cause it needs a lot of leads, etc.
b) bus termination gets weird with potentially variable reflection and distortion at high speed
c) clock skew is hard to manage on wide busses at high speed. You gotta make sure all the leads are the same length and/or add compensation logic, and it's a lot harder than a serial connection which might have a clock right next to the data line, or might be self-clocking, so the data line includes the clock signal and it can't be out of sync.
About the only parallel bus left in modern computers is RAM. Rambus was a disaster, and probably soured the industry on a serial ram interface for some time. Note though that a RAM channel now usually has a max of two slots, but boards with only one slot per channel are common too; 4 slots per channel was common many years ago, when speeds were slower.
If you think about PCIe we moved from a 32-bit bus (32 wires bidirectional) to x16 slots with 64 wires in total (16 lanes, 2 directions each, +ve/-ve for each). Maybe you consider AGP as the true precursor of x16 PCIE, but that was also 32 I/O lanes. So there wasn't a wire saving.
The truth is if you want very high data rates over medium or long distances then you have to go with 'serial' (the three points above) and you can get anywhere between 2.5Gbps and 54Gbps per differential pair depending on the technology. If you are willing to have shorter distances then you can get up to something like 6.4Gbps per pin and it'll use less power (vastly less power if the data rate goes down to <1Gbps per pin), but you won't be able to reach as far. Technologies like HBM use single-ended (non-differential) signaling and make up for this with huge amounts of very short wires (up to a thousand or so).
Given the constraints of the technology if I was designing RAM right now with no legacy baggage I'd chose single-ended signaling, and if I was designing a high-speed I/O slot I would choose differential signaling. So it's all quite sensible and rational.
You had SPI, which has a clock and data pin. That got faster and faster. Then there was QSPI (quad SPI) with four data pins, so you could shovel more data per clock, e.g. when loading an initial FPGA configuration, or when a fast CPU wanted to XIP from external flash.
Now apparently there's Octal SPI? Which is... arguably a parallel bus!
DRAM consumes less power than SRAM. Far less. My understanding is, that simply changing the parameters of DDR4 towards lower latency would yield better results.
I suspect it's because the trade-off for the increased performance would be more than just the price, and that's not even considering the CPU changes and motherboard design necessary to implement SRAM main memory - no CPU maker is going to go to the expense of a separate production line just for enthusiast CPUs that support SRAM core memory.
SRAM density is significantly lower (cells/mm^2) then DRAM, because a DRAM cell can by implemented as a single capacitor, where a SRAM cell is six MOSFETs. So very generally speaking the amount of SRAM you can pack into a chip is about 1/6 (or fewer) the amount of DRAM cells in the same space.
Given that modern CPUs typically have only four DIMM slots populated in pairs and you can't add more without changing the memory interface, that means you've cut the total amount of memory possible in a system by about 80%, so max 6.4GB instead of 32GB.
So.. a memory type for "enthusiasts" that costs e.g. 5x as much as DIMMS, only allows you 20% of the memory size, and also needs a motherboard and CPU design that's incompatible with DIMM based motherboards. Not attractive to too many people.
If you want an easier way to make an enthusiast product with the performance benefits of SRAM, just put the memory on die on a multi-chip module so it's a very short distance from the processor cores (like Apple's M1). Short distance equals good signal integrity, low latency, and on die means simpler motherboard design. The trade off is that you can't add memory later, you have to buy a whole new chip.
> Probably the most notable SoC containing a Synopsys DDR4 memory controller and PHY is the NXP i.MX8M series, which was for example selected by Purism to form the heart of their Librem 5 phone.
Is this (still) true? I just tried to find any binary blobs in U-Boot source tree [0], but didn't find any.
At least for some Rockchip SoCs (RK3399), which I assume also use Synopsys IP DDR PHYs, I know that it's not necessary to provide anything else apart from upstream U-Boot and Arm Trusted Firmware (ATF) [1]. And AFAICT current mainline U-Boot can initialize NXP i.MX8M's DDR PHY, too.
[0] https://gitlab.denx.de/u-boot/u-boot
[1] https://stikonas.eu/wordpress/2019/09/15/blobless-boot-with-...
I wonder how I could have missed that issue with the Librem 5 till now.
I've personally created a BSP and did board bring-up on a platform based on the NXP LS1046A (one of their ARM based QorIQ SoCs). That chip uses DDR4 and didn't require any blobs (aside from the Ethernet hardware). All of the memory controller registers were fully documented too. The DDR4 controller there was very similar to the DDR3 controller on their older T2080 PPC chip that I've worked with.
In fact, one of the most annoying parts of the board bring up was determining what data strobe delay values to use via a process of trial and error.
I'm not sure why NXP would use a different DDR4 controller for the iMX that requires a blob while their other chips don't.
I've confirmed the presence of these blobs in i.MX8M BSPs myself, as has Purism, see their blog post.
https://github.com/u-boot/u-boot/blob/master/arch/arm/mach-i...
I don't know whether RK3399 uses Synopsys. There are other DDR4 PHY IP providers, like Cadence, which don't require firmware.
That would enable error detection circuitry to redirect some reads and writes around defects in the silicon without incurring a latency penalty for the majority of reads and writes.
I suspect it's not a more popular scheme because it makes it hard to pipeline your activity if there's unpredictable timing somewhere.
One of the most interesting things which will happen in the near future in computing is the adoption of serial memory interfaces.
For pretty much the entire history of modern computing, RAM has been attached to a system via a high-speed parallel interface. Making parallel interfaces fast is hard and requires extremely rigorous control of timing skew between pins, therefore the routing of PCB traces between a CPU and RAM slots must be done with great precision. At the speeds of modern parallel RAM interfaces like DDR4, what is theoretically a digital interface in practice must be viewed as practically analog (to the point that part of a CPU DDR4 controller is called the “PHY”). Moreover, the maximum distance between a CPU and its RAM slots is extremely tight. The positioning of RAM slots on a motherboard is largely constrained by these physics considerations. For these reasons you have never seen anything like the flexibility with RAM attachment that you can get with, for example, PCIe or SAS. PCIe and SAS are serial architectures which support cabling and even switching, allowing entire additional chassis of PCIe and SAS devices to be attached to a system via cables.
Parallel RAM attachment methods like DDR4 and DDR5, by comparison, are both inflexible and pushing the physics to the limit (it is unclear to me whether there will even be a DDR6). Due to the complexity of the analog concerns when running parallel interfaces at such high speeds, the size of a “PHY” IP block gets larger and larger with each successive iteration of DDR, taking up more room on a CPU's silicon die. For a CPU with eight memory channels, the amount of space taken up by DRAM controllers and PHYs is now substantial.
For these and other reasons, the move to serial DRAM attachment is being considered by industry, most notably by IBM. The idea is that a multi-lane serial interface, not unlike e.g. PCIe would be used to attach DIMMs rather than a parallel interface."
PDS: I've foreseen the coming of serial memory for a long time now. It just makes sense.
What I would love to see is a whole ecosystem of open-source, open-hardware high-speed (RAM speed) serial bus, serial connection, serial protocol, etc., etc.
It should be one-size fits all. Sort of like you have a CPU, and now you want to attach either perhipherals or RAM to it. Well, use the same high-speed open-source/open-hardware serial interface and protocol for all of them!
It would be a thing of beauty, should this ecosystem exist in the future...
[1] - https://www.computeexpresslink.org/post/the-benefits-of-seri...