DDR4 memory interface
edn.com
edn.com
Starting in the 90s there was a trend of replacing parallel buses with serial ones (USB, SATA, etc.) but I think it's slowly beginning to reverse as people realise the difficulties of using high frequencies (e.g. multiple lanes in PCI-E).
Incidentally the term DIMM - dual inline memory module - comes from the doubled width over the existing SIMM, which was 32 bits wide. Maybe we'll see 128-bit-wide QIMMs in the future...
There is even JEDEC specification for DIMMs with 32-bit wide data bus (Such DIMM was used in Corel's NetWinder for instance).
And as for parallel vs. serial, essentially whole second page of the articles is about issue that is mostly specific to synchronous parallel interfaces. Multilane PCI-E neatly sidesteps such issues by using what essentially amounts to multiple independent serial links that are bonded together. But having serial interface to RAM chip greatly increases the complexity of the RAM chip (SDRAM protocol looks at the first glance as something that is very hard to implement in RAM chip, but in comparison with interfaces like PCI-E it's trivial state machine).
Also the general trend of weird (i.e. PCI) and serial interfaces in 90's was motivated by per-unit costs at the expense of required engineering. Primary motivator for PCI's reflected wave switching and multiplexed address and data was to limit number of pins and number of required passive components on motherboard. Serial interfaces that came after that (SATA, PCI-E) were motivated by the fact that routing fast parallel synchronous buses for any significant distances is hard problem because of propagation times, which have to be roughly equal for all bus wires, which implies that PCB material parameters have to be known and reasonably consistent (and that means significant per-unit additional expense in PCB manufacture and testing cost).
Currently only interface that requires careful routing and design on PC motherboard is memory interface, which can be made short enough that manufacturing differences are negligible. (It's funny how TI has 20 page application note on correct routing of USB2 which includes rules as "no vias preferably, at most one", "no stubs", "controlled impedance" and "as short as possible" while Intel's layout recommendations for USB2 can be summarized as "it's differential pair, discontinuities do not matter much, use sensible routing")
[edit: formatting + missing word]
In other words, RAID-0, for memory. :-)
I wonder if it would be interesting to have memory controllers managing striping/mirroring of memory modules and create read-optimized and write-optimized memory regions. 2 mirrored modules would give you half the latency for reads and the same latency as a single one for writes.
Has anyone already done this?
edit: and now I'm imagining an inter-memory-module bus to manage bank transitions to/from mirrored/striped without loading the processor bus.
its like two individual drives, one for /etc the other for /usr
and no, you cant have half the latency, latency is dictated by physical speed of actual ram inside chips, those are clocked at 200MHz for typical DDR3 1600Hz module, up to 300MHz for fasters DDR3 ones.
This is already common. Log onto Dell.com and configure a high end server, you will see an extensive number of options regarding mirroring and advance ECC configurations.
http://dmitry.gr/index.php?r=05.Projects&proj=07.%20Linux%20...
In order to run Linux on an ATmega1284p, that guy:
- bit-banged 16MB DRAM
- bit-banged SPI Flash
- wrote an emulator for a 32Bit ARM & MMU
Apparently it takes 4h to but Ubuntu.
For a ROM chip you actually had to be careful so that it actually generated the correct output for a given address, unless you were willing to burn an EPROM with the reverse transformation applied.
Other RAMs/Flash with different interfaces (e.g. SPI) require a prefix command to be sent over the wire, so it's easier for them to implement things like chip id commands.
When you do a suspend-to-ram on a current x86 machine, it has to save the "scrambler seed" to CMOS so that it can decode the data it left stored in RAM.
I think this also helps mitigate the attack where the DRAM is chilled to make its contents less volatile, the machine is powered off, and the DRAM is then dumped in search of sensitive material.