The Path Is Set for PCI-Express 7.0 in 2025
nextplatform.com
nextplatform.com
Since they will be pushing the limits of the SNR, they also build in forward error correction to keep the error rate under control.
Remember socket AM4, which added support for PCIe 4.0 with chipset X570? Originally, the idea was that non-X570 boards would also have PCIe 4.0, as it doesn't actually depend on the chipset - just the CPU. Turns out most boards simply weren't capable of handling the additional bandwidth, so it was eventually restricted to boards which were designed for it. And that's at one-eighth the bandwidth of PCIe 7.0!
In practice it's a bunch of point-to-point links, but even then the path across a PCB is pretty brutal in terms of loss. Cables are (surprisingly to me, initially!) better.
Overall though, sending some number of additional bits for error correction is often cheaper than driving the line harder to reduce the bit error rate.
Put differently: we could manage lower noise, but it would either result in a lower bitrate or (much) higher power. Instead, pick a target goodput and design for power and error correction from both sides.
That's because the cables in these application are not single wires but complex micro-coaxial assemblies. (The industry best is 3M Twin Axial.) And that will have much more shielding than an inherently unshielded PCB trace. With an appropriate price, of course.
Note that on balance I'll take something like an OSFP assembly (https://www.te.com/usa-en/products/connectors/pluggable-conn...) with a fancy jacket that pulls all of the individual twinax cables together, but it's also substantially more expensive than a simple flat tape-wrapped ribbon.
Other places where you'll see coax cables ganged together inside a jacket: high-speed USB cables. For example, here's a type C cable cross-section: https://twitter.com/tubetimeus/status/1125926941469462528
At this point the bits are millimeters long.
[1]: https://en.wikipedia.org/wiki/Peripheral_Component_Interconn...
[2]: https://en.wikipedia.org/wiki/Reflected-wave_switching
Of course, SNR is another big one. You have to actually be able to distinguish the voltage levels within a symbol period from each other.
Another thing: PCIe is meant to be low-latency, which just rules out a lot of the modulation and error-correction techniques that are traditionally used to enhance channel utilization, like having multiple levels of FEC and then interleaving FEC symbols temporally. That's done basically always where latency allows, e.g. in storage media, digital TV and so on.
With 2 levels you have low and high. I have no idea if PCIe is 1.8V but let’s be really gracious here and pretend it’s 5V.
Ok; so at 5V and 4 levels you have Sig Al’s between 0V, 1.25V, 2.50V, and 3.75V. And are encoding 2 bits per clock.
At 64 levels you have a signal every 0.078125V. It’s going to be really hard to capture that signal at a moment in time and at these speeds be certain if it’s definitely A or B.
At some point you get into modems (modulation, demodulation) but the add much more cost, latency, and complexity.
So before getting into radio and analog tech, here is a simple way to double the bandwidth for little cost. A driver only needs 0,1,2,3 levels and a receiver can make fairly confident guesses quickly.
I’m pretty sure PCIe isn’t 5V. Probably more like 1.8V, so already starting with a much smaller range.
The "easiest" way to get around noise problems is to put more signal in, but power consumption increases with the square of voltage -- so it's extremely costly to overcome noise with more signal. However, noise is a fundamental limit; there's a set of tricks you can maybe do to improve it but for example, kT/q (thermal voltage of noise at roughly room temp) is about 25mV and that plays a role in a lot of device physics.
Thus most signals engineers would steer away from higher-order PAM and go to QAM, where they use phase to encode one dimension of data, and voltage to encode the other. This gets you out of a one-dimensional line of voltages and into a two-dimensional matrix of voltage vs. time to represent information.
This helps fight against that square-law term for power, but it comes at the expense of more precise timing sources. Time also has noise mechanisms, that are typically only overcome at the expense of -- you guessed it -- power.
At the end of the day, probably the most fundamental limit all computation will run into is power -- both how much you can afford to put in, and more importantly, the limit of what you can pull out reliably and efficiently. I'd be curious to see the plot of PCI-express energy-per-bit over its generations...I suspect it improves over time, but not improving as fast as bandwidth is going up, which means in net each link should run hotter.
I'm fairly sure the most modern PCIe standards are all the way down to 1.2V. Prior to 1.2V it was 1.5V, 1.8V, 2.5V, and back in the days of the original spec it was 3.3V. I can't remember exactly what versions speced what voltages.
What the actual digital IO voltage that the chip that does PCIe driver is driven from is almost unrelated.
A flip flop can be calibrated to the difference potentials of thr input. You have zero and one as a result, no processing required. Single bit, single voltage. Simple.
Think about what it means to map a single voltage input to multiple bits. What device would do it? A complex ADC? Think of the overhead that would introduce.
Though I could imagine at one point you're just looking at larger buses and having to deal with those consequences
I wouldn't be surprised if that's what they'll be using, though I'm not an expert either so might be better options out there.
I know interleaving is another approach[2] for reaching high sample rates, though it introduces latency.
What other approaches are there for low-precision multi-Gs/s ADCs?
[1]: https://www.analog.com/media/en/technical-documentation/data...
[2]: https://www.analog.com/en/technical-articles/a-12-b-10-gss-i...
I'm amazed at the progress of terabit Ethernet, soon 1.6Tbit/s in one server.
https://en.wikipedia.org/wiki/Quadrature_amplitude_modulatio...
EDIT: nvm, you can't get extra bandwidth by breaking spectral symmetry on a baseband signal because the signal would become complex valued.
The improvement is largely being driven by the hyperscalers (Amazon et al) demanding it.
The hyperscalers have a need for it because they have proprietary software stacks that are extremely good at pounding PCIE and bottlenecking on it.
What I'd really like to see is network equipment catching up, because I remember a few short years ago it was hit or miss if your new WiFi router would be gigabit and then it's just stayed there while we've pushed into 4K HDR streaming and game downloads over 100GB.
I'd like to see home networks hit 10GbE+ in the mainstream so I can actually start throwing around the type of data I'm consuming; with my ISP being 1Gb/1Gb (symmetric), it's no faster to dig a file out of my storage server than to just download it again.
I have no idea if the cheap £300 8-port 10GbE Mikrotik switch will do the job and I'd rather not spend any more on Ubiquiti, but I'll need a bit more than Amazon reviews to be sure whether or not to go that route in lieu of an alternative in that price range.
I'm far more concerned with wifi not dropping out than mean performance.
I thought it is worth mentioning with a dedicated Storage API and SSD on the PS5, 90+% of loading time are spent on CPU vs very little on I/O.
With faster buses and SSDs, this is probably going to get worse.
I wonder if we'll see a CPU breakthrough within our lifetimes (let's say, before 2050).
So we're going to get the same number of lanes, they'll be twice as fast (more or less), and we'll have to pay more for them.
I do see that some dual 10G ethernet cards are pci-e 2.0 x8, and some are pci-e 3.0 x4, but I can't find solid pricing for like for like to see if that does actually reduce costs.
edit: The official Compute Express Link (CXL) website has an article about the benefits of this technology [3].
[1] - https://www.nextplatform.com/2021/09/07/the-cxl-roadmap-open...
[2] - https://pmem.io/blog/2022/01/disaggregated-memory-in-pursuit...
[3] - https://www.computeexpresslink.org/post/the-benefits-of-seri...
A similar thing happened with Intel Optane (which is now dying).
You can mitigate that with some in-package HBM. IIRC, Xeon Phis came with up to 16GB of it - you could in theory have a workstation without any memory sticks.
For Optane, I'd say it's doing so badly because it's barely cheaper than RAM, which means it's only useful in niche situations.
For starters, the DRAM manufacturers may possibly be the biggest penny pinchers in the semiconductor business. They're too cheap to use differential signaling even on GDDR memory. I'm not saying that PCIe RAM is never going to happen, but as somebody who works with DRAM controllers and DRAM vendors, I can not fathom a world where DRAM manufacturers put a PCIe controller on each chip, especially one capable of Gen 6.0 speeds! The extra die area, complexity, and bringup testing necessary for a PCIe controller is insane
Second of all, the latency would really destroy system performance. The DDR protocol is designed to minimize latency (well kind of, there are still speed targets that need to be met.) The whole point of using a parallel bus with unscrambled data is to eliminate encoding/decoding latency. PCIe uses a Serializer/Deserializer (SerDes) that's scrambled on one end and then de-scrambled on the other end using a LFSR. It also goes through an adaptive tap filter. After that, the ordered sets from each lane are assembled and sent through a FIFO for clock domain crossing/timing recovery. All of those steps add latency to the system, and that's before we get to the Data Link Layer and Transaction layer. While I'm not well versed in CXL, I know that it shares the same Physical layer (a SerDes Phy.)
Now you might imagine a chip with one PCIe end point that fans out to an array of DRAM chips on the back end, and you don't care that much about latency. But then that makes PCIe RC and EP in the system redundant because you could put the DRAM controller where you would otherwise put the PCIe controller. Now it might make packaging easier and allow for a cheaper version of DDR to be implemented, but is there a net savings? and does it justify the latency hit? Maybe. I'm not saying it's never going to happen, but I wouldn't bet on that tech.
But I'm biased (I work for Intel), so take what I say with a grain of salt.
[1] - https://www.computeexpresslink.org/members
[2] - https://news.samsung.com/global/samsung-electronics-introduc...
[3] - https://www.lenovoxperience.com/newsDetail/283yi044hzgcdv7sn...
Though the renderings indicate that this is a CXL controller connected to DRAMs on the back end. It makes sense, since the front end logic for CXL is die area intensive. This is basically an SSD without all of the complexity of wear leveling.
> The extra latency is estimated to be no worse than a NUMA hop
I was almost going to mention that in my previous comment and point out that this might be a thing in enterprise applications where SRAM is plentiful, but not likely to outright replace DDR in most consumer computing devices as the parent comment seemed to suggest.
With all the benefits of an enterprise SSD (EDSFF in most cases) form factor. And that's what data center folks want.
> not likely to outright replace DDR in most consumer computing devices
Right. I expect that CXL will initially be deployed as an internal transparent optimization for cloud providers and large data centers, with the latency and complexity (to an extent) hidden at hypervisor/orchestration layers. Some sophisticated software might have special support. Eventually we might get S3-like disaggregated memory services for cloud instances enabled by pooling (pure speculation on my part).
For example, exposing this as a ram disk while possibly useful is not really optimal. Treating it more like NUMA ram, that simply lacks any cores where the memory is "local", could be more viable. The OS could be programmed to shuffle out memory that is not frequently used over to there, as a faster alternative to paging to disk, etc. Some new madvise calls to let apps mark regions as not latency critical giving the OS permission to move the memory over even if it is frequently accessed could also be interesting.
Makes it quite impressive that they're able to push 20+ Gigabit/s/pin even in dual-sided / T-layouts (though that also uses PAM4). Of course, comparing historical GPU layouts it's pretty obvious how the RAM chips are creeping ever more closely to the GPU die - weird how that is.
They seem to have most success in high end server CPUs where memory capacity and throughput are important. They have cost, power, latency disadvantages, so their forays into low end servers, PCs, etc has been more sporadic.
Obviously, these goals often coincide, and demands on hardware will occasionally shift one priority over the other (a good example being DDR5, higher latency, but higher bandwidth).
That's really the nature of comparing two levels of cache though. Going from L2 to L3, you sacrifice latency for capacity. Same thing going from L3 to L4 (RAM), and L4 to L5 (Disk). The gaps between levels may shrink over time, but we have distinct levels as a form of cost minimization. If we could cheaply stick a terabyte of L1 cache onto a consumer grade CPU, we probably would.
What do they mean by latency here? Isn't just getting data through the serdes going to take a load of clocks? And then the correction will take a few clocks.
DDR5 will cap out somewhere around 8. DDR6...might use it? It was under consideration, from what I can find.