Why haven't we seen the need for a PCI-X or VLB-style PCIe interface expansion?
Why haven't we seen the need for a PCI-X or VLB-style PCIe interface expansion?
>To achieve its impressive data transfer rates, PCIe 7.0 doubles the bus frequency at the physical layer compared to PCIe 5.0 and 6.0. Otherwise, the standard retains pulse amplitude modulation with four level signaling (PAM4), 1b/1b FLIT mode encoding, and the forward error correction (FEC) technologies that are already used for PCIe 6.0. Otherwise, PCI-SIG says that the PCIe 7.0 speicification also focuses on enhanced channel parameters and reach as well as improved power efficiency.
So it sounds like they doubled the frequency and kept the encoding the same. PCIe 6 can get up to 256 GB/s, and 2x 256 = 512.
In any case, it'll be a long time before the standard is finished, and far longer before any real hardware is around that actually uses PCIe 7.
"just double the frequency" isn't something we're used to seeing elsewhere these days (e.g. CPU clock speeds for the last couple of decades). What are the fundamental technological advances that allow them to do so? Or in other words, what stopped from from achieving this in the previous generation?
It wasn't really like it was impossible before now - it's just more that it wasn't in demand. With the proliferation of SSDs transferring data over PCIe, it's become much more important - so the extra cost of better signaling hardware is worth it.
Not to dismiss it completely, it's still a hard problem. But it's far easier than doubling the frequency of a CPU.
Explaining it in detail requires more background in electronics, but that's ultimately what it boils down to.
High-end analog front ends can reach the three-digit GHz (non-silicon processes admittedly, but still).
Please do
1) transistor size (smaller = faster, but process is more expensive)
2) gate dielectric thickness (smaller = faster, but easier to damage the gate; there's a corollary of this with electric fields across various PCB layer counts, microstrips, etc. that I'm nowhere near prepared to articulate)
3) logic voltage swing (smaller = faster, but less noise immunity)
4) number of logic/buffer gates traversed on the critical path (fewer = faster, but requires wider internal buses or simpler logic)
5) output drive strength (higher = faster, but usually larger, less efficient, and more prone to generating EMI)
6) fanout (how many inputs must be driven by a single output) on the critical path (lower = faster)
Most of these have significant ties to the manufacturing process and gate library, and so aren't necessarily different between a link like PCIe and a CPU or GPU. Some things can be tweaked for I/O pads, but broadly speaking a lot of these vary together across a chip. The biggest exceptions are points #4 and #6. Doing general-purpose math, having flexible program control flow constructs, and being able to optimize legacy instruction sequences on-the-fly (but I repeat myself) unavoidably requires somewhat complex logic with state that needs to be observed in multiple places. Modern processors mitigate this with pipelining, which splits the processing into smaller stages separated by registers that hold on to the intermediate state (pipeline registers). This increases the maximum frequency of the circuit at the cost of requiring multiple clock cycles for an operation to proceed from initiation to completion (but allowing multiple operations to be "in flight" at once).
That being said, what's the simplest possible example of a pipeline stage? Conceptually, it's just a pair of 1-bit registers with no logic between the input of one register and the input of the next one. When the clock ticks, a bit is moved from one register to the next. Chain a bunch of these stages together, and you have something called a shift register. Add some extra wires to read or write the stages in parallel, and a shift register lets you convert between serial and parallel connections.
The big advantage that PCIe (and SATA/SAS, HDMI, DisplayPort, etc.) has over CPUs is that the actual part hooked up to the pairs that needs to run at the link rate is "just" a pair of big shift registers and some analog voodoo to get synchronization between two ends that are probably not running from the same physical oscillator (aka a SERDES block). In some sense it's the absolute optimal case of the "make my CPU go faster" design strategy. Actually designing one of these to reliably run at multi-gigabit rates is a considerable task, but any given foundry will generally have a validated block that can be licensed and pasted into your chip for a given process node.
Does that make sense?
Transmitting and receiving high-speed data is actually mostly an analog circuits problem, and the circuits involved are very different than those in doubling CPU speed.
Things like packet encoding etc. Then a quick look at the signalling change of NRZ vs PAM4 in later generations.
Gen1 -> Gen5 used NRZ, PAM4 is used in PCIe6.0.
[0] Understanding Bandwidth: Back to Basics, Richard Solomon, 2016: https://www.synopsys.com/blogs/chip-design/pcie-gen1-speed-b...
They made a significant signalling change once, with 6. How did they manage to take the baud rate from 5 to 8 to 16 to 32 GHz?
PCIe 1.0 & PCIe 2.0:
Encoding: 8b/10b
PCIe 2.0 -> PCIe 3.0 transition:
Encoding changed from 8b/10b to 128b/130b, reducing bandwidth overhead from 20% to 1.54. Changes here in the actual PCB material to allow for higher frequencies. Like changing away from PCB material like FR-4 to something else [2].
PCIe 3.0, PCIe 4.0, PCIe 5.0:
Encoding: 128b/130b
There is plenty to dive deep on, things like:
- PCB Material for high-frequency signals (FR4 vs others?)
- Signal integrity
- Link Equalization
- Link Negotiation
Then decide which layer of PCIe to look at:
- Physical
- Data / Transmission
- Link Layer
- Transaction
A good place to read more is from the PCI-SIG FAQ section for each generation spec that explains how they managed to change the baud rate as you mentioned.
PCI-SIG, community responsible for developing and maintaining the standardized approach to peripheral component I/O data transfers.
PCIe 1.0 : https://pcisig.com/faq?field_category_value%5B%5D=pci_expres...
PCIe 2.0 : https://pcisig.com/faq?field_category_value%5B%5D=pci_expres...
PCIe 3.0 : https://pcisig.com/faq?field_category_value%5B%5D=pci_expres...
PCIe 4.0 : https://pcisig.com/faq?field_category_value%5B%5D=pci_expres...
PCIe 5.0 : https://pcisig.com/faq?field_category_value%5B%5D=pci_expres...
PCIe 6.0 : https://pcisig.com/faq?field_category_value%5B%5D=pci_expres...
PCIe 7.0 : https://pcisig.com/faq?field_category_value%5B%5D=pci_expres...
[0] Optimizing PCIe High-Speed Signal Transmission — Dynamic Link Equalization https://www.graniteriverlabs.com/en-us/technical-blog/pcie-d...
[1] PCIe Link Training Overview, Texas Instruments
[2] PCIe Layout and Signal Routing https://electronics.stackexchange.com/questions/327902/pcie-...
>> To achieve its impressive data transfer rates, PCIe 7.0 doubles the bus frequency at the physical layer compared to PCIe 5.0 and 6.0. Otherwise, the standard retains pulse amplitude modulation with four level signaling (PAM4), 1b/1b FLIT mode encoding, and the forward error correction (FEC) technologies that are already used for PCIe 6.0.
Nothing else changed, they didn't move to a different encoding scheme. PCIe 6.0 already uses PAM4. Unless they moved to a higher density PAM scheme (which they didn't), the only way to increase bandwidth is to increase speed.
The future NVIDIA B100 might be PCIe 6.0 but hopefully will support 7.0 and maybe NVIDIA (or someone) gets a NIC working at those speeds by then...
That is, they recon long traces on a motherboard just won't cut it for the strict tolerances needed to make PCIe 7.0 work.
[1]: https://www.amphenol-cs.com/connect/news/amphenol-released-p...
The standard is expected to finish in 2025, and hardware / IP for PCIe 7 are already in the work. Since there is so much resemblance of it and PCIe 6. In terms of schedule it perhaps may be the least lead time of recent PCIe from Standard 1.0 to IP available. The HPC industry is really pushing for this ASAP.
It's not really the same physical interface. The connector is the same, but quality requirements for the traces have gotten much more strict over time.
> Why haven't we seen the need for a PCI-X or VLB-style PCIe interface expansion?
x32 PCIe does exist, it's just rarely used.