How PCI-Express works (2020)
ovh.com
ovh.com
So in theory one could install a GPU or HBA in one of those x1 slots on the motherboard if one isn't terribly concerned about bandwidth. Except most of the motherboards use x1 connectors which doesn't have an "open back" for larger cards... so much for that[1].
Another thing I didn't see it mentioning was this. Say you have a x16 slot on your motherboard, but you really want to install 4 NVME m.2 drives. Well each of those need a x4, but you have an x16 so per above you should be fine right, just plop in a "dumb" four socket m.2 PCIe card?
Well, no, not unless the motherboard supports bifurcation[2]. If not, you'll need a card with a PCIe switch which can turn those four x4's into one x16.
[1]: yes I know about risers
Also don’t use a Dremel or even a file a hobby knife or flush cutters work just fine. If you have a steady hand then a hot knife attachment for a soldering iron is fine too. The cutting wheel on a dremel tool can run off and the dust it produces can be conductive and short something.
(People are very very good with pretending at intelligence and bad at the practice of it)
You could also use a razor blade or other such manual cutting devices...
Absolutely hilarious to me; I believe they were running one direction at 115.2kbps and everything was just a-ok about it.
Is that true in practice or is it like USB where manufacturers can wildly violate specs and just expect consumers to deal with it?
Although graphics cards always generally should go in a 16x slot for best performance, they all work in the external 4x, even the latest most high end power/bandwidth hungry GPU parts. performance may suffer in some workloads though due to reduced bandwidth of the 4x slot though.
You can stick any PCI express device in an external GPU enclosure, will pretty much “just work” despite being a 4x slot - these things are really external PCIe enclosures despite the GPU focus.
Why was 3 capped? Nobody seems to know.
It's not so simple, there's another factor other than the number of lanes: power. According to Wikipedia (https://en.wikipedia.org/wiki/PCI_Express#Power), a x1 slot is limited to 25W on the 12V pins, while a x16 slow has a higher limit of 66W; the power pins are the same, but the maximum current is higher (2.1A vs 5.5A).
On the enterprise side of things loading up a board with bootable raid and 10g cards will lead to situations where the raid card needs to go topologically first or early so the raid bios has memory to load.
File:PCIe J1900 SoC ITX Mainboard IMG 1820.JPG - Wikipedia https://en.wikipedia.org/wiki/File:PCIe_J1900_SoC_ITX_Mainbo...
Because they bothered to read the PCIe spec.
Open back slots are against the PCIe specification ("up plugging" is the term IIRC). That is for two (three?) reasons, first its a mechanical thing, if you noticed plugging a big heavy x16 GPU into a x1 slot is an exercise in getting it just right. Any bump/whatever will move the card into a bad angle, even if its screwed into the back panel. Second, in the past, with older pcie specs the amount of current a slot was required to supply was dependent on its size. So a big x16 card might need the full slot rating while a smaller slot might only provide a smaller amount of current. (IIRC the limits were something like 75W for a x16 and only 25W for a x1). Thirdly, down plugging _IS_ part of the spec and the correct way to handle this situation. That means the MB vendor provides a larger, say x16, mechanical slot which is provides a fraction of the signal lanes, say x4.
So, the vendors which were providing cut back slots were either ignorant of the spec, too cheap to pay the couple extra cents for the larger connector, or had some fundamental constrain on providing it and were willing to provide a non-conformant part (generally unlikely). I've seen these cut back slots a few times in parts of the arm ecosystem where the vendor can't be bothered to read the spec much less implement it correctly. This is also how one has the pile of problems that are frequently seen in the SBC market where the boards won't actually work with some PCIe card, USB device, whatever because the HW/SW is only implementing the convenient parts the relevant standard.
I tried this with a SAS controller. They make 16-port SAS PCI-e x8 cards, so on a x1 slot I should get full speed for 1-2 SAS devices, and the SAS drive(s) will only slow-down if I add more... right?
Wrong. Even with a single device, you get about 1/8th the speed you should. So, yeah, PCI-e cards properly downgrade to fewer lanes, but not necessarily in a very useful way.
I think this should really read that each lane has two differential pairs (interestingly the diagram shows twisted pairs). One for inbound, one for outbound. Which sums to four wires (or traces) in total for each lane.
I’m also a little confused about the whole
> A lane is composed of 2 wires: one for inbound communications and one, which has double the traffic bandwidth, for outbound.
Not really sure what the author is trying to say. But the ideas that outbound has double the bandwidth doesn’t really make much sense, as outbound is different depending on perspective as a device. So I read this statement as saying that each device transmits at twice the speed the receiving device can receive data. Which is clearly non-sensical.
But yeah, I spotted the "two wires" stuff and a bunch of other problems. Credit for enthusiasm, perhaps, but it's not terribly accurate.
Based on other articles by the author https://www.ovh.com/blog/understanding-the-anatomy-of-gpus-u... it seems all of them are quite terrible word salads with no substance.
I might be wrong, but I believe PCIe lanes are differential. Meaning you have to have 2 pairs of wires (1 for each direction), not 1 totaling 4 physical wires per lane.
Also both pairs in the lane should have the same bandwidth but since they are "full duplex" system, you might say it doubles the overall bandwidth of the lane but it does not mean that one of them is twice as fast.
It literally starts with a discussion about add-in cards form factors like it even matters. Then has tons of blunders, like the full/half duplex one, calling transfers / second frequency which it isn’t how many transfer a second has little to do with the clock frequency because it’s dependent on many other factors such as the width of the channel and how many times per clock cycle you can transfer data. It doesn’t even begin to go over the basic concepts of the PCIe root complex and the structure of a switched fabric.
Then I has some gems like “having a nice GPU with 16 PCIe Lanes and having a CPU with 8 PCIe Bus lanes will be as efficient as throwing away half your money because it doesn’t fit in your wallet.” which really isn’t true…
The data is transmitted and received as differential pairs because this helps in signal integrity issues.
You are also correct that you can transmit and receive at the same time but whether this helps our now depends on what you are running.
The Serdes used for PCI-E are also used for other protocols like SATA and USB3. They are extremely similar at the physical layer but the higher layer protocols are very different.
USB 1/2 has a single differential pair so you can only transmit or receive but not at the same time.
USB 3 has differential pairs for RX and TX so you can do both at the same time.
Just in case this leaves some readers thinking it's only possible to transmit or receive on a differential pair, not so:
Copper gigabit ethernet (1000Base-T) transmits and receives at the same time over all four differential pairs, using adaptive equalisation and echo cancellation to separate the outgoing and incoming signals on each pair.
The signal processing is quite power hungry, and doesn't work so well at higher speeds. So it makes sense that PCIe wouldn't do this.
http://xillybus.com/tutorials/pci-express-tlp-pcie-primer-tu...
http://xillybus.com/tutorials/pci-express-tlp-pcie-primer-tu...
These days all of these interconnects (and high speed ether) are in the process of a bit of a convergence into what essentially just versions of raw serdes interfaces on chips - the same wires coming off the die can be 10Gb ether, PCIE, etc etc - they all just self clocked differential pairs
A USB3 port is already pretty dangerous: you can plug in something that will generate keystrokes or mouse movements and also present storage, so a malicious device can mount itself, copy over a payload, run it, and then pretend to be a cup-warmer again.
Plug in a PCIe device and it gets to control your system.
System compatibility
Kernel DMA Protection requires new UEFI firmware support. This support is anticipated only on newly-introduced, Intel-based systems shipping with Windows 10 version 1803 (not all systems).
So, it's CPU specific, motherboard specific, firmware specific, and OS version specific.
That's not really solved, is it?
Anyway mitigation is considered to enough TB/USB4 to be adopted for new devices.
The problematic part is that the UEFI needs to support it, it seems most systems with Thunderbolt have enabled it since 2018 and systems without Thunderbolt still don't bother.
It's mandatory for thunderbolt 4.
* NVMe 2.0 going to support HDD
* Thunderbolt, USB 4.0 supports PCIe based connection
* CFexpress, SD express is based on PCIe
The only reason SATA fell behind on a per-lane basis is because they stopped updating it.
PCIe has no special sauce for running faster, just more lanes. It's taking over drives, but mostly for other reasons.
DVI was replaced by HDMI. HDMI 2.1, which came out in 2017, is about 2/3 as fast per lane as PCIe gen4, which also came out in 2017. For an external passive cable that's really good.
USB 3.1 jumped to 1.2GB/s in 2013, which means it was faster than PCIe for four years. And USB4 is twice that speed per lane, making it faster than PCIe gen4.
So the constraining factor is not the type of signalling, it's the number of wires and transceivers. You can make anything go faster with more wires, but that's expensive.
But the big issue you face is that PCIe is fast because the physical layer is so tightly controlled. Lane lengths to the same port are all approximately the same length. All the lanes are carefully shielded with a complex mix of grounded traces and large ground planes above and below. The length of the traces is limited to make sure the speed of light doesn’t introduce too much delay.
So PCIe speed doesn’t come from fancy algorithms, but from fancy electronics, and a very tightly controlled physical layer. Extending that layer outside of a computer is tricky.
Just like trains are much faster than cars, PCIe is much faster the DisplayPort. But trains require tracks, that are long and straight, and car don’t. Equally PCIe requires electrical connections that are short and straight, and DisplayPort doesn’t.
Thunderbolt kinda splits the difference. It’s basically PCIe, but heavily downrated so it can handle being outside of a motherboard.
Thunderbolt goes for even longer cable runs by using active cables, while long PCIe runs use redrivers or retimers that aren't integrated into the cables. Thunderbolt is also not really a downrated PCIe in any way; Thunderbolt transmits data faster than PCIe 4.0 lanes, but doesn't scale to as many lanes as internal PCIe links.
It allows the CPU (or any other device in the PCI bus) to write/read data to/from the device when it’s convenient for the sender, without having to interrupt whatever the receiving device is currently doing.
You still need to coordinate with the device to tell it that you’ve written to its memory, or read from it. But that’s a pretty cheap operation.
The alternative is the both CPU and GPU would have to stop what their doing and manage the data copy while doing nothing else.
So it’s basically the difference between sending someone an email vs giving them a call.
With DMA your emailing a big document, the later calling to make sure they received it. Without DMA it would be like calling them and reading the document out over the phone. One is clearly better for everyone’s productivity.
I could still read this two ways however: one where the memory is on the peripheral, and one where the memory is main memory, where the peripheral is copying to/from using DMA. Which one is it?
You send the device a circular list of descriptors (pointers) to a region of main memory.
In order to send data to the device, you write your network packet to the memory region associated with the pointer of the current ‘head’ of the descriptor list.
So far, you have a ring of pointers, one of those pointers points to a location you just wrote to in ram.
You then tell the device that the head of the list has changed (as you just wrote some data to the region that the head of the list is pointing to - so it can consume that pointer), the device then goes ahead and copies the data from ram into an internal buffer on the card. Once the data is consumed, the tail pointer of the ring buffer is updated to indicate that the card is finished with that memory region.
> __padding 45 minutes ago [dead] [–]
> Typically with devices like network cards (that also operate over PCI-E) You send the device a circular list of descriptors (pointers) to a region of main memory. In order to send data to the device, you write your network packet to the memory region associated with the pointer of the current ‘head’ of the descriptor list. So far, you have a ring of pointers, one of those pointers points to a location you just wrote to in ram. You then tell the device that the head of the list has changed (as you just wrote some data to the region that the head of the list is pointing to - so it can consume that pointer), the device then goes ahead and copies the data from ram into an internal buffer on the card. Once the data is consumed, the tail pointer of the ring buffer is updated to indicate that the card is finished with that memory region.
But equally a peripheral can expose its own memory and ask the host to write into it.
Cheap devices tend to do the former because it avoids the need to have expensive memory built in. They can just “borrow” system memory. More expensive, performance optimised, devices tend to do the latter.
It’s also worth mentioning that DMA tends to work between every device attached to the PCIe bus. So Microsoft’s DirectStorage API seems to be using this feature, by having the GPU directly read data from an SSD, without the data ever touching the CPU or main memory.
It's interesting to compare to embedded processors without a memory management unit, like this STM32 reference manual, see p.68 and following:
https://www.st.com/resource/en/reference_manual/dm00124865-s...
Everything looks like a memory address. Note that it's not actually memory, the processor just diverts requests for that address to the peripheral instead of memory. But on that little ARM processor, if you want to write to actual RAM, that's memory addresses 0x2001 0000 to 0x2001 BFFF. Data in the onboard Flash memory is at 0x0020 0000 - 0x002F FFFF. If you want to talk to something on a serial port, write to registers from 0x4001 1000 - 0x4001 13FF. If you want to show something on an attached LCD, or pull a buffer from the USB or Ethernet peripherals, or work with GPIO, or do anything at all, really, it's at some memory offset. This chip has some DMA, you can set it up to automatically push from one peripheral memory space to actual RAM or vice versa. But everything happens at a region of memory.
This is perhaps the DMA scenario you describe in the last two lines, my thinking is that it would make sense to do this all the time, at least when transfers are large.
As others have said, you'd normally configure the MMIO space for uncached access, or you'd need to be careful to force the memory ordering you need. The device specific interfacing requirements would be the guide there. Devices can indicate if their MMIO ranges are prefetchable or not, which should indicate if stray reads would cause side effects or not.
One bonus of MMIO is DMA could interface with other devices, whereas I don't think devices are allowed to drive the I/O bus like that.
It ends up being the same as any other interface: there's a limit to how many devices can be connected before the limit for non-blocking simultaneous bandwidth is reached, but that merely means that bandwidth can be a bottleneck if you go above that.
This holds especially true for consumer motherboards and PCIe breakout boards.
Multiplexers just lets you switch between connected devices - it won't let you use both simultaneously. It's hot-swapping without having to physically change devices.
And yes multiplexing is basically “time sharing” and the PLX chips operate on the physical layer only. PCIe switches have an internal bus, buffers and actual decode the packets to know where to send them too, they can also mediate between different versions of PCIe connected to the same switch so connecting a PCIe 2.0 device to a switch would not impact other 3.0/4.0 devices whilst a multiplexer would always operate at the lowest “speed”.
But PLX chips were pretty much the only thing you can get on consumer motherboards at least back when multi GPU setups were common and SLI/XFire motherboards we’re a thing (that said PLX did help quite a bit more with XFire setups than SLI due to the lack of an external cross bridge between GPUs in later XFire revisions).
The chipset on your motherboard can also support its own PCIe lanes however at least on Intel chipsets it’s not a classical PCIe switch its closer to the silicon in the CPU that runs the root complex.
Massive amounts of I/O such as this are instrumental for high and extremely dense I/O applications whether it be for storage, network, GPUs, etc (or combinations of all, of course). I know it's been a consideration for companies such as Netflix, Cloudflare, and Nvidia.
However, AMD Epyc uses half of the 128 lanes on each socket to talk to the other socket, so on two-socket systems each socket has only 64 PCIe lanes available, for a total of 128 lanes, the same as with a single socket. On the other hand, newer AMD Epyc processors can (depending on the motherboard) use only 48 lanes to talk to the other socket, freeing another 16 lanes per socket (for a total of 160 PCIe lanes) at the cost of slower inter-socket communication; and they also have an extra single PCIe lane per socket to be used to connect to a BMC. The article at https://www.servethehome.com/why-amd-epyc-rome-2p-will-have-... has a great explanation of all that.
I still remember building a two socket Xeon workstation some years ago and being puzzled that video wasn't initializing. Turns out one CPU wasn't quite seated correctly and the GPU was in a slot wired to it. I wonder if this architecture avoids that?
It's more than just that. The memory is attached directly to each socket, so for instance you could have 64GB of memory on each socket for a total of 128GB; for a core in one socket to access memory which happens to be attached to the other socket, it has to go through these inter-socket links. More than that, for a core in one socket to access memory in the same socket but which has been cached somewhere in the other socket, the cache coherence traffic has to go through these inter-socket links.
> Turns out one CPU wasn't quite seated correctly and the GPU was in a slot wired to it. I wonder if this architecture avoids that?
No, each PCIe link is wired to only one of the CPU sockets (PCIe is point-to-point, not a bus like classic PCI), or to an auxiliary chip which is wired to only one of the CPU sockets; if that CPU is not seated correctly, what you saw could happen. The architecture which avoids that is the older one in which all CPUs were wired together in a bus, with PCI and memory attached to an auxiliary chip also on that bus.
For people that aren't reading the whole article, I want to make it clear that "slower" is only relative to other newer chips. The lanes on the old chips were half as fast, so even 48 lanes beats them by a factor of 1.5x.
There's a reason gen4 took so long to roll out - it has been finalized for 4 years (with he manufacturers working on support for much longer), and it's still not ubiquitous. I'm sure you'd rather have a gen4 or gen5 board when gen6 is rather than be stuck on gen3.
(Plus, for gaming loads there are diminishing returns - outside reducing texture upload times, GPUs in gaming scenarios do not use a lot of PCIe bandwidth.)
What's nice is that it still keeps everything good about PCIe - like the electrical and mechanical spec, but does away with the old legacy things that nobody likes about PCIe. However, I'm not super familiar with CXL, so maybe it just looks nice from far away.