5Gbps Ethernet on the Raspberry Pi Compute Module 4
jeffgeerling.com
jeffgeerling.com
For some reason I couldn't break that barrier, even though all the interfaces can do ~940 Mbps on their own, and any three on the PCIe card can do ~2.8 Gbps. It seems like there's some sort of upper limit around 3 Gbps on the Pi CM4 (even when combining the internal interface) :-/
But maybe I'm missing something in the Pi OS / Debian/Linux kernel stack that is holding me back? Or is it a limitation on the SoC? I though the ethernet chip was separate from the PCIe lanes on it, but maybe there's something internal to the BCM2711 that's bottlenecking it.
Also... tons more detail here: https://github.com/geerlingguy/raspberry-pi-pcie-devices/iss...
At what point are you saturating the poor little ARM CPU (or its tiny PCIe interface)?
Excellent reading on this available here :
http://www.intel.com/content/dam/doc/application-note/82575-...
and here :
https://blog.cloudflare.com/how-to-achieve-low-latency/
Edit : with the inbound 10Gb card referenced
The chip in some of the 2GB RPI4s is rated for only 3.7Gbps.
https://www.samsung.com/semiconductor/dram/lpddr4/K4F6E304HB...
https://medium.com/@ghalfacree/benchmarking-the-raspberry-pi...
LPDDR cannot sustain anywhere near the max speed of the interface. It's more of a hope that you can burst something out and go to sleep rather than trying to maintain that speed. In a lot of ways DRAM hasn't gotten faster in decades when you look at how latency clocks nearly always increase at the same rate of interface speed increases. And LPDDR is the niche where that shines the most, because it doesn't have oodles of dies to interleave to hide that issue.
[1] https://www.raspberrypi.org/forums/viewtopic.php?t=271121
Well yes because 5Gbps Ethernet is actually a thing ( NBase-T or 5GBASE-T). So 1Gbps x 5 would be more accurate.
Cant wait to see results on 10GbE though :)
P.S I really wish 5Gbps Ethernet is more common.
1: https://www.amazon.com/UGREEN-Ethernet-Thunderbolt-Converter... (USB-C)
2: https://www.amazon.com/2-5GBase-T-Ethernet-Controller-Standa... (Don't get the knock-off version of this, the brackets aren't the right sizes.) (PCIe)
The expensive ones I'm waiting to arrive:
3: Either a second hand Intel X520-DA1 card or the "refurb" from AliExpress
and https://mikrotik.com/product/crs305_1g_4s_in with RJ10 SFP+ modules. Then cry at how much you just spent.
edit: Oh its 3Gbit across 5 interfaces, one of which isn't PCIe, so the PCIe side is probably only running at about 50%. It might be interesting to see if the CPUs are pegged (or just one of them). Even so, PCIe on the rpi isn't coherent so that is going to slow things down too.
edit: run `perf top` to see if that gives you a better idea.
15.96% [kernel] [k] _raw_spin_unlock_irqrestore
12.81% [kernel] [k] mmiocpy
6.26% [kernel] [k] __copy_to_user_memcpy
6.02% [kernel] [k] __local_bh_enable_ip
5.13% [igb] [k] igb_poll
When it hit full blast, I started getting "Events are being lost, check IO/CPU overload!"You have to set the irq affinity to utilize the available CPU cores.
There is a script included with the source you used to compile drivers called "set_irq_affinity"
Ex (Sets IRQ Affinity for all available cores) :
[path-to-i40epackage]/scripts/set_irq_affinity -x all ethX
I wish I had the cycles and the kit on hand to play with this!
I would also suggest using taskset[1] on each iperf server process to bind them each to a different cpu core.
Finally, I would suggest UDP on iperf and let the sending Pi's just completely saturate the link.
If you do all that, I think you have a good chance at achieving 3.5Gbps over just the Intel card.
This is very likely the answer. I see a lot of people who think of the Pi as some kind of workhorse and are trying to use it for things that it simply can't do. The Pi is a great little piece of hardware, but it's not really made for this kind of thing. I'd never think about using a Raspberry Pi if I had to think about "saturating a NIC".
But I like to know the limits so I can plan out a project and know whether I'm safe using a Pi, or a 3-5x more expensive board or small PC :)
First off, thank you for doing this kind of 'r&d', it is really exciting to see what the Pi is capable of after less than a decade.
Would you be interested in someone testing a SAS PCI card? I'm going to pick up one of these as soon as they're not backordered...
https://www.broadcom.com/products/ethernet-connectivity/netw...
They don't use copper, they use fiber. It wouldn't be a mystery if you searched for '100gbs pci ethernet'.
(He also mentioned 390MB/sec write speed to nvme, which is suspiciously close to the same ceiling)
Note that combining the internal interface with the 4 NIC interfaces, and overclocking to 2.147 GHz got it up to 3.4 total Gbps. So the IRQ interrupts are the main bottleneck when it comes to total network packet throughout.
> I think the PCIe link hits a ceiling around there.
You're trying to shove 10 gallons of shit into a 5-gallon bucket!
--
I'm not sure how high you can set the MTU on those Pi's (the Intels should handle 9000) but I'd set them as high as they'll go, if I were you. An MTU of 9000 basically means ~1/6th the interrupts.
No, it is not. That NIC is a PCIe Gen2 NIC. By using only a single lane, you're limiting the bandwidth to ~500MB/sec theoretical. That's 4Gb/s theoretical, and getting 3Gb/s is ~75% of the theoretical bandwidth, which is pretty decent.
I mean, before this the most I had tested successfully was a little over 2 Gbps with three NICs on a Pi 4 B.
What direction are you running the streams in? In general, sending is much more efficient than receiving ("its better to give than to receive"). From your statement that ksoftirqd is pegged, I'm guessing you're receiving.
I'd first see what bandwidth you can send at with iperf when you run the test in reverse so this pi is sending. Then, to eliminate memory bw as a potential bottleneck, you could use sendfile. I don't think iperf ever supported sendfile (but its been years since I've used it). I'd suggest installing netperf on this pi, running netserver on its link partners, and running "netperf -tTCP_SENDFILE -H othermachine" to all 5 peers and see what happens.
Modems used to do this too. The 'cheat' is that they report Layer 1 bandwidth, which is a completely useless number to the end user. The bulk of the loss occurs between Layer 1 and Layer 2 (with dribs and drabs for packet headers and so forth)
See: https://github.com/geerlingguy/raspberry-pi-pcie-devices/iss...
You can also use ethtool -C on the NICs on both ends of the connection to rate limit the irq signal handeling allowing you to optimize for throughput instead of latency.
According to the Intel 82580EB datasheet[1] it supports an MTU of "9.5KB." It's unclear if that means 9500 or 9728 bytes.
I looked briefly for a datasheet that includes the ethernet specs. of the Broadcom 2711 but didn't immediately find anything.
Recent versions of iproute2 can output the maximum MTU of an interface via:
# Look for "maxmtu" in the output
ip -d link list
Barring that you can try incrementally upping the MTU until you run in to errors.The MTU of an interface can be set via:
ip link set $interface mtu $mtu
Note that for symmetrical testing via direct crossover you'll want to have the MTU be the same on each interface pair.[1] https://www.intel.com/content/www/us/en/embedded/products/ne... (pg. 25, "Size of jumbo frames supported")
Get a USB 3.0 2.5G or 5G card. With a fully functional DMA on the USB controller it can get quite close to PCIE option.
A setback for all Linux users at the moment:
The only chipmaker making USB NICs doing 2.5G+ is RealTek, and RealTek chose to use USB NCM API for their latest chips.
And as we know Linux support for NCM now is super slow, and buggy.
I barely got 120megs from it. Will welcome any kernel hacker taking on the problem.
QNAP QNA-UC5G1T uses Marvell AQtion AQC111U. Might be worth a try.
Why not loop the ports back to themselves? IIRC, 1gbit ports should autodetect when they're cross connected so it wouldn't even need special cables
How to force the packets through the external wires depends on the operating system. On Linux you must use namespaces and assign the two Ethernet interfaces that are looped on each other to two distinct namespaces, then set appropriate routes.
I might be making things up, but I believe you can also run code on DPDK nics? i.e. beyond straight networking offload. If that's the case you could try compressing the data before you DMA it to the nic. This would make no sense normally, but if your bottleneck is in fact the pcie x1 link and you want to saturate the network, it would be something worth trying.
I mean really the whole thing is at most a fun exercise as the nic costs more than the pi.
Overclocking can get the total net throughput to 3.4 Gbps.