Faster Linux 2.5G Networking with Realtek RTL8125B
jeffgeerling.com
jeffgeerling.com
One thing to keep in mind about the CRS305 is that it can't handle more than one or maybe two SFP+ copper modules. They use a lot of power and put out a lot of heat. I think the CRS309 has fewer restrictions.
For the fiber ONT (what I realize now you may be calling their fiber modem) I would love to see a similar attempt to get their hardware out of the loop.
They can work in that configuration (with up to 4 in that enclosure) but I'd only do it if you have a fan on the thing. It relies entirely on passive ventilation, and the enclosure gets hot, too, like 60°C or more!
I haven't bothered to do much actual diagnosis on this because I'm satisfied with the speed as-is, but it's another illustration of weirdness where you'd least expect it.
The modern versions (ie: ConnectX 6) are much more expensive of course, but the older hardware might do better with FreeBSD / Linux anyway. Of course, your mileage may vary so do some research before buying (or buy for $30 and hope for the best).
You'll also either want Optical Transceivers + Fiber, or Active Fiber (aka: Fiber that has transceivers soldered onto them already), or DAC (direct-attach copper, the cheapest but only works over small distances).
Optical Transceivers + Fiber is the most flexible of course, different transceivers are spec'd for different wires and different costs though, so it gets a bit complicated.
Active Fiber and DACs are is probably easiest, but if you buy a 10m cable but need like 12m, you need to buy a new cable + transceiver combo (rather than just purchasing a bunch of different fiber lengths).
Length and cost is the big difference from Active Fiber and DAC, but otherwise are very similar in use. You plug the modules into the SFP+ port, hope its compatible and let them rip.
I'd start there and if you find you need other features, think about spending more.
If you want to go all-out fast and not need a switch you can use a QSFP+ <-> 4xSFP+ cable or splitter by getting a QSFP+ card on the server end.
Ethernet is more costly, SFP is the more approachable-but-unusual way. Lengths of the runs and how many clients changes the decision making a bit, so keep that in mind.
I didn't mind cost so much but I wanted to stick with typical Ethernet. Mainly a comfort thing, I don't know the limits of SFP well.
I went with this switch: Mikrotik CRS312-4C+8XG
... and these cards: ASUS XG-C100C
I upgraded my home lab to 10GbE around a year ago mainly to speed up my SSD-based backups. All told I think it was about $800 to upgrade my storage cluster and a couple clients.
We support 10G-BaseT as well for devices that are 10G-PoE-BaseT but otherwise standard here is SFP+. Gets even cheaper when you consider power consumption, heat dissipation, and multi-port cards (2xSFP+ NICs are extremely affordable and nice when you have segregated networks or VM hosts).
EDIT: 10G-BaseT also has ~2 orders of magnitude higher latency compared to SFP+ Fiber, and ~1 order of magnitude higher latency than SFP+ DAC. The numbers are small across the board, but it's relevant. 5-10W draw for 10GbE compared to 1-2W draw for SFP+ Fiber, for power consumption numbers.
What am I missing?
Redundancy is an easy way to spend a ton of cash poorly, but if done well I recommend it for anyone!
The virtualization aspect for example, that's just a bullet-point for how I use KVM and some of the fancier gear to prove out certain implementations... then automate using them. Lately, SR-IOV.
For those like me, it's often fairly practical but the trivial stuff may be what gets shared
Does anyone know why they are not contributing to the kernel directly? (Just a guess: their drivers are from the same codebase as the Windows driver and use a bunch of abstraction layers that wouldn't be acceptable in Linux)
[1] https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
[2] https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
[0] https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
https://www.realtek.com/en/component/zoo/category/network-in...
Supported mainline by the r8169 driver.
> Use Intel if you want working hardware
Sadly Intel totally dropped the ball on this.
i225-V (Foxville) NIC early steppings were awful. And no driver can fix that, it required Intel to produce a new stepping... countless motherboards are affected by that one...
Realtek NICs are fine at this point in time.
They worked fine for 10 Gbps, but I had to switch to Mikrotik transceivers for 2.5G connections.
Meanwhile there are notable recent perf speedups in the modern ethernet too: https://git.kernel.org/pub/scm/linux/kernel/git/netdev/net-n...
Tl;Dr 1/2.5gig works for consumers needs, 10gig is not price competitive for our needs.
> If it was available, there would be apps to use it.
To an extent, since the inverse is also true: there is as of yet no market demand for such large amounts of bandwidth either.
Once popular apps/usecases with heavy bandwidth requirements see wide adoption (at first limited to specialty high bandwidth networks), we will start seeing growth in the consumer bandwidth space.
This is pretty expensive considering that you can get a 10G switch instead for around $40 more.
10 -> 100 -> 1000 -> 10G -> 100G
I know there are some oddball 25G, 40G, and 2.5G out there too, but always wondered this.
I suspect it's about us having 10 fingers and using arabic numerals.
Each generation gets introduced with new switches, new NICs, and new cabling. Since 100Mb it has been possible to have link aggregation groups, where two or more cables on the same number of NICs on each side are bundled together into one bandwidth device. At some point you need to decide that it's worth doing the transition -- and 10x is a pretty good number for that.
The recent oddballs are 50G, 40G, 25G and 2.5G. All of them come from the idea of taking a higher speed NIC and adding a little more hardware to get multiple PHY transceivers -- a 100G becomes 2x50 or 4x25, and a 10G becomes 4x2.5.
Oddly, 40G comes from 4x10G in the other direction, and is more expensive to produce than 50G equipment. It's not very popular.
I started to write that there are $30 single-fiber-pair CWDM pluggables for 40GB making it the fastest speed you can do cheaply at distances longer than realistic for DAC cables... but checking ebay I see that there are now 100G pluggables for $39, so maybe its time to update the last of my 40G hosts to 100G.
40G QSFP is just 4 lanes of 10G and can usually be broken out into separate ports
Same for 100G QSFP28 (4 lanes of 25G SFP28)
You see this all the time with interfaces. Like with PICe bifurcation. Or how 56G infiniband is 4 lanes of 14G tho you can’t split it, etc
https://en.m.wikipedia.org/wiki/SerDes https://en.m.wikipedia.org/wiki/XAUI
First 10G was 4x3.25G SerDes, so 4 traces to route on the board to the switch chip. There is some overhead in the signaling so you get 10G. If you plugged in a 1G, it just used one lane. Move to 40G, each lanes speed went up, 12G IIRC. At this point there was room for MAC/PHY so you go break each out to 4x10G. 256 traces to route to switch ASIC. Board routing is black magic. PHYless (no separate physical PHY chip required each set of 4 traces to be the same length down to the nm). Arista 7050 was first switch that shipped this. First 100G was 10 lanes, but lots of traces so smaller number of ports. Then they got the lanes up to 28G so you got 100G port for 4 lanes again, or 4x25G. So paired with MAC/PHY you could get 4x25G, 2x50, etc. and so on and so on as the speeds on the SerDes goes up. This is a simplified write up but mostly correct.
10G requires upgraded cabling, and lower cost 10G devices may use SFP+, but most consumer gear uses RJ45.
And as far as I understand it 10G on cat5e will work most of the time as long as you're not going super far.
NCM protocol is very lightweight, but has support for DMA, and some offloading
This is what happens when the kernel community is confrontational and generally a pain in the neck to deal with. After enough friction caused by the netdev folks, the vendor just gives up on upstreaming and self-publishes their driver.
But honestly, Drew, all serious vendors of networking hardware have fully functional upstream drivers and good relationships with Jakub and Dave. However, one can't just throw mess of a code to netdev and expected it to be accepted as-is because they are a big vendor or something. Which I believe is good for users!
In my case: Company A submits a driver with a feature in their driver that basically triples performance for our class of device. It was implemented in a horrific, unreadable way. Companies B and C (mine) each independently re-implement that feature in a less crazy, far more readable and maintainable way. Company B gets their driver accepted with this feature (I think it was part of the initial submission). We get ours NACKed and are told to implement it for the entire kernel. This would be fine, I suppose, if companies A and B were also required to remove the feature and help implement it, but they weren't.
This might not have been such a big deal if we had been a big company. However, we had 1 dev (me) doing drivers and support for Linux, Solaris, OSX, FreeBSD, ESX, etc. So in the short term, management directed that we have customers ignore the driver in the kernel and use the one from our website so we didn't have to stop development on other OSes to implement this feature for our competition. This feature was eventually implemented in a generic way (with my help), but for several kernel releases the driver on our website outperformed the in-kernel driver by a factor of 3 or more.
FWIW, FreeBSD had no problem with the feature, and another driver author eventually ported it out of my driver into a general layer where it still exists to this day, and is used by almost every driver in the OS.