Sysadmin friendly high speed Ethernet switching
blog.benjojo.co.uk
blog.benjojo.co.uk
https://www.kernel.org/doc/html/latest/networking/switchdev....
https://github.com/dentproject/dentOS
Switchdev support means you don't need hardware-specific userspace tools (with their own bizarre syntax to learn) in order to configure the switch.
DentOS support means the device uses a sane bootloader (uboot or grub) and the only binary blobs on the device will be the ones built into the bootloader (IntelME, Arm Trusted Firmware) and the switch firmware which will be part of linux-firmware (and therefore very easy to manage/update).
In particular, looking for these two keywords is how you make sure that the hardware vendor is staying on "their side of the line" between hardware and software. Violations of this line are endemic to 10G+ switching.
For 1gbit: Ubiquiti Edgerouter ER-4 and ER-6; both have SFP cages for fiber optic transcievers https://github.com/NixOS/nixpkgs/pull/194153 https://store.ui.com/us/en/products/er-4
I am curious though what configuration option prevents this from ending up with software switching. I understand the mellanox kernel module was compiled and loaded, but certainly that doesn't mean that anything you do in the network stack gets converted to switch fabric code and uses hardware packet switching. How do you make sure that you don't errantly wind up with poor latency and capped throughput?
But also, short of making and selling your own networked device for whatever reason, what are the real benefits of going this approach? I can see crazy use cases for where you have full control of the network stack (but again, see my point above — how do you guarantee you are not doing this in software?) but for most purposes, especially with qsfp and fiber, how much are you really gaining by doing it on-device? What is the killer use case here?
EDIT
Upon rereading, it seems the switch is hard-coded to be hardware-switched and cannot end up in a situation where you are accidentally using software packet switching in the first place (i.e. it does not just optimize to a hardware packet switching state). But that limits what you can do considerably, to the point that an off-the-rack Juniper or CSCO or whatever probably has more features than you can do here without writing your own code to hook into the mellanox sdk?
I mean, I'm not touching any mellanox sdk here, I am using the a very similar stack that someone on a "software router" would use, on a switch that can automatically accelerate it to 800G+ throughputs, while hitting a 60W power target.
You can hit some of those performance/power numbers in vendor hardware like Juniper/Cisco/Arista, however you have to also put up with their software, I (and others in my group of peers) have not had great experiences with vendor software, and in this setup I am able to patch/fix the software on my own terms.
If there is a security vuln in one section, I can fix that, and call it a day, I won't be forced to upgrade parts of the system I do not want to. I cannot do this with Juniper/Cisco/Arista always.
*) IP Routing that would normally "fit" in a vendor switch
*) Bridging
*) VRFs
You will be fine
If you try and do some weird stuff then it's best to check with "ip route" to see if it was actually installed into hardware or not, but I would simply not do anything weird on such hardware
If you really want software switching you have to use the management port (there's only one or two of these) whose name is "eth0" or "eth1" or something like that. So avoiding "accidental software switching" is really easy -- if you're typing "eth" you're doing it wrong. You can even explicitly delete this interface if you don't need the CPU to be able to snoop/inject traffic to/from the switch ports.
This is common with small home routers - the WLAN and wired LAN ports (all typically appearing as one NIC) will be made part of a bridge `br0`. The four LAN ports aren't typically exposed as separate NICs so there is hardware switching going on there (some devices do let you split them out because they are VLANed internally though).
If `ip link show type bridges` doesn't show any bridges then you aren't software switching unless your drivers are lying to you.
The answer is "switchdev":
https://www.kernel.org/doc/html/latest/networking/switchdev....
The Linux switchdev driver is the awesome magic that says "make hardware-offloaded switching ASICs look just like software switching". It's beautiful and amazing, as you'd expect from Mellanox.
How disruptive are switch config changes? If I edit /etc/network/interfaces to add a new port, vlan, etc - does a `systemctl restart networking` (or whatever equivalent) bounce ports or halt switching for a moment while changes are applied?
IIRC the /etc/network/interfaces does a reconfiguration that's pretty disruptive.
Things like brctl and ethtool worked on the fly without issues (note though that I mostly used Arista years ago).
It is usually non-disruptive if it gets applied as deltas. If your config tool does a teardown/recreate then that's disruptive. Within the bounds of ethernet and routing protocols (OSPF DR/DBR changes are disruptive, STP can be fun, ....).
They are seamless as long as your configuration does not do something stupid like tear down interfaces to reconfigure them. The switch takes no noticeable time to program from the Linux space to the ASIC
> - does a `systemctl restart networking` (or whatever equivalent) bounce ports or halt switching for a moment while changes are applied?
"systemctl restart networking" will typically blow away even the well most well configured systems in my experience.
In the post I suggest using ifupdown (the first one), since it's the most "easy" to debug, but I'm sure networkctl works too, with a healthy amount of systemd restraint
Even though I have been running HPC clusters for years, I still really don't like core switching infra[1]. Its just a pain in the arse to investigate and run.
[1] I realise that this is very much not a core switch, or even a TOR switch.
(Remember... this is a hobby where people persue their interests and experiment with things they want to know more about, not a proscribed set of exercises for just one path towards devops engineer).
Is there an OS or another Linux distribution that matches Debian's performance in this respect, without the complexity of an entire Linux system? Could Debian be stripped down (and then how are updates applied)?
I don't really know why you would want do that (though, RHEL based would be a reasonable 2nd option)
Performance of the distro isnt really a big deal because the OS has so little to do with the day to day packet shuffling, My switch runs at a load average (crappy metric I know, but to give you a idea) of 0.04.
I apply updates with apt update and apt upgrade, as you would a normal Debian server. Only the kernel is pinned since it is special for the switch, however you can rebuild the kernel deb's when you want, as the driver is in the mainline kernel repo.
This could be done with any Linux. I'm not sure what 'complexity' you are referring to, but the complex open modular nature of the Linux kernel allows us the flexibility to build a custom kernel for this interesting application without unnecessary modules. To be frank all modern operating systems are complex because modern hardware is complex and user expectations are high. Some just do a great job of hiding the complexity from the user. Linux makes an attempt to expose it's complexity, but it's complexity is about the same as iOS, macOS, Windows, etc.
If you wanted to do without Debian or any distribution you could compile your own kernel, your own user space tools, and then install it on the hardware. At this point updates would need to be applied by pulling code from upstream and compiling it yourself. Look into Linux From Scratch to get an idea of how Linux without a distribution works. However, I don't think the juice is worth the squeeze in this situation.
EDIT: My answer to "is there another OS", to which the answer is yeah, probably a dozen. I think a blog post of doing this with openBSD would be interesting because I'm not sure what the exact steps are to install custom drivers on a BSD and openBSD has a very high standard for security. I think the reason it's done on Linux is because there is a lot of expertise about Linux and this type of project is relativity straightforward with Linux.
All you need is a driver for the switch chip that serves it as a set of files you can read and write configuration commands to. With the right driver setup you could then configure the chip using human readable textual commands and use a script and standard cmd line tools to configure the switch. We could PXE boot plan 9 and leave the switch diskless. PXE booting plan 9 networks is brain dead simple and works out of the box.
Familiarity with plan 9 and its kernel along with a solid understanding of c. Plan 9 had its own c library which is close to c90 but much cleaner imo. Networking and threading libs are really nice.
Driver might be in two parts like usb where the kernel driver serves up the usb controller and attached devices while user space file servers open the devices and serves them. e.g. a webcam is served as a video stream file.
> Also, is Plan 9 maintained sufficiently?
A community fork known as 9front is maintained and receives patches almost daily. http://git.9front.org/plan9front/plan9front/HEAD/info.html and http://9front.org/releases/
If you like videos here is a really nice channel of a plan 9 hacker: https://www.youtube.com/channel/UC7qFfPYl0t8Cq7auyblZqxA
Switch like OPs with all traffic passing to and from a DPU to make a poor man's Aruba CX10000
All other DPUs seem to be NDA-encrusted.
See: https://github.com/na-son/nvidia-air for bootstrapping a non-EVPN topology.
Not sure that's a fit for home network enthusiasts, and for business, just pay the bucks to be able to pass the buck - ie get a packaged solution not homebrew.
For what it's worth I did not pay 5k USD for it, that is significantly overpriced for it on the 2nd hand market.
> for business, just pay the bucks to be able to pass the buck - ie get a packaged solution not homebrew
There is a significant cost difference here,
The overall point of the post is that the "packaged solution" is not viable when most switch vendors have what I can describe as "crap" software quality and/or support.
So if you are in the land of simple L2 switching and L3 routing, this switch is amazing because you can escape the crap vendor software.
If you're really into homelab stuff or trying to run a small business where networking is critical, they're quite a good deal.
The "packaged solutions" as a new product from a known vendor with a support contract (because honestly if you're not buying the support contract you're going to spend a LOT of time messing with the switch no matter what vintage it is) are at least 10x the price, which very well could be much too much for a home user or small business.
I'd love to have managed 2.5Gb switches with 10Gb uplinks in my house using a custom linux OS that I can use standard config management tools with...
Me too. I have been eyeballing "SparX-5" based switches, which do run Linux ("SMBStaX"), but you'd need something like million bucks to get anywhere :( Could be candidate for a Kickstarter style project maybe...
[0] https://conclusive.tech/products/rchd-sparx-networking-som/
Have you found any for sale other than the Microchip dev board?
I stumbled across this recently: https://www.servethehome.com/insane-48-port-2-5gbe-2x-25gbe-... but I haven't been able to find where you can buy one these days.
Sadly my PCB design skills are not quite up to dense BGA escape + 25gig lanes...
100%, this has been the most frustrating thing about following the various "open" switching worlds. There's a massive gap in the middle between the sorts of 4-8 port switches that end up in OpenWRT-compatible routers and this sort of enterprise switch that's barely accessible to the "homelab" class user.
I would absolutely love to have some open switching in the "Ubiquiti" class, desktop and 1U rackmount devices with gigabit through 10G as their primary interfaces. I'm personally in the VoIP world and if I could install Asterisk directly on a 48 port PoE switch I'd be deploying them by the dozens.
Check out DENT NOS: https://dent.dev/ There's a Delta 32x1G (PoE+) + 16x2.5G (PoE++) + 6x25G SFP28 switch that can run DENT.
1. 48-port 10/40 GbE SFP+ with mostly copper UPoE/++
1.5. 2-4 100 GbE QSFP28 uplinks
2. Doesn't sound like a jet engine
3. Doesn't phone home to a cloud in another country
4. Doesn't use an obscure, fragile configuration language
5. Doesn't cost $5k
1. 8-port nbase-t (2.5/5/10G supported in all ports)
5. Doesn't cost $500
This would suffice for the fast part of my home LAN... the rest is handled well by a gigabit ethernet switch with one 10g uplink.
https://www.ebay.com/itm/386453004153 Arista DCS-7050SX-64-R 48P 10GbE SFP+ 4P 40GbE QSFP+ RA Switch $180
That thing retails new for 10k. You got awesome friends if they have such a thing lying around unused and willing to sell it to you at a discount!
For those of us with less fortunate bank accounts: what's the smallest and reasonably affordable Mellanox model that has a similar featureset in terms of native Linux support?
Tactical eBay (or whatever is the robust 2nd handmarket in your region) can yield similar discounts on such hardware. Retail price is often list price, and that is often a large mark up regardless. The 32x100G port version of the same switch goes for around £2000 in the UK 2nd hand market
> what's the smallest and reasonably affordable Mellanox model that has a similar featureset in terms of native Linux support?
The SN2010 is likely the smallest, the SN2700 is likely the cheapest
Until I can buy one of these things or something similar, which I am now highly motivated to do, if you have any experiences or insights you can share, I’d love to see them.
Ultimately vyos/vtysh consumes this single configuration file and uses it to template commands or configuration files for other underlying software. For example if you configure ipsec in the vyos configuration file, it will produce a configuration file from a template for FreeSWAN to implement the options you have set.
I don't personally know offhand how things like bridges are configured in vyos and how vyos maps bridge configuration onto the physical interfaces, but I do know that there is more than one way to do it in Linux, and not all ways of doing it are compatible with switchdev interfaces.
Without vyos having native awareness of switchdev device types or support in its configuration templates for applying features in ways that are tested and compatible with switchdev interfaces there's no way to know if a particular vyos configuration will result in something that works with switchdev or not.
Yes, you can always drop to running manual commands to perform fixups or additional configuration but at that point you're losing most of the sauce that vyos provides and you might as well just go back to vanilla linux where you have full control of everything.
It is my understanding that all the linux-based switch os's (Cumulus, OS10, etc.) work similarly -- custom shell + custom configuration file mapping to templated backend configurations. It would be nice to see vyos gain official templates tests and support for hardware accelerated switchdev
The 10GbE ones are silent, and the 100GbE one isn't exactly loud, unlike lots of second-hand kit which has come from a DC...
> what did this guy connect to several 100 GbE
They have 100GBASE-PLR4 optics in them, that allow the 100G ports to be split up into 4x25G ports (or actually in this switch, 2x25G ports due to a hardware limitation with this switch)
> how does the upstream connections he mentions look like, and from which provider
They are just normal 10G-LR Single mode optics, in a data center.
> The device is second hand, so the likely use-case is a Home(?)Lab type of setup.
Nope, this device now runs my business bgp.tools
The high speed (25G+) does not have good solutions for multimode, and the length limits that physics enforces with multimode mean that it's not "no-brainer" applicable for going any more than between the same rack row.
So if you are dealing with datacenter cross connects that can exceed the max distance for MultiMode then you might spend hours debugging broken stuff for no real gain. MultiMode is slightly cheaper, but it's a false economy the moment stuff does not work correctly. I've spoken to people with DC cross connects that go into the 5km+ of cable distance. So it's easier to just stock one kind of optic per speed, and call it a day.
Equinix I think already phased out MMF XCs
Single-mode fiber is future-proof. Utilities bury that stuff and depreciate it over a 30-year lifetime. There has been like one spec change since the 1970s.
Single-mode fiber. Always single-mode fiber. Nothing else, ever.
You may be surprised at how quiet and low power these Mellanox switches can be.
No need to have switches and servers in your rack anymore, every server is a switch and every switch is a server with a 192 threads CPU. Insane.
https://ipng.ch/s/articles/2023/11/11/mellanox-sn2700.html
The cable has a SAS (SFF-8087) connector on each end. I bet you can replace it with an SFF-8087-to-Oculink cable:
https://www.amazon.com/chenyang-SFF-8611-SFF-8087-PCI-Expres...
and a PCIe-to-oculink card like one of these:
https://www.amazon.com/Ableconn-PEX-OL153-OCuLink-SFF-8612-A...
(none of the links are affiliate links)
I find understanding how Linux networking really helpful in uunderstanding mikrotiks (which I use a lot in prod, although tend to shy away from for the most critical and demanding of services)
As impressive as it is, switchdev is a big departure from "Linux networking" though. This is not a great platform if that is your main objective. The interfaces are not "normal" in that respect and huge subsystems of the Linux network stack do not apply.
I remember 10base*, so I find 100000base* pretty amazing.
-The SN2010 retailed for about 10K (Euro/Dollars)
-It has never been truly available, as in: you could go somewhere, order it, and expect it to turn up in 2-3 days
-Even though small, these units tend to be loud, with at least 2 tiny fans making a lot of high-pitched noise
But the most salient point is that, even with "Linux on a router/switch", there's no guarantee that you'll get decent performance, as that entirely depends on how well the kernel understands the (proprietary) onboard chipset, which usually means that you're squarely in "well, here's this blob that works on certain kernel versions, and good luck with that!" territory.
As long as the ASIC is programmed the performance is the same between kernel versions. If the ASIC does not get programmed then it's like the entire device does not work at all (so you likely need to roll back the update you made).
> which usually means that you're squarely in "well, here's this blob that works on certain kernel versions, and good luck with that!" territory.
As mentioned in the post, the switch in question has blobless drivers, unlike the broadcom stuff you likely have experience with based on what you are saying