The CPU Cost of Networking on a Linux Host
people.kernel.org
people.kernel.org
In the last 15 years there has been a hard move away from these architectures. Almost no packets are forwarded by the same processors running management and control plane functions anymore. This is mainly because the required traffic rates today need dedicated silicon purpose built for the task (the Broadcom Tomahawk3 can do 12.8 Terabits/sec above a relatively small packet size).
I don't know how things will shake out for the Linux world and x86 packet forwarding given the trend and lack of real performance in the kernel. Right now, your best bet when it comes to Linux and high network throughput/packet processing requirements is to just bypass the kernel entirely with DPDK, a "smart" NIC, or XDP.
That's pretty much settled isn't it? It was settled in the same way the "it got too much for the general purpose CPU" problem always gets settled - you do it in custom silicon, and define a standardised interface.
That's pretty much what happened with 3D graphics where the standardised interface was OpenGL (but now seems to be up in the air). It's also pretty much what's happened in AI with standardised libraries interfacing to custom hardware, and it's what happened with networking.
For networking, the interface standard is OpenFlow. So, if you think it's possible you will need to handle links of about 1Gbs or over in the future, you do your networking using an OpenFlow implementation like Faucet. If it's not much above 100Mb/s the Linux kernel module that implements OpenFlow, called openvswitch, will be fine. Otherwise use some custom hardware.
Openvswitch has been around since 2009 - so it's not exactly a new thing.
https://www.broadcom.com/company/news/product-releases/52756
I don't know much about the technical details, but the pitch I've heard is that it gives you ASIC level performance with more flexibility to reprogram the chip (not full FPGA).
That said, there's only so much you can do in a chip before considerable tradeoffs are going to be made. They're not going to offer the same level of flexibility you get out of a general purpose CPU, but may not have same the restrictions of most fixed pipeline chips - their product sits somewhere in the middle. Also, P4 seems to sit in a space complex enough to make it unreasonable for most network shops - it's not for your average enterprise or service provider network.
It's been longer than that for router vendors. The move to hardware based forwarding was in the 90s.
https://blogs.cisco.com/sp/a-bigger-helping-of-internet-plea...
While they still obviously make custom ASICs, they've moved towards more software where possible.
To be clear, Linux has a very robust networking stack. But it will never come close to the natting and routing performance of an actual router.
And so we develop things like DPDK to spend even more CPU just to keep things usable, but it still feels like a big step backwards.
A typical k8s deployment runs in containers that are in VMs. So each packet you want to send to or from a container needs to touch a cpu and traverse a networking stack six times. That's dumb.
Is it? It's certainly inefficient compared to dedicated hardware. But so is anything relying on a CPU - we could just use ASICs for everything. But then every logical change requires weeks/months/years of development and manufacturing.
The goal of k8s, VMs, etc is flexibility. I can set up a 100-node k8s cluster with less-than-perfectly-efficient networking stack in mere minutes. Good luck matching that with dedicated hardware.
Right, except the asics already exist and you're actively choosing the less efficient, more expensive option.
> The goal of k8s, VMs, etc is flexibility.
I don't think this is necessarily bad, as long as you understand the tradeoff. You're choosing a fundamentally slower architecture to make management easier. It's a choice of prioritizing the developer experience over the user experience.
On the other hand: the cost for development goes down if you don't need to pay for extra steps taken by an extra person. If you take a 10-step process that is run by 5 people and reduce it to a 5-step process run by 3, you have 2 more people to do other stuff, or roles that you don't have to create/fill in the first place.
It’s hard to imagine any dev task being more expensive than organizing and programming tables into a trident or tomahawk chip using the broadcom’s SDK.
The transition from physical host memory -> vnic is analogous to just sending the packet over a wire. That's not the slow part of modern networks.
It's possible to optimize basic packet forwarding in a hypervisor/containervisor using leaner options like ipvlan or even SR-IOV. If you replace hardware firewalls with something like k8s network policy then you are technically replacing hardware with software but it isn't that slow if configured right (hello Cilium) and you're probably implementing a more sophisticated policy anyway.
If you need nonstop CDN levels of bandwidth, then you're probably going to go with a setup that has fewer layers, but CRUD applications don't exactly need terabits of bandwidth.
https://www.intel.com/content/dam/doc/application-note/pci-s... is a good overview.
Also, at least on AWS, you can directly attach VPC network interfaces to containers, which (among other things) obviates the need for software bridging, veth pairs, etc.
Finally many platforms like docker, in the past, will spin up userspace proxies (docker-proxy) which is even more expensive.
https://en.wikipedia.org/wiki/Channel_I/O
I always thought PC's could be setup to do something similar. The cores used for that could even be small and simple compared to main cores. Folks in early 2001 could be Bittorrenting their Linux distros or whatever they use it for with no user-visible lag. I'd like at least one each for user-input devices, graphics, storage, and networking. Other good reasons to isolate them a bit from each other.
There's already ARM SoC's in embedded that have a good core for apps and weak one for I/O. Recently, the RISC-V chip with the minion cores. If embedded can do it, then there seems to be no technical limitation holding back desktops. Just marketing, backward compatibility, etc.
Do you know a good source, to learn about the lower-level Linux/container/docker/k8s networking?
https://blog.packagecloud.io/eng/2016/06/22/monitoring-tunin...
https://blog.packagecloud.io/eng/2017/02/06/monitoring-tunin...
The illustrated guide to receiving data is also solid:
https://blog.packagecloud.io/eng/2016/10/11/monitoring-tunin...
Second, router ASICs can be designed to sped transistors on a predictable packet flow path, instead of intelligence like out-of-order execution to make general purpose code fast. For example, routers have big expensive content-addressable memories called TCAMs that are used to store things like routing tables and ACLs. The ASIC, moreover, can implement a highly tuned pipeline designed around the latencies in the underlying memories. E.g. you get a packet, grab the destination address, look up the next hop in the routing table, etc. Each step takes a predictable amount of time that you can account for and optimize.
Third, parallelism is much cheaper in hardware than in general purpose CPUs. It takes a lot more transistors to be able to execute a second general-purpose instruction stream than to have a single-purpose circuit that does some work in parallel with something else.
Yes. Most routers + switches are "line rate" meaning the packets go through the switch at the speed of electricity, as if the switch wasn't there are all.