Bypassing the Linux kernel for high-performance packet filtering
blog.cloudflare.com
blog.cloudflare.com
Packet BRICKS is a Linux/FreeBSD daemon that is capable of receiving and distributing ingress traffic to userland applications. Its main responsibilities may include (i) load-balancing, (ii) duplicating and/or (iii) filtering ingress traffic across all registered applications. The distribution is flow-aware (i.e. packets of one connection will always end up in the same application). At the moment, packet-bricks uses netmap packet I/O framework for receiving packets. It employs netmap pipes to forward packets to end host applications.
(Credit goes to Asim Jamshed, who pulled this off as part of an internship at ICSI.)
bricks> lb = Brick.new("LoadBalancer", 2)
bricks> lb:connect_input("eth3")
bricks> lb:connect_output("eth3{0", "eth3{1", "eth3{2", "eth3{3", "eth2")
bricks> pe:link(lb)
This binds pkteng pe with LoadBalancer brick and asks the system to read ingress packets from eth3 and split them flow-wise based on the 2-tuple (src & dst IP addresses) metadata of the packet header. The "lb:connect_output(...)" command creates four netmap-specific pipes named "netmap:eth3{x" where 0 <= x < 4 and an egress interface named "eth2". The traffic is evenly split between all five channels based on the 2 tuple header as previously mentioned. Userland applications can now use packet-bricks to get their fair share of ingress traffic. The brick is finally linked with the packet engine.We're using packet bricks primarily for high-performance network monitoring in environments with more than 10 Gbps aggregate upstream traffic.
Thinking of the future: Recent experience suggests that we are able to do around 50-100 Mpps of traffic dispatching in Snabb Switch using one CPU core. I suspect that dedicated software traffic dispatchers will displace hardware (RSS, VMDq, etc) in the immediate future.
We are planning to explore this soon in the context of software dispatching for 100Gbps ethernet ports.
That is one solution: take a 10G port into Snabb Switch, filter and sort the traffic, then feed a slice to the kernel e.g. via a tap device with multiqueue /dev/vhost-net acceleration (same interface that QEMU/KVM uses).
The risk I see here is that maybe it is even harder to tune your kernel when it is using a software interface for I/O instead of a hardware one. The kernel can depend on so many things (multiqueue, TSO, LRO, checksum, encap offload, etc) that can behave differently between hardware/software NICs and you would need to be confident that this will work out well. Otherwise the risk is that you take your hardest problem - tuning the kernel - and make it even harder.
If we are only talking about 2x10G ports per server then one alternative would be to connect Snabb Switch to the kernel with a physical 10G port instead of a software one. That is, separate the DDoS-protecting frontend (Snabb) from the backend (kernel) with a network cable. You could still run the Snabb application on the same server but with a dedicated network card cabled directly to the kernel. (You could also run it on a different server if you prefer.)
End off-cuff braindump :-)
Or just do everything in userspace. Reinjecting into the kernel is not strictly necessary, although obviously you may be constrained by existing code.
http://daemonkeeper.net/781/mass-blocking-ip-addresses-with-...
TLDR: http://daemonkeeper.net/wp-content/uploads/2012/05/ipset4.pn...
https://www.usenix.org/system/files/conference/osdi14/osdi14...
There are two problems with doing the "take over the nic" techniques:
1) I don't believe you can actually push, say 2M pps back to the kernel with any of this techniques. There is a reason RSS exists, and even if you can process 10M pps on one CPU, it doesn't mean it's easy to insert them back to kernel.
2) I don't think putting a piece of custom code between CloudFlare kernel and network card is feasible on the architectural level. You really want to stand in the way and have to actively forward all these packets?
The example in the article found a single core handling 1.4M packets per second. If you're running a web-server shoveling data out to clients those packets are going to be close to the maximum size which, if I haven't screwed up the math, looks something like this:
1.4M * 1400 bytes (assuming a low MTU) * 8 (bytes -> bits) = 15Gbps
That's not to say that there isn't still plenty of room to improve and, as lukego noted, there's a lot of work in progress (see e.g. https://lwn.net/Articles/615238/ on work to batch operations to avoid paying some of the processing costs for every packet) but for the average server you'd find bottlenecks on something like a database, application logic, request handling, client network capacity, etc. before the network stack overhead is your greatest challenge. The people who encounter this tend to be CDN vendors like CloudFlare and security people who need to filter, analyze, or generate traffic on levels which are at least the the scale of a large company (e.g. https://github.com/robertdavidgraham/masscan).
Improving the kernel would greatly speed up many of our applications with no real downsides.
With 40G to the edge, we also need to be able to properly firewall/filter/and all that fun stuff the traffic!
I am kind of amazed that the Linux kernel did not become the dominant data-plane for the networking industry ahead of proprietary implementations from Cisco, Juniper, etc.
Hopefully Snabb Switch will have better luck there... ;-)
Cisco IOS is variously hosted on a proprietary real-time OS, QNX or Linux. Juniper JunOS is FreeBSD-based. Arista AOS is Linux-based.
With the advent of SDN, the packet-rate limitations of a general-purpose OS lead to things like DPDK and Snabb, but they both run within a Linux host environment (DPDK can use FreeBSD as well; not sure re: Snabb).
And also the article is talking about handling just 10G, not 100G!
You're about to see a whole bunch more 10G in the coming 5 years.
Also, 10gb servers are not rare by any means. Take a look around the next time you walk in a colo. 10 gb servers everywhere.
Mostly I just wish you'd read it again: you appear to have missed the part where I said that this is a real problem which needed working on.
The point I was making is that it's not a problem for most Linux users. Linux includes millions of devices attached to sub-100Mb networks but even if you want to look solely at things in modern data centers ask yourself how many of them are running network-limited applications or are providing services to internet users over an uplink which is actually fast enough to stress modern hardware. For all but the most demanding users the Linux vs. BSD decision will be made on other factors.
Just to use an example which you see a lot today: how many of the developers jumping on Docker care that much about network performance at this level? I would argue that continued development of the container system has done more to boost Linux usage than low-level networking performance, even though both are entirely legitimate and worthwhile concerns worth developer sponsorship. (ZFS has pulled in the opposite direction for people who care about storage)
From the other direction, imagine if *BSD had gotten serious about package management by the mid-to-late 90s when it was obvious how much better the experience was on Debian so that a generation of developers wasn't trained to favor Linux to avoid getting sucked into dependency management. That doesn't have anything to do with the kernel but it mattered more for many, many people. This would have been really interesting if kfreebsd had hit critical mass and made the cost of switching that much lower.
You do sound both mean and factless. Meaningful apple to apple BSD to Linux comparison needed.
BSD numbers > Linux numbers.
That doesn't mean it's a (Free)BSD issue, just that Netflix chose a particular architecture and are doing the work to make it go fast.
Netmap is in FreeBSD, and available for linux. I think it might be in Dragonfly at this point as well.
There are three main reasons the kernel is slow for networking: per-packet dynamic memory allocation, lots of memory copying, and system call overheads.
The first two can be improved by modifying the kernel, and I think people are attempting to do this. The system call overheads arise naturally from having the networking code in the kernel. Basically every time you perform a system call the kernel has to save the userspace context, do the system call, and restore the context. This takes time and is bad for cache locality.
But as others noted, for most people who aren't cloudflare this doesn't really matter.
Aren't most web applications I/O bound? The Arrakis team sped up Memcached and Haproxy quite a lot by bypassing the kernel. It seems like there could be a large market for these techniques as they become easier to use.
http://people.inf.ethz.ch/troscoe/pubs/peter-arrakis-osdi14....
However, even after a few years now I don't think they are seeing light at the end of the tunnel.
Also eg. http://vger.kernel.org/netconf2009_slides/LinuxCon2009_Jespe... from 6 years ago show forwarding at 4 Mpps.
(Forwarding is, of course, receiving AND sending instead of just receiving so these should translate to higher RX-only numbers.)
that-said...slightly-longer: for what they do there are alternatives like ipset for other things its not as clear-cut, hence things like PF_RING. its not that great thought, you're sacrificing all features for fast sniffing.
technically a good zero-copy implementation of packet mmap w/ a userspace ring would achieve +- the same thing, too.
Yet you never hear about Google or AWS using kernel bypass in their load balancers, for example (possibly a trade secret, possibly the result of Linux monoculture).
ISPs are where I expect to see the disruption of HPC-oriented x86 servers being supremely capable of handling work previously done by specialized hardware.
Pretty much all high-frequency trading today uses FPGAs, for instance, and that's often with teams of fewer than ten people.
https://www.youtube.com/watch?v=uDy_8Q0GdTk
IIRC, the FPGA has been incorporated into a switch. When a market data packet starts to arrive, the system starts sending a response packet before the input packet has completely arrived and before the system has actually made a decision. While the input packet is read, the system decides whether or not it will cancel the response market order by intentionally corrupting the checksum of the output packet at the last possible instant.
The issue is that Solarflare has patents on some of the underlying techniques. They were very open to licensing those patents for a reasonable figure when I last talked with them about it, but that wouldn't work for a general-purpose OSS project.
I am curious about packet filtering in Windows. Anyone with experience in HN?
Now, in my company, we are doing some tests using different methods: WinPcap, WFP, NDIS, and WinPcap is the winner in a VM but we will start to test with real 10gbps ethernet cards next week.
This isn't one of these "Windows sucks!" posts. Windows is a very good endpoint for many services (DNS, DHCP, AD, IIS/ASP.net, VPNs, etc) and of course an extremely popular client.
But that being as it may, packet storms should never be allowed to hit endpoints, that the entire purpose of packet filtering. So you'll want to be taking them out on front-line appliances, and appliances based on the Windows NT kernel simply don't exist.
So this is why CloudFlare cares about this, they're utilising Linux on an appliance in front of their endpoints to try and drop as many "bad" packets as they can detect. Both Linux and various BSD variants are used commonly on networking equipment, so trying to optimise them seems to make a lot of sense.
Windows on the other hand? If you're trying to do packet filtering on the endpoint itself then you're fighting a losing battle. For example, they're low-level hooking network traffic, and while that works wonderfully for filtering, it is a terrible idea if the machine is used for other things as it can disrupt normal legitimate machine traffic.
Also, this is oriented to internal network endpoints not visible from Internet, so I don't expect to receive massive network-intensive attacks.
It is completely OS-bypass, and should be good to handle full line rate (14.8Mpps per port).
In our tests WinPcap was faster than an NDIS driver, so it will be interesting to compare.
That said I think the Wireshark guys (linked from the Wireshark site anyways) might have some answers. I know for WiFi capture they had fully functional Windows devices.
And I have performance numbers with OVS-dpdk that make kernel bypass compelling since its off-the-charts while comparing with kernel datapath.
For those interested, my ovs-dpdk experiments which also include patches, README etc. for others to carry it themselves can be found here:
https://www.dropbox.com/sh/nfe70cgksmy543k/AABD_0qsQ15e2GItX...
The perf results directory has the ovs-dpdk perf for all the use-cases in the dataplane performance pdf that you might be interested in.
Has use-cases covering up to 11VM or 11 containers with 11 IP flows to measure dataplane performance.
Also you really need a server with 1 gig hugetlb support (and also enable that for guest) to extract maximum performance.
Expected I guess ...
Or how about a faster kernel? :)
I've done high performance nginx, in the million request per second range (there are situations that benefit from these, though unfortunately such discussions always get waylaid by people insisting that performance doesn't matter), but there is enormous system overhead at this rate that I'd like to get around.
I don't know if a complete drop-in solution is the right solution though. If your application is performance sensitive enough to require embedding a full networking stack, you might as well make use of better APIs. For example it'd be silly to indirect the event dispatching through something poll/select-like. Instead you'd much rather just have the core IO loop call the handlers directly. Or as another example, zero-copy will be impossible with a recv()-like interface where the client provides the buffer that data needs to go to, but will be trivial with an API where it's the network stack giving the client a buffer that already has the data.
If you want to experiment with this, mTCP (http://shader.kaist.edu/mtcp/) is probably the right starting point.
presentation: http://www.openonload.org/openonload-google-talk.pdf
http://people.inf.ethz.ch/troscoe/pubs/peter-arrakis-osdi14....
BSD has had userland networking for a while:
http://www.bsdcan.org/2014/schedule/events/447.en.html
I'm not sure if there's anything comparable for Linux.