BPF at Facebook and beyond
lwn.net
lwn.net
Our L4 load balancer is implemented entirely in BPF byte code emitting C++ and relies on XDP for "blazing fast" (comms approved totally scientific replacement for gbps and pps figures...) packet forwarding. It's open source and was discussed here at HN before https://news.ycombinator.com/item?id=17199921.
We discussed how we use eBPF for traffic shaping in our internal networks at Linux Plumber's Conference http://vger.kernel.org/lpc-bpf2018.html#session-9
We presented how we enforce network traffic encryption, catch and terminate cleartext communication, again, you guessed, with BPF at Networking@Scale https://atscaleconference.com/events/networking-scale-3/ (video coming soon, I think.)
Firewalls with BPF? Sure we have 'em. http://vger.kernel.org/lpc_net2018_talks/ebpf-firewall-LPC.p...
In addition to all these nice applications we heavily rely on fleet wide tooling constructed with eBPF to monitor:
- performance (why is it slow? why does it allocate this much?)
- correctness (collect evidence it's doing its job like counters and logs. this should never happen, catch if it does!)
...in our systems.One of the pieces of fleet-wide tooling that heavily uses eBPF is PyPerf, which we talked about publicly at Systems@Scale in September ("Service Efficiency at Instagram Scale" - https://atscaleconference.com/events/systems-scale-2/ - video also coming soon, I think).
On top of these which is common for every machine, service owners can deploy their own BPF programs for specific use cases. In fact our self service tracing tooling is also a BPF program. We talked about it back in 2014 when it did not use BPF https://tracingsummit.org/w/images/6/6f/TracingSummit2014-Tr...
Re: Virtual Switch implemented in BPF: see Cilium’s work to connect containers with their BPF based connectors.
Also for GP, is anyone using vpp in production?
http://events19.linuxfoundation.org/wp-content/uploads/2018/...
But I agree that once it's more mature, it will be better overall.
"Current implementation only supports single queue, multi-queues feature will be added later.
Note that MTU of AF_XDP PMD is limited due to XDP lacks support for fragmentation."
That's a non-starter for many use cases, since multiple queues is a fundamental feature of DPDK.
I was chatting to a Facebook engineer on their use of BPF this summer and heard the same thing, which was surprising to me. There seem to be a number of companies that take advantage of Linux being licensed under GPL and keep their own forks/patches of the kernel that they use internally (anecdotally, I’ve heard Google does this), and the they stay on some old version, which apparently Facebook doesn’t do.
Getting things upstream can be more time consuming, but it's a worthwhile investment IMO.
I've seen this happen with Linux kernel stuff as well as other OSS projects.
$ git log v5.3...v5.0 | grep "^Author:" | cut -d"@" -f2 | sed 's/>//' | sort | uniq -c | sort -n | tail -50
167 samsung.com
188 sang-engineering.com
190 acm.org
192 broadcom.com
196 collabora.com
206 ingics.com
207 infradead.org
207 lixom.net
219 pengutronix.de
223 glider.be
232 c-s.fr
232 microchip.com
247 mediatek.com
251 roeck-us.net
329 st.com
341 fb.com
341 socionext.com
346 netronome.com
360 ti.com
380 canonical.com
384 linuxfoundation.org
387 nvidia.com
388 codeaurora.org
398 arndb.de
424 embeddedor.com
460 renesas.com
469 chris-wilson.co.uk
498 baylibre.com
504 suse.com
516 chromium.org
539 suse.de
540 lst.de
554 oracle.com
652 nxp.com
682 bootlin.com
690 davemloft.net
697 linutronix.de
774 arm.com
781 linux.ibm.com
948 google.com
1052 linaro.org
1230 huawei.com
1315 mellanox.com
1333 linux-foundation.org
1477 linux.intel.com
1501 kernel.org
1851 amd.com
2229 redhat.com
2548 intel.com
4373 gmail.com
There's the usual suspects in there (after a quick double check, it turns out microsoft.com just missed the cut-off with 140 commits, plus 2 from linux.microsoft.com, Amazon has authored 71, and alibaba have 139)No one wants to maintain out-of-tree patches, it's a complete pain. Google was doing it extensively with the Android project for the longest time, but they've been working hard at getting those all in to upstream to drastically reduce the work involved in updating the Android kernel.
:::cough::: Amazon :::cough:::
One challenge is that bpf upstreaming is much harder. We had to add support for relocations, spilling, multiple return values, and a few other things that might not be needed by the C bpf folks.
I'm curious why C BPF programs wouldn't need this.
To elaborate on what aey said in the sister comment, since eBPF allows user applications to generate code that runs in the kernel, the language is kept pretty simple so that it's easier for the kernel to verify that the generated code isn't doing something bad. If you look at the documentation, there are only two registers:
https://www.kernel.org/doc/Documentation/networking/filter.t...
There are also limitations on the size of eBPF programs the kernel will allow, and looping and such, so that a single user is less likely to DOS the entire system with a bad program (whether purposefully or accidentally). It's not really that similar, but I had the amusing thought that writing eBPF assembly vaguely reminded me a little of writing TIS-100 code.
Original bpf is a much simpler bytecode. Ebpf extended it and made it essentially x86_64 assembly (semantically). Then we all decided to call ebpf bpf for reasons.
(I'm one of the maintainers of said library.)
Meanwhile, the LWN kernel index (https://lwn.net/Kernel/Index/#Berkeley_Packet_Filter) will lead you to more information about BPF than you ever wanted.