Notes on BPF and eBPF
jvns.ca
jvns.ca
I must finally becoming a security pessimist when I read those sentences and the first thing I think is: these statements will not age well.
We definitely hit this one at some point: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/1763454
We also ran into a couple of similar but unrelated panic bugs on much newer kernels on non-Ubuntu distros.
It'll be valuable to learn this, so that we might be able to proactively address them. Ebpf is the core of our product.
I've personally never found obtaining a working kernel tree to be difficult, certainly easier than a working BCC toolchain. Or all of the various compiler flags needed for clang not to emit code incompatible with the verifier.
The verifier is definitely annoying, especially at first, but I found myself sort of quickly working out the verifier's expected idiom, and a lot of it can be wrapped with macros.
All of this drama pales in comparison to writing freestyle C code in the Linux kernel without causing random panics.
They should probably put it behind its own capability like CAP_LOAD_EBPF.
Forcing signed ebpf will also help though.
"Isn't supposed to be able to" is a lot longer and distracting vs the oversimplification-for-sake-of-understanding of "can't". As far as it being proven wrong though - that's already happened, eg CVE-2021-29154
https://blog.kernelcare.com/vulnerability/specially-crafted-...
If there's a physical possibility, it's just a matter of time before someone finds a way, as was proved by the CPU cache bugs leaking information.
Perhaps "is designed not to" but that's a mouthful.
I think we should accept "can't" and yet know the limitations of certainty.
I wouldn't be surprised if it becomes entirely root-only by default soon, if it isn't already.
Scroll down to "The ultimate ROP"
First: eBPF code is JIT'd in the kernel; at runtime, it is simply native code running at CPL0 alongside the rest of the kernel. Running eBPF code working with pointers is... working with raw pointers. There's no interpretation layer to bounds check or otherwise provide safety.
But eBPF code is meant to be safe: you can get a handle to some kernel structure that's passed to you from trusted code, but you can't bounce from it to a random offset in kernel memory. The way eBPF does this is by verifying the CFG of your eBPF program before it's translated to amd64 or arm. eBPF programs are generally just C programs (the simplest and best way to write an eBPF program is just to write a C program and compile it with the right LLVM flags), and verifying C programs is a hard problem; eBPF gets around this by only accepting a subset of all possible programs (those where memory accesses are simple enough to prove safe, that don't jump anywhere outside of known narrow range of program text, and that don't have unbounded loops).
The tricky thing here is that the eBPF verifier is pretty complicated and lives only in the kernel. People have found bugs in it. If you find a good verifier bug, you can launder an untrusted pointer into your eBPF program (in the end, these bugs end up looking sort of like the browser Javascript RCEs that finagle a bad pointer out of some part of the browser API).
The biggest mitigating factor for these bugs is that Linux systems generally don't expose eBPF to any user other than root, so the upside to these kinds of bugs is limited (it gives you root->kernel, which is not nothing, but not the top of most people's priority list).
The other big issue is that eBPF is a huge source of in-kernel flexibility about runnable code. Modern exploit mitigations are in large part about making sure that instructions running at CPL0 are all known, so that if you manage to corrupt allocator metadata or write an arbitrary 8 byte value at an arbitrary 8 byte offset you can't easily turn that into remote code execution. But, of course, eBPF is an in-kernel JIT; it's there to run essentially random code inside the kernel. eBPF code is normally constrained, but if you have a kernel memory corruption bug, you can aim it at the eBPF subsystem and violate the kernel's assumptions.
Also videos from eBPF Day KubeCon 2021: https://www.youtube.com/playlist?list=PLj6h78yzYM2Pm5nF_GmNQ...
The author's mentioned that you can trace MySQL with USDT, which is a tracepoint inserted by the developer at select locations in the code. This kind of tracepoints form a "stable interface" for tracing/performance debugging, whereas uprobe, which hooks into select userspace functions, are unstable as the binary is recompiled. Unfortunately, the USDT tracepoints (via DTrace) have been removed in MySQL 8.0. This makes it significantly more difficult to trace MySQL, although it's not impossiblhttps://news.ycombinator.com/item?id=29772927e. I've done a proof of concept of tracing MySQL with uprobe instead of USDT in this repo[1], which can kind of give you the same results (and possibly more stuff, as I can more easily read arbitrary memory address due to how the old USDT tracepoints are structured). This is not stable tho, as any MySQL upgrade may introduce incompatibility with the trace script, as I read memory address based on offsets (whereas with USDT this can be kept pretty stable). My appeal to Oracle to re-add this functionality[2] has unfortunately been rejected, which I think is a mistake given the wide range of possibilities unlocked via BPF.
[1]: https://github.com/shuhaowu/mysqld-bpf
[2]: https://bugs.mysql.com/bug.php?id=105741
Another thing that I've been recently thinking of is using BPF to validate programs written for real-time Linux (via PREEMPT_RT). To my understanding, one of the main thing to avoid is page faults [3]. With the proper BPF tracing scripts, I think we can validate that programs indeed avoids page faults in integration testing. I'm not sure if it is super useful yet, but as I'm trying to write a few RT programs, it's something that came to my mind.
[3]: https://lwn.net/Articles/837019/
In addition to tracing (so bpftrace-based/bcc-based tools), I've recently discovered that there there are:
1. ebpfsnitch (https://github.com/harporoeder/ebpfsnitch): which is an application-level firewall without kernel modules.
2. ebpf-traffic-monitor (https://source.android.com/devices/tech/datausage/ebpf-traff...): which appears to be using BPF to account for traffic for different apps on Android.
3. kubectl trace (https://github.com/iovisor/kubectl-trace): Run tracing on k8s.
There are apparently also use cases in the context of security, but I'm not familiar with it.
Sorry, I'm a little confused why this would be necessary? Like, sure, it's a nice to have on a CI as a basic sanity check but if you just invoke mlockall you'll end up with everything wired down and you're good to go regardless?
While on that thought experiment, perhaps in general if you can use BPF/USDT to help trace/debug RT programs? I'm thinking of being able to verify/visualize timing for better tracing? Or maybe there are already tools that existing that I don't really know how to use (like ftrace + trace compass, maybe)
Whereas ebpf allows instrumentation free enforcement. Plus, app devs do not need to be aware of this fact. This facilitate separate of responsibility in code and organization.
>...
>seccomp / landlock security things
Landlock does not use *BPF.
Seccomp can only use BPF at this point, not eBPF (though there has been some work on it).