A Beginner's Guide to eBPF
github.com
github.com
> eBPF (often aliased BPF)[2][5] is a technology that can run sandboxed programs in a privileged context such as the operating system kernel.[6] It is used to safely and efficiently extend the capabilities of the kernel at runtime without requiring to change kernel source code or load kernel modules.[7] Safety is provided through an in-kernel verifier which performs static code analysis and rejects programs which crash, hang or otherwise interfere with the kernel negatively.[8][9] Examples of programs that are automatically rejected are programs without strong exit guarantees (i.e. for/while loops without exit conditions) and programs dereferencing pointers without safety-checks.[10] Loaded programs which passed the verifier are either interpreted or in-kernel JIT compiled for native execution performance. The execution model is event-driven and with few exceptions run-to-completion,[2] meaning, programs can be attached to various hook points in the operating system kernel and are run upon triggering of an event. eBPF use cases include (but are not limited to) networking such as XDP, tracing and security subsystems.[6] Given eBPF's efficiency and flexibility opened up new possibilities to solve production issues, Brendan Gregg famously coined eBPF as "superpowers for Linux".[11] Linus Torvalds expressed that "BPF has actually been really useful, and the real power of it is how it allows people to do specialized code that isn't enabled until asked for".[12] Due to its success in Linux, the eBPF runtime has been ported to other operating systems such as Windows.[4]
You can crash a kernel with a BPF program. But it's overwhelmingly likely that the crash will arise from buggy pre-existing kernel code that just hadn't been seriously exercised before eBPF gave people new tools to push that code with. What's much, much less likely to happen is a segfault or NPE in your own BPF code.
I do wonder if this is the case for many people. It seems verifier is a bit unpredictable and makes eBPF programming quite painful.
Or do you mean something else?
But in this case I think they mean on the same machine. "In production" would be more accurate than "in a cloud environment". And yeah I wouldn't load custom kernel modules in production just to do observability.
The bpftools maintainers tell you to learn the bytecode format when you ask them what the errors mean, because they expect you to understand what the verifier means when it tells you "unknown scalar" on every single goddamn line of code.
Something like "ebpf coding rules, what to use and what not" would be very helpful.
All ebpf examples that are older than say, 3 months, already don't work anymore. Not even the official ones from the XDP tutorial project (and the libxdp maintainers because the kinda are splitting off a lot of headers into a separate xdp library as it seems).
Most userspace code still relies on the 5 years old bpf-helpers.h, which meanwhile is not supported anymore because it doesn't use the __helper methods from the kernel (they also refactored the kernel in the meantime, and force you to use e.g. __u128 instead of native data types).
Oh boi, did I underestimate what "bytecode vm" means when kernel developers talk about it.
Also, always use llvm, and remember to build two bpf files for each endianness, and use -g for debug symbols. Otherwise you will try to find out what the bpftool errors mean for days, because of shitty mailing list answers.
Unsolvable problem: reject any loop that is provably infinite.
Solvable problem: reject any loop that isn't provably finite.
The trade off is that there will always be some loops that in fact always do terminate, but that the verifier can't prove do.
while (condition) { … }
Do: #define MAX 1000 for n = 0; n < MAX; n++ { if !condition break; … }
Unroll all loops, don’t allow any backward jumps and limit to (say) 1m instructions.
while(condition)[1000 label]{…}
What could possibly go wrong.
Way better than running a denial of service attack on your own systems or those of your customer's.
That certainly depends on what eBPF is used for. If your load balancer errors out at [greatest number of connections envisioned] and an adversary manages to establish [greatest number of connections envisioned] then the result is a denial of service.
Not every operator is confident in making code changes in 3rd party software or might even be allowed to make such changes. Increasing resources o.t.o.h., e.g. adding RAM, is rarely banned. I sure would want a system to make best use of available resources.
If the code is intended to use as a library or the binary distributed to third parties one will have to handle it differently. For libraries taking a parameter indicating the maximum expected is common, for example. See e.g. man 3 read.
What custom usage do you have for it?
"My report "What is eBPF?" and in-depth book "Learning eBPF" are both available for download [0] from Isovalent or with your subscription to O'Reilly's learning platform. You can buy "Learning eBPF" from any good bookstore (support your local bookshop by ordering it there!)"
IIUC, you need to give contact info on that page to get the PDF. So this [1] might be a better starting point.
[0] https://isovalent.com/ebpf/
[1] https://ebpf.io/
I clicked the first link so I got the gist but how many other people just give up and disengage with their post? Or click the link then just close the tab without reading further.
It runs sandboxed kernel extensions? Or it's a VM? Something like that? The writing of the page itself doesn't inspire a lot of confidence, but then maybe it's just over my head.
Yes to both of those.
The kernel has a bunch of extension points that can run eBPF code in a VM. That code can make decisions for the kernel and/or track events.
eBPF code can do basically any calculation you want, but it can't have infinite loops.
It's loaded as bytecode, with a spec for how it's formatted and what the instructions do and what data structures are built in to the VM.
The main benefit is that it runs in the kernel, so it can be triggered very very often with minimal performance impact.
I can highly recommend it if you're eBPF curious =D I guess my only gribe is that the latter parts of the book, is a bit product heavy, but done in the most tasteful manner it could probably have been done, i.e. using it as examples on how to use eBPF.
If you want full control then kernel module is the way to go, but this doesn't have the same security and stability guarantees.
it's like a reverse backronym
https://www.lastweekinaws.com/podcast/screaming-in-the-cloud...
If my goal is to reduce the cpu cloud bill, is BPF good enough compared to making a kernel module ?
Lots of network drivers and NICs support offloading XDP programs ("xdp_prog") to the network controller's chipset, which results in zero CPU i/o interrupts if you e.g. use an XDP_DROP to block traffic.
Being able to block network traffic _before_ it reaches even kernelspace is a game changer.
As for this repository - the README triggered a false alarm in my bullshit sensors, but the code example is pretty nice.