How to add eBPF observability to your product
brendangregg.com
brendangregg.com
Iirc the real challenge was with writing kprobe ebpf functions that access native structs. I don't think we ever found a good solution for that because you need kernel headers for each machine which we didn't have.
Of course if I'm missing something obvious, do tell!
(I'm the same person who left this comment on Brendan's site as well. Not a random copy paster)
Of course, tracepoints are better, as they are stable and don't need BTF (although raw_tracepoints do), but there's often cases where they don't exist for the thing I am tracing (and arguably shouldn't exist, if it's too niche/hot-path).
As someone who has been building observability software for the last few years I'm ridiculously excited about eBPF. Looking forward to seeing what the Netflix dashboards mentioned in the post look like & the data pipelines that support them.
Any time there's some popular public API, people will copy other people's projects built around this API and make them "better". I don't think Daniel Stenberg (author of cURL) is upset that there's HTTPie. I don't think it hurts cURL and I don't think it hurts the developer community as a whole.
I guess one could argue it would make life easier if there was just one HTTP library (or one eBPF library I guess). But that's just unrealistic, that's not how people work. Plus it removes the competition aspect which is nice to have for any project.
Is there any examples of open source projects that died because people were making too many ports?
Yes, Linux tracing.
I entered in 2014 when there were 14 different tracers, none of them finished (ie, does everything and works on all distros). It wasn't a lack of engineers -- there are plenty of smart engineers in tracing. But they were scattered among those 14 tracers. It was a mess, and no one could understand why Linux didn't have a mature tracing solution. (I think this also answers your question: Why cares if people make shitty ports?)
So I started thinking I could fix it by creating my own 15th tracer...Just kidding! What a terrible idea. Instead I joined one of those 14 existing tracers (a new one called BPF) and worked hard to get others on board, and do what we hadn't done well before -- collaborate!
Early on I wanted to do my own front-end, but Alexei convinced me to collaborate on bcc. I also wanted to do my own language, but got people to collaborate on the existing bpftrace instead.
The entire journey of BPF observability has been one of us choosing NOT to keep spinning up competing projects and instead to collaborate together on one. And that's how we succeeded in delivering a tracer that does everything on all distros.
I should add you might be missing a bit of context here too, and that's that tracing is maintenance heavy: You don't develop it once and then it works forever. It needs constant updates to match kernel changes etc.
Thank you for your explanation.
Modern BPF (eBPF) was created by Alexei Starovoitov and Daniel Borkmann, who are still maintainers but are joined by over a hundred contributors. I've spent most of my time contributing to the observability frontends, bcc and bpftrace. Apart from the people, the major companies contributing include Facebook and Isovalent; Microsoft are becoming its own major contributor to its own Windows implementation.
Wait, what now !?
After checking : he wasn't joking[0]. Good job MS.
I think the real problem is that BPF was built by engineers without professional marketing help, who would have come up with a better name!
In other words, not every site is designed around link aggregators like HN, and some things will be assumed, just as if you open a book to some random chapter.
> eBPF is an extended BPF JIT virtual machine in the Linux kernel
It's really quite an amazing tool, especially since you can use it to run sandboxed programs in kernel without changing kernel source or loading a kernel module. These programs can be written using a limited dialect of C.
Some examples of its use can be found here: https://github.com/iovisor/bcc/tree/master/examples
Brendan Gregg, the author of the posted blog entry, works with BPF on observability at Netflix and delivered a keynote at UbuntuMasters 2019. The video is on his blog is a great intro. [0]
I've been watching BPF for a couple of years now, and it seems slow on the uptake, but I hope Gregg is right that people will eventually start writing new drivers, firewalls, observability and security tools, loaded from userspace, but all running safely in a kernel vm, maybe even written in new async-first programming languages.
[0] https://brendangregg.com/blog/2019-12-02/bpf-a-new-type-of-s...