New Relic to open-source Pixie’s eBPF observability platform
blog.pixielabs.ai
blog.pixielabs.ai
eBPF is indeed a part of the puzzle. It allows us to access telemetry data without any manual instrumentation when running on Linux machines.
Pixie itself is extendable and currently ingests data from many other sources as well. Joining forces with New Relic will allow us to focus on expanding the open-source project, but also expand our capabilities by plugging into other open APIs and frameworks such as OpenTelemetry, Grafana, Prometheus.
https://github.com/pixie-labs/pixie/tree/main/demos/simple-g...
p.s. The slides would be nice to have too :-)
Slides: https://www2.slideshare.net/ZainAsgar/no-instrumentation-gol...
Write up: https://blog.pixielabs.ai/ebpf-http-tracing/ https://blog.pixielabs.ai/ebpf-function-tracing/post/
Why do you need the kernel support if you modify the binaries, why not insert a function to write your logs and then insert a call to that function rather than relying on kernel support via an int 3?
Curious really.
But I was explaining the most common and obvious (for a user) use of eBPF.
They should just look at the io visor project to see some of the stuff that can be done with it (disclaimer, I work with one of the io visor maintainers)
https://developer.apple.com/library/archive/documentation/An...
* Non intrusive: meaning one can snoop info of application without changing application code.
* Deep visibility: function level and syscall/kernel functions reveal more context and are more accurate in a lot of cases.
* Low overhead: everything runs inside kernel space, no context switching compared to other kernel based/aided tracing.
* Expressiveness: eBPF is fairly expressive, can do many things that usually are exclusive to high level programming languages.
We've been eagerly awaiting some customers to adopt newer kernels so we can start leveraging eBPF because of the performance gains in these type of scenarios.
Getting down the the kernel often can help find problems with disk access or network issues.
One of our staff engineers is exploring it now for NFS stats: https://gitlab.com/wchandler/tracing-tools/-/blob/master/nfs...
In Support Engineering we often straddle the line of 'SRE style stare at graphs and configuration as code' and 'log on to the box and look at syscalls'. We are very very excited about eBPF.
Deploys eBPF kprobes (based on bpftrace) and uprobe (based a custom front-end language) and instantly get rich data (arts, return value, latency), query data in a Pandas-like scripting language and visual dashboard.
On the other hand - it already is overhauling service meshes, VPNs and firewalls and network security policies, etc. The stuff fly.io is debuting now is probably going to be standard in few years.
This is outside my area of expertise so maybe I'm just missing some deeper insight - and I can see where you're coming from for traditional "throw the artifact over the wall, good luck running it" system operations kind of stuff. Or debugging services in production, which tracing is a key part of - but again just a small part of observability, and one that the rest of your dev process should be actively trying to minimize. If you have any degree of DevOps going on, many key SLIs will much higher-level (p99 of HTTP requests, MB of storage per customer), and I don't see how eBPF addresses that better than existing instrumentation, or in some cases at all.
https://techcrunch.com/2020/11/24/splunk-acquires-network-ob...
[blatantly promoting my substack] Been following this area about a year, first wrote about some of startups using eBPF here in late 2019: https://monitoring2.substack.com/p/ebpf-a-new-bff-for-observ...
In general, academic rule applies: Never write an article without defining terms beforehand.
https://docs.newrelic.com/docs/apm/new-relic-apm/getting-sta...
In the current form Pixie Cloud is required since it's a freemium product.
We are planning to open-source a self-managed version of Pixie early next year, and it can be used completely self-hosted.
This definitely is an unprecedented and forward looking investment by New Relic. While ambitious, it became evident in our conversations with them they we committed on standardizing on open source telemetry standards such as Prometheus, Open-Telemetry, Graphana etc.
We're not yet in a position to speak for them in detail but we believe this bet on our project reinforces their plan to open-source telemetry layer to accelerate the adoption of observability practices by developers.
Hope that adds some color! would be great to hear more thoughts.
Sysdig was a pioneer in harvesting data from the kernel. Their original solution required installing a kernel module and they are now moving to eBPF based approaches. The Falco project is really exciting.
Since we're a relatively new project (started 2 years ago) we started with eBPF and built our platform around it. As we open source we'll share with groups like Falco and hopefully collaborate.
We are preparing to open source early next year, so stay tuned. The code will be posted on our repository, that currently hosts our community scripts: https://github.com/pixie-labs/pixie.