So this is a solution for a problem, and it arrives just at the moment that people have solved the problem more generically ;)
(Prodfiler coauthor here, we had solved all of this by the time we launched in Summer 2021)
So this is a solution for a problem, and it arrives just at the moment that people have solved the problem more generically ;)
(Prodfiler coauthor here, we had solved all of this by the time we launched in Summer 2021)
But I think you're only thinking about CPU profiling at <= 100 Hz / core. However, Brendan's article is also talking about Off-CPU profiling, and as far as I can tell, all known techniques (scheduler tracing, wall clock sampling) require stack unwinding to occur 1-3 orders of magnitude more often than for CPU profiling.
For those use cases, I don't think .eh_frame unwinding will be good enough, at least not for continuous profiling. E.g. see [1][2] for an example of how frame pointer unwinding allowed the Go runtime to lower execution tracing overhead from 10-20% to 1-2%, even so it was already using a relatively fast lookup table approach.
[1] https://go.dev/blog/execution-traces-2024
[2] https://blog.felixge.de/reducing-gos-execution-tracer-overhe...
https://gitlab.com/freedesktop-sdk/freedesktop-sdk/-/issues/...
So not a perf issue there, but they don't think the workflow is suitable for whole-system profiling. Perf issues were in the context of `perf` using DWARF:
https://gitlab.com/freedesktop-sdk/freedesktop-sdk/-/issues/...
So it's both similarly fragile, but one is almost never disabled.
The broader point is: For HLL runtimes you need to be able to switch between native and interpreted unwinds anyhow, so you'll always do some amount of lifting in eBPF land.
And yes, having frame pointers removes a lot of complexity, so it's net a very good thing. It's just that the situation wasnt nearly as dire as described, because people that care about profiling had built solutions.
Things get even more complicated because context switches can mean CPU migrations, making many of your data useless.
> Things get even more complicated because context switches can mean CPU migrations, making many of your data useless.
No it doesn't. If a user space thread is blocked on doing kernel work, its stack isn't going to change, not even if that thread ends up resuming on a different thread.