Generally, the “compile with frame pointers” request is a bit underspecified on x86-64. For example, the perf trampoline that is cited as the reason why future Python versions are going to use frame pointers itself does not have a frame pointer. (It does not clobber it, either, but backtraces through it will skip a frame.) Similarly, if you just compile glibc with frame pointers without patching a handful of assembler files, then for typical workloads, 5% to 10% of your samples will have an incomplete backtrace because the immediate caller of glibc string functions is not recorded (they don't have a frame pointer, either).
I expect this x86-64 debate will be obsolete Really Soon Now because everyone will simply copy out the shadow stack on every sample. It's even faster than frame pointer traversal because it's just an array copy. We are still looking at making DWARF backtraces a bit faster (or even the more complex unwinding case), but I doubt that this work will be impactful, given the politics involved. The shadow stack will have some performance impact, too, but it can be switched off on a per-process basis if necessary.
See for instance Apple's ARM64 calling convention: https://developer.apple.com/documentation/xcode/writing-arm6...
"The frame pointer register (x29) must always address a valid frame record. Some functions — such as leaf functions or tail calls — may opt not to create an entry in this list. As a result, stack traces are always meaningful, even without debug information."
[0]: https://fedoraproject.org/wiki/Changes/fno-omit-frame-pointer
[1]: https://lists.fedoraproject.org/archives/list/devel@lists.fedoraproject.org/thread/OOJDAKTJB5WGMOZRXTUX7FTPFBF3H7WE/
edit: whoops, this was already mentioned in the blog post. Should read before commenting :)BTW, I'm a big fan of your work, thanks for Flask and Jinja!
That said, we should still try to convince people!