What everybody in this discussion seems to miss is that you don't need to unwind the DWARF data structures during profiling time, you are free to convert DWARF to a fast-lookup data structure on the machine.
DWARF needs to support every CPU under the sun. Every unwinder on the other hand is CPU-specific. For prodfiler.com's continuous in-production unwinding, we convert DWARF into something compact and fast-to-lookup that is then placed in eBPF maps.
It all works like a charm. We can have our cake (e.g. use RBP as GPR) and eat it too (e.g. use .eh_frame, converted at runtime into a fast-to-lookup format) to do reliable whole-system unwinding in production.