Yeah, doing an 8KB memcopy for every profile sample sounds like a lot of overhead. Is DWARF unwinding so slow that that is actually faster?
Virgil doesn't use a frame pointer, and I sort of regret it. It uses custom unwinding information that is used during GC or throwing an exception (i.e. controlled crash). I spent considerable time optimizing both the space and time of that lookup, to the point where it's only 32 bits of metadata per call site and a few dozen instructions to walk each frame. But that was a major, major pain to debug and I found a bug in it at late as last year.
A frame pointer is also required for stack allocation of objects that aren't fully scalar-replaceable. Virgil doesn't do that yet, but aspires to someday.