Since then I gave a short (15 min) talk about producing and understanding flame graphs: http://oirase.annexia.org/tmp/2023-03-08-flamegraphs.mp4
Since then I gave a short (15 min) talk about producing and understanding flame graphs: http://oirase.annexia.org/tmp/2023-03-08-flamegraphs.mp4
- ARM frame pointer handling is more compact (relative to the non-frame pointer code) than that on RISC-V.
- ARM code does less frame pointer handling.
The first might be true because ARM has the stmdb instruction which (https://developer.arm.com/documentation/ddi0406/b/Applicatio...) stores multiple registers to consecutive memory locations using an address from a base register. The consecutive memory locations end just below this address, and the address of the first of those locations can optionally be written back to the base register.
So, code can save registers on the stack _and_ decrease the stack pointer in a single instruction.
The second might be true because, of the compilers/compiler settings used for those measurements the ARM one did more aggressive function inlining.
My comment on that article still stands: https://news.ycombinator.com/item?id=34660474
tl;dr: don't make code generation worse for everyone else just to appease your one tiny use-case; fix your goddamn tools instead.