BOLT: Binary Optimization and Layout Tool
github.com
github.com
> Inspired by the performance gains and to address the scalability issue of BOLT, we went about designing a scalable infrastructure that can perform BOLT-like post-link optimizations.
> Our experiments on large real-world applications and SPEC with code layout show that Propeller can optimize as effectively as BOLT, with just 20% of its memory footprint and time overhead.
Neither is made to do the optimizations you are talking about, that is what regular llvm does.
[1] http://lwn.net/1998/1029/als/rope.html
(An aisde: I was at that talk in 1998 at Atlanta Linux Showcase and remember it fairly well. I'm the guy who asked Nat if he'd considered using a genetic algorithm for the optimization processing.)
[0] www.hungry.com/~shaver/gropt.tar.bz2
Why not hook this into LLVM's compiler infrastructure, instead of recovering the CFG from a binary - and not any binary, but one that has to built with special flags anyway? This seems backwards.
i think this paper (cant find a non-paywalled copy) https://dl.acm.org/doi/10.1145/12276.13338 got around 10% just by attempting to align the allocation of the callers and the callee.
maybe memories are large enough now, and the whole proprietary binary-only library business model is gone, that we should really be abandoning the idea of seperate compilation units as the default.
For example if your code calls a particular runtime function shortly after starting, that function (or the basic blocks from that function that are executed frequently) can be placed close to the call site.
LTO requires that you’re compiling all the code to get the benefits.
I’m also not sure if either gcc or clang’s LTO does layout at the basic block level or at the function level. There is a lot of benefit from doing it at the basic block level. You can literally lay out all the code that runs at application start time so that it is mostly adjacent and fall-through from block to block during execution.
This isn’t something you necessarily do instead of LTO, but rather something you can do in addition to it.
What kind of perf benefits can you expect?
What kind of application have you tested this with?
"For datacenter applications, BOLT achieves up to 7.0% performance speedups on top of profile-guided function reordering and LTO."
This is used for large binaries, like hhvm. There's a lot more details here: https://research.fb.com/publications/bolt-a-practical-binary...
Propeller looks much better now.
http://www.bitsavers.org/pdf/mips/RISCos/3204DOC_RISC_os_Use...
I wonder if those have been open sourced.