Link-Time Optimisation (LTO)
convolv.es
convolv.es
LTO was available in llvm forever, and well predates GCC, clang is just a particular frontend.
RMS finally allowed the GCC IR to be saved to disk in part as a response to llvm doing it and becoming popular
(I worked on both GCC and llvm forever)
It also misses the function summary vs non summary modes, which is fairly important.
For the initial version of this guide, I chose to focus on features as they are exposed in specific tools, as it's easier to pin down a "first available" date that way. Basic LTO for Clang is listed as first available in Clang 2.6, which is also the first LLVM release to officially include Clang.
> It also misses the function summary vs non summary modes, which is fairly important.
Are you talking about the module summaries LLVM uses in its parallel / thin LTO mode, or something else?
JVM implementations (and ART), alongside CLR do such optimizations.
Furthermore, nowadays PGO is part of the process as well, so the information can be carried over across runs, used as means for the JIT to quickly achieve the optimal execution point, and carry over from there instead of starting over from scratch every time.
In fact, it's even more wrong than that. The ability for managed runtimes to perform LTO across dynamic code loading means you're even able to get LTO across code your compiler has never witnessed - e.g. plugins written by 3rd parties.
https://devblogs.microsoft.com/dotnet/performance-improvemen...
Does anyone know of articles that go into a bit more on precisely the kinds of LTO they offer? I'll update the guide once I understand the situation a bit better.
Here is a 2015 paper for OpenJDK,
https://cr.openjdk.org/~vlivanov/talks/2015_JIT_Overview.pdf
Modern ART has a bit of everything, Assembly written interpreter for fast startup, followed by a JIT stage with PGO capabilities, followed by an AOT compiler that AOT compiles (with LTO) when the device is idle, and uploading PGO data into the Play Store, so that incrementally the same devices collaborate to the optimal PGO data set.
https://source.android.com/docs/core/runtime/jit-compiler
IBM and Azul JVMs have similar approaches with their cloud based JIT infrastructure.
https://developer.ibm.com/articles/jitserver-optimize-your-j...
https://www.azul.com/products/intelligence-cloud/cloud-nativ...
0) https://docs.oracle.com/javase/specs/jvms/se7/html/jvms-5.ht...
[1]: https://learn.microsoft.com/en-gb/cpp/build/reference/gl-who...
[2]: https://learn.microsoft.com/en-gb/cpp/build/reference/ltcg-l...
These two options also control profile-guided optimisation.
MSVC's LTCG (LTO) pipeline has been heavily in use at Microsoft for decades now. PGO and LTCG go together. LTCG is the default way most of our binaries are built.
A lot of work has been done to make MSVC LTCG fast and scalable. In particular, the compiler backend is multithreaded on a per-function basis, and there is also support for incremental LTCG that reuses previously compiled machine code and summary information.
https://www.sqlite.org/amalgamation.html
> And because all code is in a single translation unit, compilers can do better inter-procedure and inlining optimization resulting in machine code that is between 5% and 10% faster.
I prefer to use unity builds for various reasons.
You don't need to if you're using the LLVM linker (lld) - by default it will use all of your HW threads (nr of cores).
lld --help | grep thread
--threads=<value> Number of threads. '1' disables multi-threading. By default all available hardware threads are usedFrom https://clang.llvm.org/docs/ThinLTO.html#id10
By default, the ThinLTO link step will launch as many threads in parallel as there are cores. If the number of cores can’t be computed for the architecture, then it will launch std::thread::hardware_concurrency number of threads in parallel. For machines with hyper-threading, this is the total number of virtual cores.
But maybe I misunderstood you.CMake CMAKE_INTERPROCEDURAL_OPTIMIZATION variable, when set, opts in for LTO so, yes, due to LTO design you will inherently lose concurrent linkage.
Invariably, though, the executable is at least a few MB larger (a bad trade on systems like the Nintendo Switch where memory is scarce) and the incremental build time is intolerable.
I've always wondered why it's never brought the promised free performance; I can only assume our frequent profiling and judicious definition of hot functions inline has gotten us almost everything LTO would bring anyway.
For our (Rust) project, thin LTO doesn't cost much extra in compile time and gives a few percent speedup. But it's not a terribly optimized code base, and we don't have any manual inline annotations etc.
My current understanding is: just add -flto to CFLAGS and CXXFLAGS, it does wonders and there is no downsides(other than longer build time). correct me if I am wrong please
It does more than optimize.
If anything you might want to explicitly do -fno-lto to not use LTO when you have fat objects.
My understanding is that LTO itself is not breaking anything. In these cases the software that it was running on was broken in some subtle way that didn't show up before.
Linux distros like Gentoo that can try to compile lots of stuff with LTO find a lot of bugs related to this.
There are cases where it causes runtime errors as well.
Sometimes software is written incorrectly and has "undefined behavior" (you probably know this already). The thing with UB is that it may or may not cause a visible bug at runtime. Sometimes these bugs go unnoticed for a long time.
It's really common for code with UB to run without visible errors when compiled without optimizations, but when you compile it with optimizations it segfaults at runtime. LTO is adding another form of optimization and sometimes the transformations the compiler made just happen to cause already existing UB to "surface".
This is what I meant by LTO "causing" runtime issues to appear!
There are cases of this happening in the Gentoo bug tracker, and in the Linux kernel!