LLVM 3.9 Release Notes
llvm.org
llvm.org
LTO is the compiler feature I'm waiting to become mainstream. Without it you have to make compromises with translation units (put more stuff in headers than you'd want to) to get the performance you want.
[0] http://blog.llvm.org/2016/06/thinlto-scalable-and-incrementa...
If LTO is similar that sounds like a massive win and something to look forward to.
[edit] Looks like LTO is more about performance than perf+binary size. Still sounds like awesome stuff.
Previously, before Google built ThinLTO, we built LIPO (and WHOPR) for GCC, https://gcc.gnu.org/wiki/LightweightIpo Google isn't going to use GCC as it's compiler anymore, and in fact, has mostly moved off it. We needed to replace LIPO, so we built ThinLTO.
The problem with LTO models in general (as the thinlto slides go over) is how to make it scalable. Part of that is memory, part of it is parallelization, part of it is optimization/analysis speed.
THe memory issues are about how memory efficient the IR is and how demand driven IPO is. IE can your whole-program analysis go function by function, without it taking a lot of time and memory to import individual functions (and throw them away if you need to).
Most LTO is slow because it tries to keep large parts of the program in memory at one time, and has no good way of being as lazy as possible about pieces of the program.
The parallelization issue is "what is being serialized to make LTO work". Some compilers parallelize the backend code generation, some parallelize nothing. There is often still a lot of serial stuff (global analysis, etc).
The CPU time issue is how fast you perform analysis and can generate code.
ThinLTO is meant to allow you to parallelize basically every step, import exactly the parts of the program you want at a given time in a quick and memory efficient way, etc.
Back when we'd use LTCG it was a ~30 minute link time penalty which prohibited it from all but the final release builds. Seeing this stuff move forward to be more effective to use is a great improvement.
So I've been having some "fun" with LLVM targeting AArch64, and was rather, er, disappointed when I discovered the toolchain didn't use GNU ld, but Mach's ld64.
Now in theory, ld64 is a next-gen linker that can link code not at the "module" (one compiled C file) level, but at the individual function and (global) variable level. Cool! More granularity! Except it doesn't support anything equivalent to GNU ld's extremely useful linker scripts, the stuff that says "put this code in SRAM, that data in Flash. etc..." and its command-line is... poorly documented. Also, it only works with Mach-O binaries, which is incompatible with objdump, that swiss-army knife of object code [1].
And anyhow, what does LLVM refer to as LLD? Because there are actually two separate projects under that umbrella: The ELF/COFF one (http://lld.llvm.org/NewLLD.html) and the ATOM-based one (http://lld.llvm.org/AtomLLD.html)
As I understand, the ATOM-based linker is still experimental, so they're referring to the ELF/COFF one?
TL;DR: LLVM doesn't hold a candle to binutils
[1] I recommend trying out using objdump to turn your app data into object data you can link directly into your application. This is what modern toolchains call "resources", but it's really interesting to test it out at the C/"bare-metal" level to see how it works. Bonus: you get to really learn the difference between pointers and arrays ;)
My understanding is that the darwin flavor uses the atom-based architecture, while the gnu and windows/link flavors use a section-based architecture. There was a thread on the llvm-dev mailing list that explains this in more detail. [2]
I don't know much about the linker script side of things, but I was under the impression that they were working on that.
[1] http://lld.llvm.org/Driver.html [2] http://lists.llvm.org/pipermail/llvm-dev/2015-May/085088.htm...
You can of course use GNU ld with Clang/LLVM if you so choose; this is the default on FreeBSD/arm64 today. The new point with 3.9 is that the full Clang/LLVM + lld toolchain can self-host (using the lld ELF support). ELF lld supports a significant subset of the linker script syntax, and additional functionality is being added as actual uses are found. (It's sufficient to link the FreeBSD kernel today.)
...but not Mach-O, Apple's format.
[1] https://en.wikipedia.org/wiki/Binary_File_Descriptor_library
And if you do need to build for multiple targets, you just run the binutils tools multiple times with the right options for each, because they're never identical.
Compiling per-target toolchains is an annoying chore, and binutils is the last holdover bugging me with this when I work with Rust.
The question here is whether you can run the same tool (the same binary) for two different targets, or if you have to run a different version of e.g. ld built specifically for each target.
While clang -flto can produce up to ~15% faster binaries than with gcc, it's still not usable for bigger projects which put a lot of their API into a shared library.
Some of those functions in a shared library are eventually inlined, but with clang -flto the exported copies of the functions are optimized away, whilst gcc keeps the copies additionally to the inlined variants. So you can try to keep those inlined API calls in the shared library with __attribute__((used)), but then they are not inlined anymore with performance regressions up to %50.
Not fixed in clang 3.8 and not in 3.9. I hope it will become usable in 4.0 as the performance benefits would be dramatic and the implementation efforts to keep the ((used)), i.e. exported copy, minimal.
How could foo() had ever printed "X"?
The code itself isn't well-formed because %ptr isn't actually defined anywhere—just pretend that it was. ;)
[0] Technically, LLVM does not have variables (as in, names you can assign to more than once), since it's in SSA form, but I don't know what a better name would be here. :)
Edit: simplify
Looks like D is gaining some traction recently :)