What's new for RISC-V in LLVM 15
muxup.com
muxup.com
Relaxation basically means having the linker redo a bunch of work the assembler otherwise would have done. Having them separate is then counterproductive.
The assembler is often spliced into the back end of the compiler to avoid the round trip through text on the way to machine code.
My conviction is that lowering to machine code in pieces then stitching them together is no longer necessary or advisable. Emit files containing some richer format - most obviously a compiler IR - and combine those before lowering the single blob to machine code.
I am not a compiler expert but this notion that there are optimisations to be had in the more expressive higher-level compiler IR and then delaying the emitting of machine code rings true. I would be grateful were anyone able to direct me to some papers etc on the subject.
Relocations are instructions executed by the linker (or loader). Write the address of this symbol to this offset from the start of the data section, stuff like that. Relaxations cover things like a branch target turned out to be closer than it might have been so change the instruction sequence to a smaller/faster one.
So if you don't eagerly convert to machine code, you don't need the link time relocation data, so the assembler and linker don't need the code to deal with it.
Of course if you stay in a friendly IR you can also do optimisations easily at link time, so most articles on this approach call it 'link time optimisation'.
Being able to link pre-compiled object files to a new program is certainly a bit more than "a kind of incremental compilation", in my opinion. It's also a fairly clean solution to the problem of not recompiling code that hasn't changed.
Linking sounds like a clean solution, but it really, really isn't. It's messy as all hell.
I think maybe, if I was to design a solution from scratch, I'd make object files basically a binary encoding of assembly code rather than true machine code; notably, I would keep jumps as abstract "jump to this label" instructions rather than proper machine code. This would put the assembler in the linker, where we actually have enough information to do instruction selection, rather than in the assembler where we don't.
We could also just put straight-up compiler IR in the object files, which puts a whole compiler back-end in the linker. This opens the door for LTO. Obviously, there are performance implications of this; if most of the hard work happens in the linker, linking can't be a stupid single-threaded serial operation anymore. But we could just make multi-threaded linkers and we get back our parallelism. We do lose some of the incrementalness of incremental compilation though.
Just a reminder that we're talking about LLVM here, whose whole point from the beginning was exactly what you described ;)
Linkers actually don't change instruction lengths. If they need to jump further away than what the instruction allows, they add trampolines, which are chunks of code with the appropriate instruction for the long jump, and place the trampoline at a good distance of the relocatable short jump.
Edit: well, except in the proposed case of relaxation, where they'd want to shorten jump instructions.
https://en.wikipedia.org/wiki/Architecture_Neutral_Distribut...
With executable stack and all that jazz?
Red Hat seems to have a =3
This is literally not what is happening in the real world though. Counter example: NEC SX Aurora (literally developed in the open https://github.com/sx-aurora-dev/llvm-project), Fujitsu A64FX (free LLVM based compiler, $$$$ proprietary home grown compiler), AMD with hipSYCL. The reason for this if there was no LLVM available, the compiler would be totally proprietary. This from scratch compiler would have been more expensive to make leading it to not being available for free but instead licensed at huge cost. There has been no uptake of GCC for this role since it's inception. We both know why.
And even if a $$$$ using LLVM is made i would prefer that over a fully custom compiler. I can link against it like a LLVM compiler and you can learn a lot about a unknown platform from how LLVM compiles for it, making reverse engineering a lot easier compared to a proprietary compiler.
Give it time. You know what happened with GCC on the NeXT computer; the only reason the NeXT compiler became free is because NeXT had to release it, since they based it on GCC.
> There has been no uptake of GCC for this role since it's inception. We both know why.
We do not. I was under the impression that the GCC stance on modularity has mellowed in recent times.
> using LLVM is made i would prefer that over a fully custom compiler.
False dichotomy. I would prefer GCC.