I think there's a lot of value in LLM compilers to specifically be used for superoptimization where you can generate many possible optimizations, verify the correctness, and pick the most optimal one. I'm excited to see where y'all go with this.
https://github.com/SuperOptimizer/supercompiler
There's code there to generate unoptimized / optimized pairs via C generators like yarpgen and csmith, then compile, train, inference, and disassemble the results
What do you mean the tech isn't there yet, why would it ever even go into that direction? I mean we do those kinds of things for shits and giggles but for any practical use? I mean come on. From fast and reliable to glacial and not even working a quarter of the time.
I guess maybe if all compiler designers die in a freak accident and there's literally nobody to replace them, then we'll have to resort to that after the existing versions break.
What they missed is to mention verification (they probably don't know about alive2) and comparison with other compilers. It is very likely that LLM Compiler "learned" from GCC and with huge computational effort simply generates what GCC can do out of the box.
The problem with using alive2 to verify LLM based compilation is that alive2 isn't really designed for that. It's an amazing tool for catching correctness issues in LLVM, but it's expensive to run and will time out reasonably often, especially on cases involving floating point. It's explicitly designed to minimize the rate of false-positive correctness issues to serve the primary purpose of alerting compiler developers to correctness issues that need to be fixed.
Additionally, they train approximately half on assembly and half on LLVM-IR. They don't talk much about how they generate the dataset other than that they generated it from the CodeLlama dataset, but I would guess they compile as much code as they can into LLVM-IR and then just lower that into assembly, leaving gcc out of the loop completely for the vast majority of the compiler specific training.
It seems like somehow build systems were invoked given the different targets present in the final version?
Was it mostly C/C++ (if so, how did you resolve missing includes/build flags), or something else?
> (A <=> B) < 0 is true if A < B
> (A <=> B) > 0 is true if A > B
> (A <=> B) == 0 is true if A and B are equal/equivalent.
TIL of the spaceship operator. Was this added as an april fools?
It's useful for stable-sorting collections with a single test. Also, overloading <=> for a type, gives all comparison operators "for free": ==, !=, <, <=, >=, >
In C++20 the compiler will automatically use the spaceship operator to implement other comparisons if it is available, so it's a significant convenience.
In practice though, correctness even over ordering of hand-written passes is difficult. Within the paper they describe a methodology to evaluate phase orderings against a small test set as a smoke test for correctness (PassListEval) and observe that ~10% of the phase orderings result in assertion failures/compiler crashes/correctness issues.
You will end up with a lot more correctness issues adjusting phase orderings like this than you would using one of the more battle-tested default optimization pipelines.
Correctness in a production compiler is a pretty hard problem.
- foundation model is pretrained on asm and ir. Then it is trained to emulate the compiler (ir + passes -> ir or asm)
- ftd model is fine tuned for solving phase ordering and disassembling
FTD is there to demo capabilities. We hope people will fine tune for other optimisations. It will be much, much cheaper than starting from scratch.
Yep, correctness in compilers is a pain. Auto-tuning is a very easy way to break a compiler.
What would be more interesting is training a large model on pure (code, assembly) pairs like a normal translation task. Presumably a very generalized model would be good at even doing the inverse: given some assembly, write code that will produce the given assembly. Unlike human language there is a finite set of possible correct answers here and you have the convenience of being able to generate synthetic data for cheap. I think optimizations would arise as a natural side effect this way: if there's multiple trees of possible generations (like choosing between logits in an LLM) you could try different branches to see what's smaller in terms of byte code or faster in terms of execution.
ChatGPT does this, unreliably.
> What would be more interesting is training a large model on pure (code, assembly) pairs like a normal translation task.
It is that.
> Presumably a very generalized model would be good at even doing the inverse: given some assembly, write code that will produce the given assembly.
Is has been trained to disassemble. It is much, much better than other models at that.
Personally I thought we were way too close to perfect to make meaningful progress on compilation, but that’s probably just naïveté
Even just looking at inlining for size, there are multiple recent studies showing ~10+% improvement (https://dl.acm.org/doi/abs/10.1145/3503222.3507744, https://arxiv.org/abs/2101.04808).
There is a massive amount of headroom, and even tiny bits still matter as ~0.5% gains on code size, or especially performance, can be huge.