Compilers and IRs: LLVM IR, SPIR-V, and MLIR
lei.chat
lei.chat
Our goals with this pipeline are to enable static analyses that can choose the right abstraction level(s) for their goals, and using provenance, cross abstraction levels to relate results back to source code.
Neither Clang ASTs nor LLVM IR alone meet our needs for static analysis. Clang ASTs are too verbose and lack explicit representations for implicit behaviours in C++. LLVM IR isn't really "one IR," it's a two IRs (LLVM proper, and metadata), where LLVM proper is an unspecified family of dialects (-O0, -O1, -O2, -O3, then all the arch-specific stuff). LLVM IR also isn't easy to relate to source, even in the presence of maximal debug information. The Clang codegen process does ABI-specific lowering takes high-level types/values and transforms them to be more amenable to storing in target-cpu locations (e.g. registers). This actively works against relating information across levels; something that we want to solve with intermediate MLIR dialects.
Beyond our static analysis goals, I think an MLIR-based setup will be a key enabler of library-aware compiler optimizations. Right now, library-aware optimizations are challenging because Clang ASTs are hard to mutate, and by the time things are in LLVM IR, the abstraction boundaries provided by libraries are broken down by optimizations (e.g. inlining, specialization, folding), forcing optimization passes to reckon with the mechanics of how libraries are implemented.
We're very excited about MLIR, and we're pushing full steam ahead with VAST. MLIR is a technology that we can use to fix a lot of issues in Clang/LLVM that hinder really good static analysis.
Should be:
LLVM dialect in MLIR, which can be translated directly to LLVM IR
Otherwise, great project! We're also using MLIR internally and it's been awesome, game-changing even when considering how much can be accomplished with a reasonable amount of effort.
I think the next big problems for MLIR to address are things like: metadata/location maintenance when integrating with third-party dialects and transformations. With LLVM optimizations, getting the optimization right has always seemed like the top priority, and then maybe getting metadata propagation working came a distant second.
I think the opportunity with MLIR is that metadata/location info can be the old nodes or other dialects. In our work, we want a tower/progression of IRs, and we want them simultaneously in memory, all living together. You could think of the debug metadata for a lower level dialect being the higher level dialect. This is why I sometimes think about LLVM IR as really being two IRs: LLVM "code" and metadata nodes. Metadata nodes in LLVM IR can represent arbitrary structures, but lack concrete checks/balances. MLIR fixes this by unifying the representations, bringing in structure while retaining flexibility.
HN: two net downvotes for asking a question.
https://github.com/llvm/clangir/blob/main/clang/lib/CIR/Dial...
Our understanding of ClangIR is that it has side-stepped the problem of trying to relate/map high-level values/types to low-level values/types -- a process which the Clang codegen brings about when generating LLVM IR. We care about explicitly representing this mapping so that there is data flow from a low-level (e.g. LLVM) representation all the way back up to a high-level. There's a lot of value and implied semantics in high-level representations that is lost to Clang's codegen, and thus to the Clang IR codegen. The distinction between `time_t` and `int` is an example of this. We would like to be able to see an `i32` in LLVM and follow it back to a `time_t` in our high-level dialect. This is not a problem that ClangIR sets out to solve. Thus, ClangIR is too low level to achieve some of our goals, but it is also at the right level to achieve some of our other goals.
I knew people that wrote backends for gcc, and they pretty much all agreed it was a nightmare.
Most compilers in the 80s already used some form of non-ast IRs: Generalised portable instructions, cfg representations, etc. I'd say that back then people were experimenting with various forms of IRs for optimization purposes.
By 90s it was typical for compilers to have several IRs. Gcc belongs to this compiler generation.
2000s is the time when IRs standardised around 2-3 SSA-like forms, which made it possible to do most optimisation work on this level only.
http://www.chilton-computing.org.uk/acl/applications/hartran...
In fact, even the original fortran had quite an involved compiler, complete with numerous optimisations and sophisticated register allocation.
Fun read, btw.
Every one and their mother has their own proprietary MLIR dialect nowadays, and the era of competitive open source compilers is sort of fading.
The open source vs proprietary decision is usually a decision taken based on exactly this difference.
The sole reason for many companies to upstream some of their stuff is that they do not have to keep up with the firehose of upstream changes on their own.
I really think most of these proprietary users are just figuring things out as they go, and doing that upstream 1. wouldn't be accepted, and 2. wouldn't make sense for them in the first place. But I'll admit that's speculation, I don't have intimate knowledge of what everyone's doing :)
Out of curiosity, why?
Also a lot of textual shaders were already unreadable generated code. But for the ones that weren't, it's definitely a loss for those of us trying to peak at the internals of games, indeed :)
One of the problems of the article is how many compiler infra is re-made. But a big reason is that if you are doing Python, for example, you don't wanna touch C++.
In fact, nobody wanna touch C++, except(maybe!) C++ devs :)
I only do Python when doing systems programming in UNIX that outgrow bash scripts, or scritping software that uses Python as their scripting language.
Then again, I enjoy using C++ despite its flaws. :)
And for the look of it, look like is not something that is viable to do without using C++?