Clang IR (CIR): A New IR for Clang
facebookincubator.github.io
facebookincubator.github.io
The question, however, is why this particular IR implementation, given there are a number of ongoing efforts to do general C/C++ MLIR implementations. The urgency for upstreaming this change is a little weird, as i'd expect the community to take real time to investigate and think through the IR choices.
I attended an LLVM conference earlier in the year, and they are a smart bunch with really deep understanding of this stuff, so I expect they will have opinions as to whether this is the right implementation, and whether this is the right time to adopt it.
On a side note, one of the dangers of MLIR as far as I can see is an unfortunate side effect that optimisations performed on the language IR will not be available to other languages which use either LLVM IR or their or MLIR variant. This will either mean there are lost optimisation opportunities, or duplicate effort to port passes to other MLIR variants. That'll be a shame.
If that's a problem then isn't currently also a issue with the other MLIR variants? I also wonder if LLVM couldn't reduce the need for a language fork of MLIR somehow.
I am working on the frontend for a toy lang, and was planning to transpile to LLVM IR. This is of interest to me. But its frustrating how opaque the dev docs for such projects are.
Maybe because it's not "another IR for LLVM", but an IR for Clang?
Anyway, if you actually care: https://discourse.llvm.org/t/rfc-an-mlir-based-clang-ir-cir/...
Quoting here for benefit of others;
"Motivation
In general, Clang’s AST is not an appropriate representation for dataflow analysis and reasoning about control flow. On the other hand, LLVM IR is too low level — it exists at a point in which we have already lost vital language information (e.g. scope information, loop forms and type hierarchies are invisible at the LLVM level), forcing a pass writer to attempt reconstruction of the original semantics. This leads to inaccurate results and inefficient analysis - not to mention the Sisyphean maintenance work given how fast LLVM changes. Clang’s CFG is supposed to bridge this gap but isn’t ideal either: a parallel lowering path for dataflow diagnostics that (a) is discarded after analysis, (b) has lots of known problems (checkout Kristóf Uman’s great survey 31 regarding “dataflowness”) and (c) has testing coverage for CFG pieces not quite up to LLVM’s standards.
We also have the prominent recent success stories of Swift’s SIL and Rust’s HIR and MIR. These two projects have leveraged high level IRs to improve their performance and safety. We believe CIR could provide the same improvements for C++."
> Maybe because it's not "another IR for LLVM", but an IR for Clang?
AFAIK, Clang also uses LLVM IR.
Clang doesn't use an IR for the frontend stuff; it emits AST straight to LLVM. It's built for speed and not very flexible. It can be difficult to see what some features like ObjC ARC are doing since there's no way to see an "intermediate" representation without debugging the compiler.
As I understand it, those language-specific IRs are closer to the language, why LLVM IR is closer to a cross-platform assembly dialect (or at least doesn't preserve much information from the original language).
[0] https://blog.rust-lang.org/2016/04/19/MIR.html
[1] https://mitchellh.com/zig/astgen#what-does-zir-look-like
Rust -> HIR (higher intermediate representation) -> MIR (middle intermediate representation) -> LLVMIR
Different stages are differing levels of complexity and do different checks
What these extra IRs afford you is more granularity. More points to choose from along that high-level/low-level gradient mean you also have a better chance at the right balance of high-level and low-level information for your optimisation pass to work.
The LLVM’s discourse RFC goes in depth about the project motivation
Link to the discourse RFC:
https://discourse.llvm.org/t/rfc-an-mlir-based-clang-ir-cir/...
Motivation:
> In general, Clang’s AST is not an appropriate representation for dataflow analysis and reasoning about control flow. On the other hand, LLVM IR is too low level — it exists at a point in which we have already lost vital language information (e.g. scope information, loop forms and type hierarchies are invisible at the LLVM level), forcing a pass writer to attempt reconstruction of the original semantics. This leads to inaccurate results and inefficient analysis - not to mention the Sisyphean maintenance work given how fast LLVM changes. Clang’s CFG is supposed to bridge this gap but isn’t ideal either: a parallel lowering path for dataflow diagnostics that (a) is discarded after analysis, (b) has lots of known problems (checkout Kristóf Uman’s great survey 32 regarding “dataflowness”) and (c) has testing coverage for CFG pieces not quite up to LLVM’s standards.
However, Clang remained unchanged. So what Facebook is doing now is applying this IR refactoring that had been successful with Swift to Clang. The result, CIR, will allow a better structured, faster and more secure Clang front end [1].
This is not some magical nonsense and you're not going to get magical results. This is just solid engineering and is based on an already successful pattern.
[1] https://www.phoronix.com/scan.php?page=news_item&px=Meta-Dev...
> In general, Clang’s AST is not an appropriate representation for dataflow analysis and reasoning about control flow. On the other hand, LLVM IR is too low level — it exists at a point in which we have already lost vital language information (e.g. scope information, loop forms and type hierarchies are invisible at the LLVM level), forcing a pass writer to attempt reconstruction of the original semantics. This leads to inaccurate results and inefficient analysis - not to mention the Sisyphean maintenance work given how fast LLVM changes.
[1] https://discourse.llvm.org/t/rfc-an-mlir-based-clang-ir-cir/...
> Our current (and initial) goal is to provide a framework for improved diagnostics for modern C++, meaning better support for coroutines and checks for idiomatic uses of known C++ libraries.
...
> In general, Clang’s AST is not an appropriate representation for dataflow analysis and reasoning about control flow. On the other hand, LLVM IR is too low level — it exists at a point in which we have already lost vital language information (e.g. scope information, loop forms and type hierarchies are invisible at the LLVM level), forcing a pass writer to attempt reconstruction of the original semantics. This leads to inaccurate results and inefficient analysis - not to mention the Sisyphean maintenance work given how fast LLVM changes. Clang’s CFG is supposed to bridge this gap but isn’t ideal either: a parallel lowering path for dataflow diagnostics that (a) is discarded after analysis, (b) has lots of known problems (checkout Kristóf Uman’s great survey 28 regarding “dataflowness”) and (c) has testing coverage for CFG pieces not quite up to LLVM’s standards.
> We also have the prominent recent success stories of Swift’s SIL and Rust’s HIR and MIR. These two projects have leveraged high level IRs to improve their performance and safety. We believe CIR could provide the same improvements for C++.