A deeper look into the GCC Rust front-end
lwn.net
lwn.net
This is, of course, not true. It allows only two things: (1) allow access through pointers and union fields (2) allow calling unsafe or external functions.
Neither of these is really a "check", more like "some operation that just isn't present in safe rust and is present in unsafe rust". They aren't conditionally allowed in safe rust; they just aren't allowed. In particular unsafe does not alter the operation of borrowck by one jot.
(And it’s more than those two things, but you’re still right in general that unsafe adds, it does not remove.)
ETA: I found your lwn comment. I forgot entirely about unsafe traits and I would say those are entirely different from my two categories. So, three things.
Can you expand on this? What are they? I would've thought rustc outputs only 1 IR?
* AST - abstract syntax tree, created from lexing/parsing
* HIR - high level IR
* THIR - typed high level IR
* MIR - mid-level IR (this is the one that enabled non-lexical lifetimes)
* LLVM IR - given to LLVM to generate the final binary.
If you want to learn more, the rustc dev guide has a bunch of details https://rustc-dev-guide.rust-lang.org/hir.html
> I would've thought rustc outputs only 1 IR?
Yeah, I mean these are in stages. One goes to the next goes to the next, and there's only one output, but that doesn't mean there aren't a bunch in the middle.
Different IRs are good at different things. MIR is based around the control-flow graph, which is also what non-lexical lifetimes are based on, so it's much easier to implement against MIR than it would have been against HIR.
1. AST 2. HIR 3. THIR (side-table lookups) 4. GCC Generic
We basically skip MIR in gccrs.
Its pretty sensible to have other IR's, we have many passes in gccrs simplifying things so the graph of what your working with is simpler and simpler each time.
I mean in GCC for C++ for example they use GENERIC and add a bunch of custom tree-codes such as LAMBDA_EXPR or TEMPLATE stuff for example then they keep substituting etc and finally as part of handing off to GCC middle-end it triggers the gimplification of all of these custom tree codes. So even the C++ front-end you could argue has two IR's.
Please correct me if I'm wrong; I'm not trying to pretend I am any kind of expert here.
> It's used for type checking and error verification; once that's done, it can be translated and handed to the GCC mid-layer.
Sounds the same serial thing to me.
Can you please share the reason behind why it should be this way ?
Many things in Rust are syntactic sugar and can be handled by desugaring the AST into another IR so you dont even have to think about it in other passes. The main issue for me is how complicate the type inference is.
So if i wanted to use GCC GENERIC for type resolution for this example:
``` let a; a = 123; let b:u32 = 1; a += b; ```
How do you resolve the type of 'a' you must use inference variables and then update the TREE_TYPE as you go so this means walking the tree's over and over again as type information is gathered over time on the inference variable. Using a separate IR and using id's and side tables makes all of this much much more simple for Rust.