mrustc: In-progress alternative Rust compiler (to C)
github.com
github.com
This is not a compiler for developing Rust programs. This is a compiler to let platforms unsupported by Rust/LLVM build rust programs from source. So it takes a Rust program which already exists, assumes it can be compiled by rustc, and translates it to C so it can be handed off to your local compiler. Skipping the burrow checker probably lets this execute in environments which are more constrained than what rustc can fit into.
Of course there's a question of how much of Rust's safety survive translation into C. Certainly some. E.g. the emitted program would include bounds checks where not specifically elided in Rust, and references are statically guaranteed to be non-null before being lowered into pointers. But there could be gaps, especially in cases where Rust semantics mismatch with C in ways that LLVM hasn't exposed yet.
Could you elaborate on what potential mismatches there currently are between the semantics of Rust and C? Where particularly do you think there might be issues/dragons hiding in the semantics of a translation?
Float integer conversions (overflows are defined to saturate in rust, are undefined behavior in C).
Maybe things like how pointer casts are treated in correct (but unsafe) rust vs C. Or similarly in transmute. Generally it wouldn't surprise me if rust was standardizing a memory model that was subtly different from C.
I'm sure the list goes on, but "that sort of thing".
Rust also has different aliasing rules. Stricter than C for references, but more relaxed than C for raw pointers.
Edit: Now that I am coming back and reading this though I am going to make an edit, and like normal for substantive edits I'm going to mark it clearly...
Edit2: Or I can't edit, so let's put it here:
A better example than any of the above might be infinite loops, `while (true) {}` (or any similar loop without side effects) is undefined behavior in C, it's defined behavior in rust.
I would assume that borrow checking uses far less memory than the actual compilation after the compiler has proven the program to be in concordance with the borrow rules.
It more so seems that it is something they did not bother to put too much effort into, as it wasn't necessary, than a legitimate way to reduce a compiler's memory footprint
As currently implemented in rustc, it operates on an intermediate representation that has already had name resolution and some simplification done. As such, i think it would be rather hard to separate from the rest of the compiler.
That alone would make rust a much easier language for newbies.
I wonder how these two projects compare.
There's a bunch of corporate/.gov envs that don't allow binaries to enter their system. All new code has to be compiled from source, on the target system (or at least on that side of the process firewall, it might be an airgapped cluster that can share code). They have C and C++ compilers that they've been compiling from source since the dawn of time, but they don't have a rust compiler. This gives them a mechanism to begin using rust within their process.
Being able to target a pretty recent version of Rust with a compiler written in C would be so, so useful for these purposes.
I would assume that on such compilation farms that these systems generally use, this could be done very quickly.
Perhaps there would be an interest for the Rustc team to provide a recursive automated setup that is capable of compiling the latest Rustc from OCaml, and finally from C as well since one must follow a similar process with OCaml.
Methinks that building this chain by trial and error is rather trivial compared to the actual work in building a compiler.
Typically, they are inversely correlated.
Since them they've bootstrapped this all the way to 1.50: https://git.savannah.gnu.org/cgit/guix.git/tree/gnu/packages...
The work done here is non trivial, this chain of versions are required for the final binary to be reproducible.
Indeed. Although it is being worked on, the whole bootstrap process of rustc is currently a major hassle, requiring to start from the oldest CaML versions of the compiler up to the most recent ones.
I hope there's some automated assistance for scanning inbound source code too. Imagine reviewing everything, line by line, millions of them.
I remember reading an interesting thought experiment that traced history that investigated what would happen if Dennis Ritchie had put malicious code in the first C compiler that was designed to detect whether the compiler compiled a compiler, and then copied the malicious code into it.
It concluded that tracing the history, that if this were to have happened, then GCC and Clang and many other programming languages would have said malicious code in their compilers that do not show up in the source code with none the wiser.
Before the B compiler written in B, there was a B compiler written in something else. I couldn't figure out what from that document.
Or would that not count as “compiling from source”?
I guess you didn't want to advertise the real reason of the project.
This takes very long and is error-prone, so having a compiler that can build even an intermediate version is already a big help.
You don't have a real language until there is a stable specification and multiple independent implementations. Until then it's just an experimental toy.
And people do see the pain here as a problem, that's why the project in this very thread exists. I think most in the Rust community see mrustc and the pressure that an independent implementation brings process wise as a very good thing for the language.
It should probably be still functional, or easily fixed if not.
That doesn't mean Python isn't useful! It is in fact a wildly successful tool that has spawned a large ecosystem of packages and a diverse group of users. But it unfortunately does not qualify as a true language, for it lacks the necessary ingredients: a formal specification of the language and more than one implementation of said specification.
This has a "I will cancel my subscription" letters to the editor feel, as if anyone cares :-)
I mean, there are pros and cons to any language and ecosystem. "I found some fault in bootstrapping approach and will condemn the whole language/effort for it" is a hardly a reasonable argument...
As if regular devs often have a need to bootstrap from scratch themselves...