Mrustc: a Rust compiler written in C++
github.com
github.com
One of the interesting things about this project is that because its initial goal is to validate the reproducibility of the main Rust compiler it doesn't actually need to implement a borrow checker, as the borrow checker strictly rejects programs and has no role in code generation.
Assuming the input is sound, mrustc will do the right thing. Ignoring the borrow checker will not make things magically work.
Which is more appropriate for this scene I think.
LOL
Do you really want that?
My point is that the borrow checker is explicitly meant to protect against data races and such. There are situations where I write code I can prove does not violate any hazards but Rust's "validation" is overly-zealous.
The few times I do have an issue I'm usually writing FFI code to interact with C, and I try to minimise that code as much as possible.
It's often a good exercise to craft your program to the language you're targetting, to take advantage of it's strengths or embrace it's traditions.
[1] https://stackoverflow.com/questions/4239121/code-compatibili...
[2] https://stackoverflow.com/questions/179492/f-changes-to-ocam...
Besides, all of this is necessary for soundness. Turning off the borrow checker won't magically get you a compiler that is less permissive, it will get you a compiler that will very likely produce nonsense if your program would have failed the borrow checker had it been present.
If you give an invalid program to the regular compiler it will be rejected, but if you give the same program to this compiler it might produce nonsense. So the developer of a program would feed it to the regular compiler for checking, but others could take the finished release of the program, assume that it's valid and feed it to this compiler instead, as I understand it.
-------------------
Jokes, aside, if you aren't there for Rust borrow-checker, then I wonder, do you really need Rust at all?
When one doesn't have a pervasive GC to dynamically clean things up, things still need to be freed but have to be freed in static locations (either explicit free calls in C or end of scopes in Rust and C++). Making this safe means defending against pointers becoming dangling which Rust does by restricting mutation. This restriction translates into some things that can't be copied, such as mutable references (that can lead to iterator invalidation and use after free, among every other undefined behaviour).
In other words, explicit/scope-based freeing being unsafe means the only way to safely manage memory is a garbage collector (tracing or reference counting), which limits how many programs can be written that are verified-safe by a computer. I believe Ada takes the approach you suggest, but this limits it to be most useful for programs that don't allocate (or only do O(1) allocation during startup).
Affineness also allows modeling things like session types better, letting programmers construct their own APIs that defend against mistakes.
You could, and people have, argue that some types could be "autoclone", as in, have `clone` calls automatically inserted where necessary, but when it is raised, a lot of people express dislike because they feel it will encourage slow code and mean people aren't guided towards the (usually) better solution of using references.
For the second one, consider Box<T> vs uniq_ptr<T>; Rust can statically prevent use after-move, C++ cannot, and you'll get a NPE. That is, you can totally get rid of a GC through only RAII, but you can't guarantee memory safety (as far as we know!) without affine/linear types.
Affine typing (plus the rest of Rust's system) on top of scoped memory management is needed to avoid problems like returning a pointers into something that is deallocated at the end of a function:
int &foo() {
std::vector<int> v = ...;
return v[0];
} // return value is dangling
And similarly, avoiding having pointers into things that are destroyed when their parent is modified: std::vector<std::unique_ptr<int>> v = ...;
int &ref = *v[0];
v.clear();
// ref is danglingI was under the impression that Rust's lifetime and ownership system is not just used to ensure safety, but also for knowing where the compiler should insert deallocation code for dynamically allocated objects at the end of their lifetimes. Is that not the case? Does mrustc generate code that leaks memory, or is there enough information in the Rust source code to do precise deallocation even without a borrow checker?
So basically lifetimes are used only by the borrow checker for intelligent linting. Once the borrow checker is done, the compiler loses interest in lifetimes entirely, and they don't affect codegen.
It makes porting things... interesting.
On a more serious note, ghc at least solved this problem by allowing you to compile Haskell to C (-fvia-C). Writing a naive C code generator from whatever your backend is using is probably not that hard.
Regardless, those compilers are usually grandfathered in; distro policies are different for other languages. It just depends.
GCC has migrated to C++ in 2008.
My language implementation compiles as C and C++. Every so often I build it as C++ (like before releases) to check for regressions and flush out any issues caught by the C++ compiler.
I don't think of it as "written in C++", though it is not technically a false statement.
All Scheme programs are really written in Common Lisp; they just need a suitable library of macros and functions ...
What's important though is that the GCC devs actively try to avoid the newest C++-features, this allows it to compile current GCC-trunk even with very old GCC-Releases. IIRC even GCC 4.3 (released in March 2008) is able to compile current trunk (GCC 8.0).
This is somewhat different for Rust, where I think the current policy is that master needs to compile with last released version (releases every 6 weeks and just to make it clear: Rust is written in Rust). Before that, the compiler was updated even more frequently. Bootstrapping Rust from 0 therefore is quite hard, since Rust is far from being as ubiquitous as C/C++-compilers. Even if you already have a rustc on your system, it's not unlikely that it is too old for compiling Rust-master. The first Rust-compiler was written in OCaml, so you need to compile that first, then compile Rust commit-for-commit until you reach current master. This could take quite some time. IMHO this is why mrustc is great, since you just compile current master (if it supports all features) with it and then use the generated compiler to compile Rust.
Sure, a pure function taking the universe as an argument.
Stuff like LLVM, and emscripten exist though so it's probably not as big of a deal as they say.
This is really something more languages should strive for.
Most languages these days obviate cross-compilation in the first place by being interpreted. Of the languages that remain, the natively-compiled ones, most haven't gone to Go's lengths of writing a custom libc, which is emphatically not recommended for both Windows and Mac (the syscall interfaces aren't stable, and this has broken Go code in the past: https://github.com/golang/go/issues/16272 ). And gc, the primary Go compiler and the one with out-of-the-box cross-compilation, supports relatively few platforms (I count 11, whereas rustc looks like it supports 50-70 platforms). You can use gccgo to get more platforms, but AFAICT gccgo's cross-compilation story isn't nearly as nice: https://github.com/golang/go/wiki/GccgoCrossCompilation .
For Rust, cross-compilation looks like this:
1. Install the libc for the target system, and, if necessary, a compatible linker.
2. Run `rustup target add foo` where "foo" is one of the target triples on https://forge.rust-lang.org/platform-support.html
3. Add the target triple to your Cargo.toml.
Which isn't too shabby. You can read more about it here: https://github.com/japaric/rust-cross
Can you name a distro or environment that doesn't already have a C compiler at arm's reach?
That's not the case with Rust compilers, which is obviously what the issue is.
Isn't there a bit of cognitive dissonance in believing that Rust as a language is an important idea (i.e. by the additional code safety and code maintainability that it conveys), but then simultaneously making the effort to rewrite the current Rust-implemented compiler in C++?
C++ is fast, but aside from a shared value around performance, it has fairly little in common with the ideas that Rust is built on.
As far as I can tell the goal of the project isn't to target more platforms (Rust targets quite a few by way of LLVM), so I don't think I'd choose any other language, including C++.
Having a compiler and standard library written in the language that it compiles has some huge benefits for increasing the pool of possible contributors.
There appears to be an attempt to leverage this project to target ESP8266, which AFAIK LLVM does not support: https://github.com/emosenkis/esp-rs
But this isn't the official compiler, this is someone's personal project?
True, but compilers are complicated machines and Rust is still changing at a fairly frantic rate.
The author seems to be doing quite a good job of development today, but if it has any hope of staying current, it probably needs to think about how to increase its bus factor (something happens like changing jobs, starting a family, or they just become interested in something else, and a single person suddenly has less time to contribute).
I don't think there's any real plans for a source-based approach. But epochs can only change a limited amount of things for exactly this reason; they minimize the compiler burden of supporting them.
It's clearly easier to make a compiler in a higher level language (Python is just an example, but Lisps are suited to this kind of thing). For example, text parsing is easier in Perl/Ruby/Python/Swift/etc. As someone who knows C++, more thought is required to do the same thing as in a higher level language, although it runs much faster. If you just wanted to bootstrap the compiler, then you'd choose the easiest route to that. It could also be easier to read and understand than a C++ compiler.
There is an open source version available in every system supported by gcc.
However C++17 is quite safe, provided one doesn't follow "write C with C++ compiler" idiom.
No, it's not. There is no use-after-free protection among other things. It doesn't matter how much you write code like C or not: it's just not safe.
Until then, those comments only annoy those of us that happen to like Rust, but don't find it mature enough to replace C++ on the use cases we happen to care about.
I don't think you'd find a single LLVM developer who would claim that LLVM is memory safe. Giving invalid IR to LLVM and not running the verifier frequently segfaults it, for example...
[0] https://www.dwheeler.com/trusting-trust/dissertation/html/wh... (previous HN discussion at: https://news.ycombinator.com/item?id=12666923 )
[1] http://www.ece.cmu.edu/~ganger/712.fall02/papers/p761-thomps...
Having a Rust compiler in C++ is a mitigation to the trusting trust attack, period. You don't need DDC for this.
DDC is necessary when you have two self hosted compilers (e.g. GCC and clang). Here we have one self-hosted compiler (rustc), and one in another language (C++). To mitigate trusting trust in rustc, use mrustc to compile rustc, and then use that rustc to compile itself, and now you have a trusted binary (provided you trust your C++ compiler. you can fix this by DDCing the C++ compilers)
The core guidelines are very similar to Rust, in spirit and intent.
That said, I always welcome tooling to make C++ safer; the end game is making programs better, not language partisanship!
[1] shameless plug: https://github.com/duneroadrunner/SaferCPlusPlus
This compiler in essence "bridges" the gap. Sorta.