GCC Rust: GCC Front-End for Rust
github.com
github.com
> A friend of mine (Luke) has been talking about the need for a Rust frontend for GCC to allow Rust to replace C in more places, such as system software. To allow some types of safety-critical software to be written in Rust, the GCC frontend would need to be an independent implementation of Rust, since the relevant specs require multiple independent implementations (so just attaching GCC as a new backend for rustc wouldn't work).
Luke:
> The goal is for the GNU Compiler Collection to have a peer level front end to the gfortran frontend, gcc frontend, g++ frontend and all other frontends.
> The goal is definitely not to have a compiler written in rust that compiles rust code [edit: unless there is an acceptable bootstrap process, and the final compiler produces GNU assembly output that compiles with GNU gas]
> The goal is definitely not to have a hard critical dependence on LLVM.
> The goal is to have code added, written in c, to the GNU Compiler Collection, which may be found here https://gcc.gnu.org 20 such that developers who are used to gcc may compile rust programs and link them against object files using binutils ld.
> The primary reason why I raised this topic is because of an experiment permitting rust modules to be added to the linux kernel: https://lwn.net/Articles/797828
> What that effectively means if that takes off is that the GNU Compiler Collection, which would be incapable of compiling that code, would be relegated to a second class citizen for the purposes of compiling the largest software project on the planet: the linux kernel.
> Thus it is absolutely critical that GCC be capable not just of having rust capability but of having up to date rust capability.
0: https://users.rust-lang.org/t/call-for-help-implementing-an-...
(You might prefer not to consider modern C++ more productive, performant, safe, and maintainable than C (and reflexively downvote), but the statement remains true: Gcc did transition to C++, for reasons. And, all the improvements that have kept Gcc competitive with Clang were done, since. And, Clang and LLVM are also coded in C++, also for reasons.)
For starters C++ does not have a stable ABI. Rust plans too have a stable ABI. In either case right now both fall back to the C ABI, but at least its a goal of the rust devs.
C++ is going to be harder to boot strap than either Rust core or C when trying to build a compiler for a new system/platform. C probably is the easiest bootstrap of the 3.
Rust and C are generally more performant than C++ code. Also in the realm of low end micro controllers C code tends to have the smallest binary image size.
In C++ if you need to interact with other languages a lot times your stuck to the C abi which which forces you to avoid a lot C++ features/constructs. As for rust its not an object orientated language and its unsafe blocks generally make it easier to at least keep rust code idiomatic.
C++ is also a massive kitchen sink of a language with many ways to hide gotchas. Probably the worst thing in C is its null terminated strings but C++ also suffers from that. C is simple, and rust has at least so far made good design choices for the most part.
There also choices in the C++ language that are even part of the standard library that seemed like a good idea but with hind sight are not that great of idea. For example operator overloading.
Although we are talking a GCC front end so I would be surprised to see both C and C++.
I'm not quite sure what systems development has to do with compiler writing? (I assume by system development, you mean something like writing OS kernels etc?)
Eg there are good reasons to stay away from automatic memory management when you are writing an OS, you want control over that. But those reasons hardly apply when you run a simple user space program like a compiler.
As for a compiler ideally its easy to port as its kinda of the bed rock to get any other system going. Although a compiler for some architecture or use cases is too large to be self hosted or does not make sense to be self hosted. So its not always a deal breaker for a compiler.
C++ is a language defined by ISO committee SC22/WG21 via the Standards 14882, most recently C++20 (superseding C++17).
C is defined by SC22/WG14, in the same way.
The most widely used implementations of these Standards -- Gcc, Clang, and MSVC -- are in fact the same project in each case (although MS's [until recently implemented] a long-superceded C Standard), and are, in fact, themselves C++ programs.
Portability of compilers is not needed, in general, to "get a system going", because cross-compilation is a mature and long-supported technique. Gcc is very frequently ported anyway, for practical reasons, via cross-compilation. Being coded in C++, and capable of compiling and cross-compiling itself, Gcc remains quite portable, as is repeatedly demonstrated by notedly frequent ports.
It is hard to imagine what could be considered a "deal-breaker" in this context. By the evidence, portability of C++ is never a dealbreaker as implementation language for a compiler, as in fact the compilers actually used in "get[ting] any other system going" are, with only rare exceptions, compilers in fact coded in C++.
(Exceptions are found not to be about portability, but about the memory footprint of the resulting compiler, which is not a product of its source language, but rather of the power of the compiler's optimizer, which is not always seen as necessary.)
In any case, the implementation language used for the compiler has absolutely no effect on the stability of an ABI for any target language. ABI stability is a choice made by language designers to favor backward link-compatibility over convenient expression of new features or bug fixes. It is hard for me to imagine the level of confusion that would produce this error.
Not anymore, as of 2020, MVSC supports C11 and C17, with the exception of the C99 features that were dropped in C11, like VLAs.
Previously they were only keeping up with the ISO C features required by ISO C++ compatibility requirements.
Yes. Though as far as suitability for 'systems programming' is concerned they might as well be Python programs or bash scripts, as long as they produce suitable output in an adequate amount of time.
OK, I can live with that definition.
But what does it have to do with the language you write your compiler in?
C has not been used for applications where performance is a priority for many years now. C lacks the expressiveness to make many software optimizations practical. I’ve written database engines in both C and modern C++ and it is no contest, C++ is much more concise while producing faster code. People primarily still use C for portability.
I'm not sure as to what capacity GCC has a similar instruction but I heard that LLVM's using of signed integers here is apparently nonstandard and what lead to Rust's decisions for vectors to be limited to a certain size: https://doc.rust-lang.org/nomicon/vec-alloc.html
I only bring it up because it pretty drastically harms readability (for me at least, and to a degree I frankly find surprising). Just thought you may want to know.
I’m going to assume that you mean that the italics don’t add value, and not their comments as a whole. I think your comment could be read either way. Anyways, even so others have since pointed out reasons why the might be used to doing so.
Some style guides make the even more inconsistent distinction that only titles of video games be italicized, but titles of “application software” not be.
I suppose that on H.N., where such software is frequently mentioned, it does stand out more.
I do not italicize Google the company either, but I do that of Google the search engine.
The titles of scientific research papers are also typically italicized.
And using signed integers is unusual yes, but again all backends (obviously) also support signed integers.
(I agree on the signed integer thing though.)
That being said, there's also cases where Rust does not have LLVM semantics, and that can cause bugs. Some famous examples being the loop optimization miscompilation, and &mut T currently not being marked noalias.
I don't think that's a case where Rust "doesn't have LLVM semantics" since the miscompilation was reproduced in standard C. Rather that rust is actively and ubiquitously leveraging otherwise rarely-exercised LLVM features, revealing a bunch of bugs (either leftovers or breakages) in them.
The former though is an area of significant divergence between Rust and C++ semantics, and LLVM directly following C++ semantics.
Ah I must have missed that, I thought you were talking about only one thing (since IIRC the noalias miscompilation is due to loop unrolling?)
TL;DR: C++ says that an infinite loop with no side effects is UB, Rust does not. Empty loops in Rust will disappear entirely when they should really loop forever.
(I edited my comment slightly because, re-reading, "blindly" sounds too negative.)
loop {
core::sync::atomic::compiler_fence( core::sync::atomic::Ordering::SeqCst);
}
There is a major effort to revamp the `noalias`/`restrict` handling going on for a while now. It takes quite long because it is hard and complex and we want to get it right.
In case you are interested, here is the new design https://reviews.llvm.org/differential/changeset/?ref=2170825 here the overall code changes currently considered https://reviews.llvm.org/D69542 and here you can find information on our monthly LLVM Alias Analysis call https://docs.google.com/document/d/1ybwEKDVtIbhIhK50qYtwKsL5...
So there is a significant difference in semantics between Rust and C/C++ here.
The standard may allow the generation of insane code, but a decent-quality compiler does not do so. The standard ought to be fixed.
Letting the compiler assume that a "switch" without a default will never go there is a great optimization too. Why not put that in the standard?
Actually, this is worse. Letting the loop be UB is like letting a true "if" be UB. Just never mind the code actually written; surely it doesn't mean what it says.
that's actually the case already. A missing default is UB if the switch condition does not match any of the cases. And C compilers already optimize accordingly.
Which ones?
If this optimizations are so important, how come Rust was designed in such a way to make them impossible? Also, how does this fit, e.g., the benchmark game results which show that Rust is faster than C for all benchmarks considered there ?
C does not have this guarantee (at least not in all cases). Also rust is compiled with the llvm backend, so my understanding is that in practice rust assumes that loops terminate. See:
https://github.com/rust-lang/rust/issues/28728
There are llvm directives that can be added to prevent the optimization, but they are rejected by the rust maintainers exactly because they would cause performance regressions.
If your loop infinitely writes the number 70 to an atomic int, it's fine.
A stable ABI is unclear.
I was already sharing templates and classes across DLLs in Windows 3.x compilers, and keep doing it with VC++ to this day.
On Linux with Qt and Gtkmm projects, and on macOS/iOS with their system frameworks.
Which is the biggest reason I cannot put up with cargo's model to compile every single project from scratch, after git clone.
As we have lengthy discussed, while it might not be a priority right now, it is certainly an adoption block among some Ada, Delphi, Swift, Objective-C, C and C++ communities.
Likewise Rust users should write to the C ABI for stabled shared libraries. (I believe there are a bunch of mechanisms to do that, though I'm not really a Rust user)
For example look at what KDE does: https://community.kde.org/Policies/Binary_Compatibility_Issu...
Yes, you have to do a lot of extra work. That's working as intended.
It's sort of an oxymoron to expect to use every C++ or Rust feature in your user-facing ABI and have it be stable. They are incredibly rich languages, with drastically different notions of "function" than C has (let alone other constructs like data layout)
That's the whole point of the KDE doc I linked. It's not a reasonable strategy to use unrestricted C++, just like it's not a reasonable strategy to use unrestricted Rust. You actually have to design an ABI, not just rely on the compiler to do it for you.
I thought there was a GNOME doc too, but I couldn't find it. The point remains: there are lots of things in C++ that people who care about stability don't use at their ABI boundaries.
Specifically KDE relies (among other things) on the ABI stability of the layout of virtual tables which is certainly not part of the C ABI.
What I expect to happen is that rather than "Rust gets a stable ABI" it will be "Rust very gradually builds upon the C ABI for selected features".
Most C++ libraries have templates in their signatures these days for efficiency, and Rust has a similar flavor (monomorphization). I think there is a fundamental tradeoff there between stability and performance, and both languages have a heavy emphasis on the latter. I could be missing something as I certainly don't know all the details of the C++ ABI.
https://gcc.gnu.org/onlinedocs/libstdc++/manual/using_dual_a...
So the ABI can leak implementation details in nontrivial ways that library authors usually don't consider. It's better to have something explicit in the code, e.g. under extern "C".
The point is not that making a stable API for every language feature is impossible; just that it's hard, fragile, and maybe not be worth the effort. If you really want stability, then use fewer features more like C. There is probably some middle ground that's richer, but templates are known to cause problems.
Also see the release history here:
The layout of std::string changed to conform to the new standard; exactly the same thing would have happened if std::string was a C POD type, there is no way around that.
In fact libstdc++ to this day still has the option to be compiled with the old ABI.
There have been other minor ABI breaks, mostly to fix bugs, which affect only a small number of programs.
Pure C libraries have exactly the same ABI stability issues caused by changes in layout of structures; to avoid this C libraries expose only opaque pointer sized handles to heap allocated objects. This is similar to C++ libraries only exposing pointers to virtual interfaces, but for high performance libraries, for example containers, this is not considered acceptable.
You can write conformant STL implementations that have completely different layouts. But the GNU one no longer had a valid layout, so it had to change.
So when you expose libstdc++ as a shared library, you're exposing a bunch of details that aren't part of the C++ standard.
If you write it in a header, then you can expect it to break callers via shared library. And templates must be in headers. This is a fundamental language issue.
I don't use Rust, but the point is "making a stable ABI" will expose it to these sorts of problems, i.e. implementation details leaking on specific architectures, outside of the language.
The projects that care about ABI stability, e.g. sqlite, don't expose layouts. They have only functions and not data in their headers.
Your original point was that for binary compatibility you write against the C ABI. I'm claiming that the C ABI has exactly the same fragile layout limitations and you can use the exact same workarounds in C++ (i.e. only expose pointers or make sure that your layout doesn't change).
I'm also claiming that templates are a red herring and have very little to do with ABI.
What people usually mean here is that Rust would have its own stable ABI, that you would get “for free,” without needing to do that work. (It’s never actually free of course... but that’s yet another detail that is usually papered over when people talk about this.)
This (among other things) was debated a lot.
A major con of not sharing is not having a backend at all. It's a lot of work to keep up with a moving target
There is also a port of the Ada frontend to LLVM backend:
As it stands, the way to bootstrap the official Rust compiler from source with just a C/C++ compiler is a few options:
* Compile OCaml (implemented in C), use it to build the original Rust front-end in OCaml and then build each successive version of the language until you hit 1.49. This option is not fun.
* Compile mrustc, which is a C++ implemented compiler that supports Rust 1.29. Use that to build the actual Rust 1.29 and then iterate building your way all the way to 1.49. That is less bad, but still not fun.
* Compile the 1.49 compiler to WASM and run it via a C/C++ implemented runtime in order to compile itself on the target system. This would also mean packaging and distributing the WASM generated code, which some distributions would refuse. I also am not sure if it's even currently feasible, as I don't follow the WASM situation closely.
A compliant, independent C++ implementation that could be built in the ten minutes it takes to build GCC itself would be a very good thing to have and would be more friendly to distribution maintainers.
How about memory safety and fearless concurrency?
Keeping a viable C++ implementation as part of GCC would be the smartest decision.
I know OpenBSD avoids rust because of the bootstrapping issue, but they also avoid LLVM because of a licensing issue.
https://bootstrappable.org/ https://bootstrapping.miraheze.org/wiki/Main_Page
Have you ever seen GCC crash with a SIGSEGV? I rarely did even when I used to be a GCC developer.
Regardless, it was worth mentioning as a potential option. I am one of the handful of maintainers for an experimental distribution where packages are either compiled or interpreted from tarballs and this would be something we'd consider. I'd MUCH rather have the GCC front-end option, however. So far, we've simply not packaged Rust and have accepted that as dead-ending our Firefox package. This may potentially revive it.
It might. There are at least two major things off the top of my head, regarding libstd:
1. specialization is needed for performance around String and &str
2. const generics are needed to support some trait implementations
We currently allow some stuff like this to leak through, in a sense, when we're sure that we're actually going to be making things stable someday. An alternative compiler could patch out 1, and accept slower code, but 2 would require way more effort.
There has been some discussion about trying to remove unstable features from the compiler itself, specifically to make it easier to contribute to, but it unlikely that it will be completely removed from the current implementation of libstd for some time.
> It's managed to build rustc from a source tarball, and use that rustc as stage0 for a full bootstrap pass. Even better, from my two full attempts, the resultant stage3 files have been binary identical to the same source archive built with the downloaded stage0.
GCC is also roughly even in compilation speed now https://www.phoronix.com/scan.php?page=news_item&px=GCC-Fast...
LLVM is much easier to work with internally but GCC is a seriously good compiler even now.
Maybe getting some new GCC devs in there with projects like this would help with that?
Running `make -j4` seems safe thus far.
With all this said, I would love for them to succeed, for multiple reasons. Including <3 GPL.
in fact, IBM uses LLVM proper as its own compiler backend for its ppc processors
This may be a vendor-specific extension though?
> The developers of the project are keen “Rustaceans” with a desire to give back to the Rust community and to learn what GCC is capable of when it comes to a modern language.
So what's the answer right now ? How does GCC measures "against" rust ?
Ada, D, Go, Modula-3, Modula-2, C++20
The only way to do this properly, if desirable, is to make GCC an official backend of the main frontend. That will defocus some progress that happens with LLVM (every feature has to be implemented on both backends) and can make dev lives hard (eg “oh this problem comes up with GCC so use the LLVM backend “). The value would be if the majority of bugs/features are in the shared frontend.
This project though seems like a parallel implementation of Rust. That’s valuable for the community and inevitable as a part of successful growth. I don’t believe it’s beneficial to the community though if this grows beyond a toy, niche project.
No. The correct way is to create a Rust language specification that describes what the correct behavior is.
Then whether LLVM, GCC, or something else is used does not matter. There won't be one implementation with defacto behavior, there will be multiple implementations that follow the spec.
This is the way mature languages work.
Just as there is a C++ standard and multiple C++ compilers that implement the standard, there should be a Rust standard and multiple implementations.
This is the way that mature languages work.
If this is the benchmark, then among popular languages you basically have C, C++, C#, JavaScript, and... is that it?
That being said, I do think, considering it longer, there are languages that I am missing, like SQL, and ones that have a spec, even if it’s not under an ECMA/ISO process, like Java and Go, that I was forgetting.
I still think that using this as a necessary condition for “maturity” is misguided.
Specs are definitely a good thing to have, but often they're just brandished as a bullet point without looking at the details: how good is the spec, what does it add over the existing tests/CI/RFCs/proofs, etc.
There is a COBOL specification, but AFAIK nobody actually implements it fully. Implementations pick and choose new features based on customer demand.
Also, COBOL is dominated by large legacy codebases, which means that if there's a discrepancy between an implementation and the specification, the users normally don't want it fixed, because they may have written code that depends on the "incorrect" behaviour and it's a lot of work to audit.
IBM built a new backend for its COBOL compiler using the optimizer from their JVM. IIRC by the time I left it generated code that was ~2x faster than the old compiler, but uptake was still slow because of migration concerns. In particular, we spent a lot of time working on features to help guarantee that code with some forms of undefined behaviour would have the same result as the old version.
Having a webapp depend on a CPython implementation detail is very different from having a kernel depend on an implementation detail of a language without a spec that was used to implement it.
And languages actually stack on top of each other. Imagine depending on a CPython implementation detail that depends on an implementation detail of a specific C compiler. Those things do happen and they make programmers' lives miserable sometimes, but imagine how often that would happen if C didn't have a spec that all compilers strive to implement.
Ada, Fortran, Cobol as well.
This is specially relevant in the industrial sector with certified compilers.
Few languages find such specification before that time.
Python also lacks it despite some competing implementations, since none actually diverge enough from CPython.
Standardization is helpful for tool building and experiments (ie here are the invariants we’ll never change). Languages don’t work that way and seeing how C/C++ have evolved (or really failed to do so at a meaningful pace), I’m under the impression (clearly unpopular due to the downvotes) that standardization and multiple competing tool chains are the cause of a lot of unnecessary complexity (not just within the language but also users of said language).
If "progress" means "rapidly add more and more new features in each release", multiple implementations will slow things down. But that problem can be addressed with editions: at some point, if the project is a success, the gcc front end will be a feature-complete version of some Rust edition, plus enough extra features to build an older version of the Rust compiler. At that point, you have a better solution to the bootstrapping problem (how to get a Rust compiler when you only have a C compiler and you want to build everything from source and not trust some binary you download from somewhere).
yet rust has a single implementation and by some metrics does better than both.
This has to happen at some point anyway otherwise we'll just get another C++. And I don't think Rust would benefit from that. New languages are designed to fix problems with the old ones, not to replicate them after all.
I just hope the designers will choose that point wisely.
Besides, if Rust doesn't become another C++, it won't fulfil the industry needs that C++ caters for, thus while it might become a success in some domain, it won't replace C++ in the OS and GPGPU SDKs.
Rust might move in alongside C++, in some, in time. Or, Rust could still very possibly fizzle. That would be the normal course of events for a new language, barring a miracle as was dispensed to Javascript, Java, C++, and vanishingly few others.
Will Dart survive and thrive? Kotlin? Scala? Clojure? Go? All doubtful, based on prior experience. Having a lot of code and a lot of users does not seem to suffice. Many other languages had those, and faded. Ada even had $billions in backing, and faded.
What we can say confidently about Rust's future is that it is not certain to fade. The miracle has come in less deserving cases.
Yet, NVidia picked it up over Rust, go figure.
I don't see why this is a bad thing. I'm a C++ developer, and I like C++. (I like rust as well)
As long as "quality implementation" means "implements the entire Rust language and not a subset".
Because otherwise, people will start getting requests to avoid using features that the non-standard Rust toolchain doesn't support.
> eg “oh this problem comes up with GCC so use the LLVM backend “
The point of having multiple implementations is making the language independent of the underlying system. A programming language is an abstraction. The only way to test whether an abstraction is a good one is trying out how well it abstracts away various underlying systems. That's why Rust needs ports to various architectures, OSes and compiler backends. Having GCC as a backend counts double here, because you get a few new target architectures for free with the port not just a new backend.
It's of course much different with Python but I can still clearly see how having an implementation defining a standard hurts the language.
One exception to this is Visual Studio's toolchain but let's not talk about Visual Studio's toolchain on a weekend...