GCC Rust Approved by GCC Steering Committee
gcc.gnu.org
gcc.gnu.org
And it's free software, for those that care, this matters.
it seems like this term is misunderstood in English to mean "no money" .. perhaps "Libre Software" is a better starting description here
I can never remember which way around this is. I think the "beer" is supposed to be free in the monetary sense and "speech" is supposed to be free in the rights/liberty sense, but the analogy doesn't actually convey this:
- Speech is normally both monetarily free and (in certain places) a right
- Beer is normally neither monetarily free or a right. But may make you "feel" free, which is another sense entirely.
Is this a cultural reference I'm missing?
Now someone tells you, “There’s free speech at Oktoberfest.” Is your first thought that there’s no monetary cost to expression, or that the expression is unencumbered by particular legal restrictions?
(Happily, Oktoberfest has both.)
Think of them as the _slightly_ longer "Free as in 'free beer' " and "Free as in 'free speech' ". When beer is described as “free,” it’s understood to be gratis. When speech is described as free, it’s understood to be libre. These are used as examples for more ambiguous situations, since software can be one, the other, both, or neither.
Former is gratis, latter is libre.
TBH the main place I've seen the phrase "free beer" is a Simpson's joke: people stand in front of a sign saying "free beer", but after Homer drinks many cups, the sign is revealed in full to say "alcohol-free beer, $5/cup".
It's understood that beer normally costs money, so the setup of the joke is that "free beer" is understood without explanation as "beer without costing money".
"free as in speech, not beer":
"free speech", the right to say what you want without legal restrictions from the state
"free beer", beer you don't need to pay for
Nothing to do with beer making you "feel free" (?) or whatever
Then again, free speech is not about being able to modify speech, release your own version of another's speech, and so on (e.g. the US has "free speech" but MLK's dream speech is still copyrighted).
Perhaps a better slogan would be:
"Free as in free sex, not beer"
And if the slogan wasn't meant for corporations too, and wasn't coined in the puritan 80s/90s, it might have been...
https://www.mentalfloss.com/article/28036/its-steal-how-colu...
They both fail to convey what it's really about.
And "free as in speech, not beer" actually does little to clarify it IMHO; it only works well if you're already familiar with the concept, but if you're not it only adds to the confusion (free speech is about the freedom to say what you want, so free software is about being able to write whatever software you want?)
Now, "free software as in right to repair" would actually clarify it, but trying to get an understandable message out seems to be "unethical" so meh :-/
Free software just isn’t enough anymore because the licenses limit your freedom too. (and even my choice of the word “limit” here will be controversial between the Apache/MIT and the *GPL crowds).
Don’t have an easy answer :-)
The problem is that merely right to repair doesn't cover all of Free Software so it will never catch on with the hard-core FSF crowd who seem to think that the only possible step is one giant leap to a Free Software utopia and than any smaller step in that direction is "unethical".
So I'm not really sure it has any more meaning than free software means no cost rather than "no one and everyone owns the software".
Price, functionality, availability, etc. all factor in. And the tricky thing with repairs is that it's a very "non-obvious" feature.
“Ability” is the wrong word here, why did you use that word? It’s ambiguous and misleading, that’s why it’s called the right to repair.
Yes most people do assume they have the right to modify and fix something they “own”, whether it’s a car or a computer or anything else. Don’t you think it’s surprising and non-obvious that you might not be allowed to repair something you paid for and own, even when you know how to repair it? Whether they ever intend to repair something, and whether they have the knowledge to repair something, are both irrelevant to whether they have the legal right to repair something.
> they keep buying products where it’s expressly not possible to repair it yourself or by a competent 3rd party.
Most people have no intention of repairing their technical purchases themselves, and that’s perfectly fine. Most people also don’t care about right to repair laws even if it affects them. The point of right to repair laws is to establish the common sense standard that consumers are allowed to legally modify their purchases, and companies will no longer be allowed to go out of their way to prevent repairs or make them difficult (also and perhaps especially wrt 3rd parties), and also that companies and lawmakers will no longer be able to make customizing and repairing purchased products illegal as has been done in the past.
Has been in the kind of news Joe Average doesn't give a fuck about and even if they watch, they forget in 10 minutes.
Techies spot mentions of it in the news because they already care for it and know the term.
That's not the same as something being in the news that reach regular people (that would be more like some high profile Hollywood divorse, some war, or gas prices).
No such thing. I'm fine with Joe Average, and I don't think everybody should know about everything, much less about the "right to repair" (much more important stuff for people to learn about, including politics and/or gas prices and other things that affect someone more).
I'm merely stating a fact: the "right to repair" is not something that has been covered in any degree for the regular persons to know or care about. Most of those that do are already familiar with the tech scene, if not as technies, then as gadget lovers and tinkerers.
Well, people understand it just fine.
It's just that the definitive allure of FOSS is not what the purists think it is.
purist: FOSS is about freedom (to work on the source, release changs as FOSS, etc)
real world: what I care about is 70% the code being available to use for free and 30% the ability for others (and me in some cases) to edit it and release my changes. Yeah, I do understand it's not the same as mere "freeware". But for most intents and purposes that the part I care most about, but I also like the FOSS aspect because it means I can get more community stuff for it, and I don't depend on a single vendor releasing it.
But we are on HN right now...
For you and me, maybe. But no, Free Software is often misunderstood. I had to explain the difference on a group dedicated to Linux the other day.
It is worth pointing out, every time.
To understand the concept, you should think of “free” as in “free speech,” not as in “free beer.” We sometimes call it “libre software,” borrowing the French or Spanish word for “free” as in freedom, to show we do not mean the software is gratis.
GCC is an engineering marvel, and the basis of two generations of engineering by hundreds of people and companies. Rust joining one of the many languages supported by GCC and stack, is welcome here.
Every build system that has added support for Rust, which aren't many, had to be radically modified to achieve that.
None of these supports the GCC Rust frontend, but all of them support the Rust frontend.
So if you actually wanted to build any >100 LOC Rust project for embedded targets not supported by LLVM, doing it with the Rust frontend is as easy as just running 1 CLI command to pick its GCC backend.
Doing it with the GCC frontend, would require you to either port one of the build systems to support it, or... give the GCC frontend a CLI API that's 100% compatible with the Rust frontend.
https://github.com/Rust-GCC/cargo-gccrs
Moreover, modules are less interesting to me in embedded development, for which I'm interested in access to Rust's borrow checker for gaining certainty of small portions of larger projects, which are written in other languages.
That said, gcc supports more platforms than llvm; including esoteric and unpopular desktop configurations.
Personally, I plan to drop this into marsdev as soon as it releases. Writing 32X games in rust sounds like silly fun.
AVR support is almost there in mainline rustc but has a codegen bug or two, last I heard.
I bet Cadence was ripping them off for the IP, which is a shame…
Any talks in Oxide about porting Hubris to RISC V? I hear getting your hands on Cortex-M*s in bulk is still pretty challenging these days.
Hubris was designed to be easy to port to RISC-V, so yes! We didn’t end up doing that though. Someone else did though! https://github.com/oxidecomputer/hubris/discussions/365
On one hand, having frontend diversity for a language helps bring new people into the language; think of the groups who can't use Rust because it doesn't support certain targets or can't integrate with their existing toolchain. There is a lot of talent that could be brought into Rust just by virtue of them being able to be included and the less barriers of entry, the better.
But on the other hand, I don't want to have to think of implementation specific behaviors, bugs, quirks, et cetera. Having an ecosystem built with thousands of people wrestling with that problem seems like a recipe for buggy software and burnout. People say that a spec would help with this aspect but I'm sceptical; a spec doesn't prevent divergence from happening, it just gives you a frame of reference for what is correct. You also don't need multiple implementations to have a spec either!
I fear that the way this is going to pan out is that we'll have multiple frontends for Rust, but the vast majority of engineers using the language will only ever think about "true" Rust: the current mainstream implementation. Issues in crates that primarily affect "alternative" Rusts will go unacknowledged, be tossed out, or have hacky patch jobs. We'll end up with either a) two separate, but high quality, ecosystems or b) one shared, lower quality ecosystem that is constantly fighting itself. I hold Rust in really high regard and it's my favorite language at the moment, so this thought is scary to me.
Though, there are a lot of smart people involved so maybe it doesn't have to be so doom and gloom. A couple of years ago I thought that cross-platform software was really messy and hard to get right but languages like Rust have the proper language features to make things like this much easier (though not perfect!). It could be that the emergence of multiple frontends could necessitate features and tooling for the Rust itself which makes all of that I said a non-issue. Maybe. Hopefully.
My concern with fragmentation is that this is essentially a reimplementation of Rust in another language (C++) targeting a different backend (GCC). It's only natural for there to be differences between the two, especially over a long period of time.
There is a separate initiative that may be more in line with what you're thinking which adds GCC backend support for the existing mainstream Rust compiler.
[1]: https://en.m.wikipedia.org/wiki/Technology_Compatibility_Kit
Doesn't this cause fragmentation in the rust ecosystem?
P.S.:I understand that people can work on any project they want. And I don't have the right to tell them not to. I'm just curious about the technical reasons for having multiple compilers.
I am out of the loop though, so if this is true, that's interesting and a bit weird.
Most contributors eventually moved into OpenJDK after it became available.
GCC folks left it around for a couple of years, because GCJ unit tests exercised parts of the compiler no one else did.
Eventually they decided it wasn't worth that maintenance cost to keep it around only for that purpose.
More likely that GCC has to follow all the bugs and quirks of rustc or no people will use GCC for Rust.
Did you mean target platforms? If so, how is this not already addressed by the rustc gcc (Via libgccjit) backend?
For stability compilers already have a large incentive not to break old programs. For longevity I don't really see how a standard affects it that much. For being more portable you do not need an entirely knew compiler.
It also allows people to design new backends (looking at CUDA LLVM backends)by finding out the right abstraction to support performance. For example, implementing a C or C++ compatible CUDA backend required the C++ committee to make changes to the memory model / consistency guarantees of C++ atomics. If C or C++ had only depended on compiler implementation for it, then there would have just been different implementations with different guarantees with no consistencies between them, and no single way to even define why they were different.
There are a couple of CppCon talks on the subject.
So you can't really produce a high quality standard with only one implementation. You'll miss important details.
Reliability at the extremes (where it may be even life or death situation) requires the developer knowing what the program (s)he is writing exactly expresses in Rust.
That just need documentation. You don't need a standard for that.
1. the standard covers everything you wished it would cover
2. every implementation implements the standard, with no bugs
3. the standard doesn't itself contain incoherent or contradictory things
Standards are a tool, not magic interoperability sauce.
The standards implying binary compatibility rules are about ABI (application binary interface), which usually depend on the ISA (instruction-set architecture) or the OS (if any) being used. You cannot have the unique one once there are multiple ISAs/OSes supported. Even when you only want to rely on some external "exchanging" representations not tied to specific ISAs, there are already plenty of candidates: CLI, JVM, WebAssembly... Plus there are more than one executable (and mostly, runtime loadable) image formats widely used (PE/COFF, ELF, Mach-O ...). You will not have the unique combination, and any attempts to ensure it "work across compilers" in that way will likely finally just add a new instance not fully compatible to existing ones, making it more fragile.
That sounds like a benefit to me :)
However, I think that's due to the general lack of interest in Lisp. You can see the C++ community has a similar ANSI standard and updates it every few years.
> Massive fragmentation in the compiler ecosystem
I wouldn't call it massive. They are pretty consistent, up until things like POSIX and FFI APIs. Let's agree there is some fragmentation. Isn't this still a better situation than if nothing was guaranteed?
Scheme is suffering from the same issues. Scheme is standardized, but implementations end up being incompatible with each other in subtle ways and the level of fragmentation is very painful. Scheme does get some updates unlike CL I guess, but all of the implementations either don't implement the the modern standard, don't have useful extensions for real world programming or are simply immature and don't have enough people working on them to get them into a nice state. In practice it's very difficult to use Scheme for anything non-trivial because of these issues.
I would much rather have no standard at all, and a single high quality implementation that everyone targeted instead of the current mess for both CL and Scheme. Until we get a new dialect that solves these issues Lisp is going to be more or less dead and irrelevant.
That would be Chez Scheme [0], maintained actively by Cisco, a company that you may have heard of - who also use the language extensively.
Racket is porting to using Chez, because it is the industry standard, it's performant, and rock solid.
The GNU alternative to Chez is Guile. Emacs can run with Guile, and Guix is built on it. It's got a fairly large community.
Outside of Chez and Guile, there are implementations and communities, but comparatively, they're tiny. Those two are the only big names you need. Like GCC and Clang for C. There are other C compilers. But you only need to know those two.
I think it already happened in 8.0
https://gcc.gnu.org/wiki/History
Thus, a new strain of development can attract people who want to do things differently, and reduce tensions all around.
On the legal front, it's unlikely, but sometimes there's legal problems with continuing to use a certain codebase.
EDIT: Clang's version is the attribute 'musttail,' if anyone is interested.
It may have a very detailed accompanying technical documentation of what the implementation is supposed to do, this may be called a specification, but a specification deserves the name with a minimum of two implementations.
This post is about a new front-end.
Rust will need a standard.
The main reason why I don't take it seriously is that code written 5 years ago will often not compile today. For a language that pretends to be a systems language that is a non-starter. If you can't guarantee a 40 year shelf life of your code then no one working on systems cares.
People working on systems in the wild don't have the brain power to learn a new tool chain every decade, let alone every year. They are solving real problems and not writing blog posts.
Or as Wikipedia puts it: "C17 addresses defects in C11 without introducing new language features."
citation needed. Yes, there's a few programs that relied on unsound things for which this is true, but that's a relatively small part of the overall amount of code.
Yes and?
Systems programming isn't front end JS work where breaking things doesn't matter. It's no surprise that Rust came out of the browser space. Only people who don't take their work seriously could ever think the above is a justification and not a red flag for never using it.
I welcome Rust becoming ossified in GCC so I can build 30 year old code without modification like I can in C. Until then, it's a toy for people with more time than responsibility.
For example if I have a trait Foo with a function bar, and I impl Foo for HashMap, and then a new version of std comes out that has named something HashMap::bar, now every call to my_map.bar() is ambiguous
And yes, there are tons of things that can subtly break code. That’s why the rust project runs the entire open source ecosystems’ tests as part of the testing process for the compiler. It’s not all of the code in existence, but it’s pretty good at flushing out if something is going to cause disruption or not.
In practice, the experience that the vast majority of users report to us is that they do not experience breakage when upgrading the compiler.
That is even worse. It means your code can silently start doing the wrong thing rather than erroring, if the inherent impl does something different from the trait.
> And yes, there are tons of things that can subtly break code.
This isn't some obscure bug in some deep edge case though, it's a completely normal and common way of using the language (implementing your own traits on foreign types) predictably leading to breakage in an obvious way. I am not sure why it should be called "subtle".
Anyway, given this issue, I think the meme that Rust is backwards-compatible is really oversold. It'd be more honest to frame it as "we hope releases are backwards-compatible, but we like adding new functions to the stdlib, so no promises" rather than marketing BC as a major selling point as is done now.
> In practice, the experience that the vast majority of users report to us is that they do not experience breakage when upgrading the compiler.
It's anecdotal for sure, but I'm personally aware of times when the exact situation I'm describing has happened and caused headaches for people.
I'll give you one, or two, depending on what you'd like to count. At some point in the conceivable future we'll be able to compile some meaningful Rust code base with both compilers and measure; a.) how long it takes to compile and b.) the performance of the compiled code.
Obviously that will induce what it always has: incentive to improve.
A big part of getting a sane environment to build rust Linux modules is having it easily integrate with the gcc toolchain.
There seems to be some contention around if this is a good idea for rust as a language itself and my counter to that would be that there are billions of people on the planet and the work integrating rust into gcc doesn't detract in any way I can see from the continued development of rust as a language, it just helps to make it more viable for many use cases and therefore increase it's footprint and overall support.
It might turn out to be a bad idea. But at this point, gcc needs to try something to stay relevant. This is the first piece of news about gcc that made me go "whoa" in the last couple years.
You're right that GCC Rust will probably be behind rustc for a while, but the sooner you start, the sooner it'll get there. The language isn't gonna "settle down" in that way any time soon, might as well just get going as soon as possible.
Actually, I'd say the language HAS settled down. Most new releases are stabilizing APIs or general tooling improvements and not actual language changes.
Not to say there aren't outstanding tweaks, adjustments, or improvements to the language that are ongoing, but rather, the target isn't moving nearly as fast as it was when 1.0 or rust 2018 were released.
> For some context, my current project plan brings us to November 2022 > where we (unexpected events permitting) should be able to support > valid Rust code targeting Rustc version ~1.40 and reuse libcore,
(current stable rustc is 1.62, 1.40 is from Dec. 19, 2019)
Which is in line with what you're saying about gcc likely being behind initially, of course.
I bet it's easier to go from version 1.40 to 1.6x than from version nil to 1.40, though - you have to start somewhere, and the later you start, the longer till you have a viable alternative.
Secondly, Rust has strong backwards compatibility guarantees and you can pin your code to a certain edition.
Thirdly, one big use case for gcc would be compiling the Linux kernel that might contain Rust in the future. So it would be enough to support the subset of Rust that gets actually used in the kernel. I would imagine they would be very conservative about which features get used and adopt at a much slower speed.
For now, it's the opposite: the kernel needs several Rust features which aren't even stable yet (https://github.com/Rust-for-Linux/linux/issues/2). But I agree that, after Rust has been in use in the kernel for a while (and it no longer needs any unstable features), the kernel developers are going to be somewhat conservative about which features are required (but not about which features are used; they probably will use a lot of conditional compilation, like they already do for gcc/clang, see the compiler-*.h files).
[0] https://gcc.gnu.org/backends.html
[1] https://doc.rust-lang.org/rustc/target-tier-policy.html
[2] https://doc.rust-lang.org/nightly/rustc/platform-support.htm...
[3] https://web.archive.org/web/20220223133124/https://people.gn...
[4] https://lwn.net/Articles/771355/
[5] https://news.ycombinator.com/item?id=26097153
[6] https://news.ycombinator.com/item?id=26203853
Does that means than (after merged) I can use gcc to compile rust code right? What are the advantages of this versus using the normal rustc compiler?
All the crate/project management would still be done with cargo? It seems to be a tool very entwined with rust itself, so I wonder if someone uses rust without cargo?
The main benefit of a GCC front-end is to move Rust beyond a single vendor, and shakeout ambiguities in the (informal) language specification.
https://users.rust-lang.org/t/where-is-the-rust-language-spe...
"For the most part though, rustc itself is the spec" - so, for implementing the GCC frontend, read the current LLVM frontend really carefully and do everything the same way? And try to keep up with all the changes which will inevitably happen in the LLVM frontend while you are implementing the GCC frontend?
At a high level, this is correct: https://github.com/Rust-GCC/gccrs/wiki/Frequently-Asked-Ques...
> If gccrs interprets a program differently from rustc, this is considered a bug.
I don't believe they are only reading the frontend, but using the reference first, then looking to the implementation second, and asking a lot of questions along the way.
(btw, there's also https://doc.rust-lang.org/reference/)
With an ideal standard document that describes the language, someone could build a compiler for the language from scratch and it should behave correctly. See https://en.wikipedia.org/wiki/Programming_language_specifica...
Some example standards:
C: https://www.open-std.org/jtc1/sc22/wg14/
ECMAScript (Javascript): https://tc39.es/ecma262/
C++: https://isocpp.org/std/the-standard
Scheme: https://schemers.org/Documents/Standards/
Ada: http://www.ada-auth.org/standards/ada12_w_tc1.html
For Rust, Ferrous Systems is working on the Ferrocene Language Specification to formally document the Rust subset that Ferrocene will use.
https://ferrous-systems.com/blog/ferrocene-language-specific...
ISO will stamp any fantasy like OOXML, but that doesn't mean the standard is useful.
Long-term, it's considered a strong signal for the health and viability of Rust if it's not strictly tied to one implementation.
For example, for arm-none-eabi, ARM provides a GCC toolchain (https://developer.arm.com/Tools%20and%20Software/GNU%20Toolc...) but makes you purchase the LLVM-based one (https://developer.arm.com/Tools%20and%20Software/Arm%20Compi...). Now I'm fairly certain there's an open-source way to target ARM microcontrollers with LLVM but you can't download it from ARM...
$ rustup target list | grep arm | grep none
armebv7r-none-eabi
armebv7r-none-eabihf
armv7a-none-eabi
armv7r-none-eabi
armv7r-none-eabihf
... which means the publicly-available open source version of LLVM also supports it.edit: looking into this more, yes, upstream clang/llvm should work, though it doesn't include a C stdlib for this target (of little import to Rust, presumably), so you'll have to find one (e.g. by taking it from the GNU toolchain) if you want to use a C stdlib on embedded. Who knows what secret sauce armclang adds to clang...
I briefly looked at the license and it seems that if you're getting it for free, then its all good, you just get no support.
(I only write Rust on ARM and so am unfamiliar with the various toolchains they offer, honestly.)
edit: My legalese is not great, but isn't the the license saying that you can't legally distribute software compiled with the LLVM-based ARM tools if you obtained it for free? Obviously the gcc-based tools can't have such restrictions.
3.2 NON-COMMERCIAL USE AND FREE OF CHARGE LICENSES:
...
(b) if you are receiving a Non-Commercial Use License or version (as applicable) of the Arm Tools:
(i) you and your Permitted Users may use the Arm Tools for internal use only; and
(ii) you are not permitted to distribute or sub-license (A) any part of the Arm Tools, or (B) Your Software,
Your Hardware, or Your Reports developed under this License using the Arm Tools. The Arm Tools shall be
used only by you and your Permitted Users, and you shall not (except as otherwise authorised in writing by Arm)
allow any other third party whatsoever to use the Arm Tools.
For the avoidance of doubt, if you are receiving a Non-Commercial Use License and the license is provided to you free of
charge, the restrictions in both sub-clauses (a) and (b) above will apply to your use of the Arm Tools.
additional edit:Here's what happens when I try to run armclang:
$ armclang
armclang: error: Failed to check out a license.
The license file could not be found. Check that ARMLMD_LICENSE_FILE is set correctly.
I did not bother to figure out if there's a way to get a license without paying (though I would likely qualify for a non-commercial license as an academic user, but it seems like a pain).But who knows, honestly they make this stuff as confusing as possible, it seems.
GCC targets more platforms than LLVM, that much is true. But beyond that, it's pretty common in the microcontroller world for a vendor to release their own specially patched version of GCC blessed for a given platform. They simply aren't doing that for the LLVM.
Those blessed GCCs aren't often seeing the hacks merged upstream.
Here's one such example:
https://www.ti.com/tool/MSP430-GCC-OPENSOURCE?keyMatch=gcc&a...
In what concerns Apple, I think they mostly care about the C++ support needed to keep LLVM going, the C++14 based dialect for Metal Shading Language, and the subset used across IO and DriverKit.
For everything else there is Objective-C and Swift.
Then Google apparently drop off clang after the ABI break votes didn't went the way they wanted, so they are now focusing on Abseil, and their style guide is anyway quite restrictive.
So now we are in this ironic situation, that VC++ from all compilers is the one with best C++20 support, closely followed by GCC, and then there is clang and the other lesser known ones still lagging in C++17 and earlier.
It's so extraordinarily, mind-bogglingly stupid. But I guess it is time for the ageing C++ queen to pushed out of the mortal coil and let Princess Rust grab her long-overdue system crown.
Did you know on MSVC a std::mutex is so huge it needs more than one cache line ? Obviously Microsoft aren't stupid, they know how to fix that... but it would change the ABI so they can't touch it.
And so the Zero Cost Abstraction promise "Don't pay more than it would cost if you did it yourself" becomes "Eh, just do it yourself" everywhere - C++ programmers learn to hand roll all the basic stuff they need, because ABI stability has ensured the standard library mechanisms are slow, or bloated, or both.
Either Rust wants to play on that field, or it doesn't.
Microsoft has broken their ABI plenty of times, and VS vNext might be when the next break will take place, which was initially planned for VS 2022.
Bashing C and C++ ABI issues on HN will do very little for those businesses to adopt source code distributions.
Some famous C++ frameworks that ship as binaries are OWL, VCL, FireMonkey, MFC, ATL, WinUI.
And WinRT has improved upon classical COM ABI to provide even more data types.
On Windows there is a component market selling libraries for those ecosystems.
Then we have middleware vendors for game consoles as well.
Maybe this isn't a market Rust community cares about and that is fine.
Microsoft is stupid enough since the first decision of the implementation impacting the ABI is made. There was no one enforcing such bad decisions shipped into the productions at the very beginning. Both libstdc++ and libc++ have more flexible rules to preventing ABI breakage.
At least from the comments I sometimes see flying by on Reddit.
GCC had laughably bad error messages before LLVM caught on. Even then, LLVM generated terrible code compared to GCC. There's a similar competition going on with open source linkers (which are finally going multi-threaded).
Both compiler toolchains currently blow pre-LLVM GCC out of the water (and GCC development has noticeably accelerated since LLVM came out).
I'm not sure which one is better at which C++ thing these days, but I'll bet GCC Rust will beat LLVM Rust at some important things in a few years (and then vice versa).
1. What the borrow checker protects against, and what constitutes valid code.
2. What the borrow checker can actually prove is valid code.
#2 tends to be much less than #1, because the compiler can only be so smart, and if it can't prove something correct (even if it might be correct), it takes the safe route and refuses to compile the code.
#1 is probably pretty well specified, or at least wouldn't be hard to write down if someone really wanted to. #2 is a moving target, because the borrow checker in rustc gets smarter with some releases, and compiles code that it used to reject (because it's been taught to understand that code better and can prove it correct). So it's a lot harder to specify exactly what kinds of code the borrow checker will accept and reject.
I could easily see a situation where gcc-rust claims to support all the language and stdlib features that a particular version of rustc supports, but has subtle differences in borrow checker behavior that causes it to reject some code that rustc accepts (or accept some code that rustc rejects). The behavior isn't incorrect, per se, but it would be frustrating for developers.
It's also possible there could be similar issues with lifetime analysis.
Rust is designed to compile fine without any borrow checking (it reduces it to the C level of safety, but valid programs generate valid code).
GCC will compile itself without a borrow checker. Then it will compile the existing borrow checker written in Rust, and then recompile itself with borrow checking.
Did they announce a change of plans? Their website just says that they have no plans for a borrow checker (i.e. it's not required to actually implement rust).
From the website for the project:
> There are no immediate plans for a borrow checker as this is not required to compile rust code and is the last pass in the RustC compiler. This can be handled as a separate project when we get to that point.
Link here: https://rust-gcc.github.io/
rust intentionally has a fairly rapid release schedule, so that you don't end up with a big release every year that introduces several possibly breaking changes.
Features aren't randomly scheduled for each release. Instead each release is a snapshot of what's been stabilized by that point.
If you prefer some arbitrary concept of stability, then you can quite easily lock to any version of rust this far and ignore any new versions till something catches your fancy.
The Rust language is fine. Great, even. Most Rust devs are bleeding edge types that always use the latest features. Hopefully this changes as Rust matures.
Perhaps you mean forwards compatibility, in which case yes perhaps that's an issue but I'm not sure why you'd want bigger, slower release cadence when it would likely make your problem worse , when libs update but updating your compiler potentially causes significant other changes as is the case with C++ for example
* Want to compile the software, not just install a pre-built binary * But, don't want to actually hack on the software and so don't care about up-to-date tools
... is probably rather small. Even for them, though, surely just an old source version will work? Or is there some reason that doesn't help?
There is no sense in which this increases the pace at which these new features are developed or released; if Rust released a new version yearly, we'd simply see extremely large releases some years (including, say, the new async/await system) and relatively small releases in others.
What it does mean is that backwards-compatible refinements and bugfixes can be quickly and easily added to the language, which I think is a good thing.
I'd be interested to know what you'd prefer, though!
Old compilers didn't use this concept. So a C->x86 compiler and a Fortran->x86 one couldn't share code as easily. And if you wanted to expand to, say, ARM, Alpha, etc. targets, it gets worse. You end up needing O(l*t) compilers[a] to be homogenous. It gets even worse if you want to add optimization. With this model, the whole system is working on an AST, and therefore, the optimizers are tailored to the individual compiler varient, so they end up being just as unportable.
However, with a modular design (front and back end), you can have C->IR and Fortran->IR front ends, then a single back end for each target: IR->x86 and IR->ARM. If you want to add, say, Ada support, you only need to write an Ada->IR module, and the back ends handle the rest. In the end, you only need O(l+t) modules. You also get the benefit of agnostic optimizers as they only need to support your internal IR.
[a]: 'l' is languages and 't' is targets
LLVM's original purpose was something analogous to the JVM, but it's obviously evolved to be more of a compiler platform, debugging toolchain, etc these days
(No direct hardware support, though, as cool as that sounds)
Meanwhile normal instruction sets are just a bytecode optimized to run quickly.
LLVM-IR at least also isn't stable, so not a good thing to bake into a chip - but you could conceivably make a stable intermediate language.
Intermediate code tends to have downsides that make it hard to work with in an actual hardware implementation. Infinite registers, for example, which makes encoding instructions challenging.
However they apparently don't want to keep fixing upstream changes and is now deprecated going forward.
As Blackthorn mentioned, specifically it can be done. There was a chip that ran JVM bytecode directly. There was an entire Lisp Machine which was very influential... but not because of its high performance.
Wikipedia has a page on this: https://en.wikipedia.org/wiki/High-level_language_computer_a...
Many people decry the lack of innovation in computer architecture over the past 50-60 years. There is a legitimate reason for that, though: Any such innovation had to outrun the exponential increases conventional architecture was experiencing. We had it for so long we could take it for granted, but while computing wasn't the first and won't be the last to see periods of exponential growth, I can't think of anything else that had the length of run that computer chips had. (And even now, it's only slowed down. There's still multiplicative advances rather than additive ones every year. The base on the exponent is just smaller than it used to be.) There's been a number of architectures based on running IR (or an equivalent directly) that didn't pan out, but there's also some that did run OK, but they just couldn't keep up with the exponentially accelerating behemoth.
I think if silicon continues to plateau, there is some hope that this could change. However, a new challenger has appeared! Now it's not good enough to just outrun what a "conventional" architecture can do with CPU and RAM and peripherals... now you also need outrun what a GPU-based architecture can do! This sucks up a lot of the oxygen. For instance, right now you see some tentative steps towards special-purpose AI silicon, but it has to do something amazing to beat out a GPU, and you also have to be really, really sure that the AI technique you are committing to silicon will still be the state-of-the-art in the couple of years it'll take you to bring it to market. AI hasn't been moving as fast as, say, JS front-end frameworks, but it's still moving fast enough that I'd be nervous in investing in that vs. the same amount of effort poured into optimizing a GPU algorithm on GPUs that are still advancing a lot every year.
The bigger issue (IMO) is that the IR for compilers tends to evolve rapidly while hardware is stuck in the mud. Moving the problem of finalizing the IR into the hardware will effectively make it so you'd change your compiler to emit IR that targets old IR.
This is why, for example, x86 will get a new instruction which will effectively go unused for 5 to 10 years (except in extreme cases where the performance gains are worth the cost of writing code to detect and use those new instructions... For example, FMA). Those who compile code want it to be able to run most anywhere in not so surprising fashions, so they'll target the LCD.
Now, imagine someone was building a JVM bytecode chip today. The JVM is on a 6 month release schedule. That's incredibly hard for a hardware manufacturer to want to keep updating their chips at that rate. Further, even if they targeted "LTS" versions, they've moved that to a 2 year cycle. Again, hard to really expect customers to want a new JVM chip every 2 years.
The likes of GCC and LLVM IR change and expand just as rapidly (if not moreso) than than JVM.
There is the option of something like FPGA or programmable hardware leaking into more places, but IMO, the HDLs and their tools simply suck too bad to expect FPGA on an ASIC to really take off. We need a first mover here and after that happens, I expect at least 5 to 10 years before the tools get to a level that doesn't completely suck. Even then, while access to GPGPUs is now pretty much universal, GPGPU programming still feels like it is in the stone ages. Despite being a thing for over 10 years. So how can we expect HDLs to evolve at a faster pace?
I'm not sure that the vaporware Mill CPU acts as an existence proof for anything, to be honest.
The problem of infinite registers is pretty easily solvable. Heck x87 already has that basic concept down with the stack registers. The only missing piece is moving overflow values onto and off of the stack. It's a fairly easy to solve problem.
That's effectively what Mill proposed in their docs. The "belt" notation was their cute way of doing that.
You still have a limited window of 8 registers. Need 9 simultaneous live values? Oops, too few registers! In order to encode infinite registers, you need to be able to support referring to an arbitrary number of registers at once, which means you need an arbitrary-length register number specifier... have fun handling that in hardware!
Already addressed and not really that hard of an issue to solve. Hardware already has all the logic embedded in it to be able to load or store memory into/from registers.
The logic would mirror the logic done by compilers when they are allocating registers and deciding what needs to be evicted to the stack.
The only missing piece is where that memory should be, but that's really as simple as having either the OS or the application logic allocate a special memory region for the purpose of register evictions and loads.
I'd be interested in reading an article from someone who has been doing it for 10 years as to why that is the case. I have theories but nowhere near enough direct experience to evaluate.
(My hypothesis is that the extreme parallelism makes it so very tiny mistakes have catastrophic performance impact by introducing accidental serialism, and as a result, it is very difficult to create an "easy to use" framework that doesn't abstract too much away and make it trivially easy to introduce even a tiny such error and crash performance. We actually make this mistake all the time in conventional CPU code, it just generally just costs you small integer multiples of performance instead of large integer multiples of performance.)
We have 4 major GPGPU manufactures (Apple, Intel, AMD, nVidia) and none of them want to make creating an open standard easy. There is no reason why nVidia should support OpenCL the same way that AMD does, they want people to write CUDA. There's no reason for Apple to support OpenAAC like Intel does, they want people to write Metal... etc. These 4 companies are trying to push everyone into their proprietary ecosystem or into an open ecosystem where they have a large say.
I had hoped that SPIR-V would be a good inroad to start fixing some of these problems, but alas, it seems like Apple and nVidia aren't big fans. It's early still, though, so maybe that changes?
OpenCL 2.x was a failure, hence why OpenCL 3.0 is basically 1.2 with everything else from 2.x marked as optional.
Short answer, it’s feasible in theory, but usually suboptimal.
https://en.wikipedia.org/wiki/Jazelle
Jazelle was an instruction set extension to allow natively running java bytecode on Arm processors. Arm never seemed to give it any proper care and attention necessary for it to make any actual market impact, they didn't even publish the ABI.
Transmeta, that used to employ Linus Torvalds at one stage, had Code Morphing Software: https://en.wikipedia.org/wiki/Transmeta#Code_Morphing_Softwa...
It was built in to the silicon, that essentially operated as a full JIT for x86 software, translating to back-end VLIW instruction set. Transmeta seemed heavy on the marketing and buzz (hiring Linus was very much a marketing move), light on the actual execution. Crusoe didn't even remotely live up to their own buzz and it just about killed off their chances of success.
I think Rust on LLVM has a major example of this? IIUC, it's a big problem to turn on strict aliasing because LLVM doesn't actually implement it correctly, it just works well enough for C/C++ etc.? (But I'm no expert.)
Still, the frontend/backend split gets you a lot closer to supporting all of the l*t pairs than the all-in-one approach. Porting is a matter of fixing bugs and filling in missing parts rather than starting from scratch.
AIUI, the solution wasn't that Rust smuggled the language-level construct through the intermediate layer, but that they simply held off on expressing these constraints until bugfixes could be merged upstream into LLVM. The only downside to disabling this was simply that compiled code would potentially be a bit less optimal than otherwise, so such tactics weren't really necessary.
It's noalias (C99's restrict code, in essence), that you're thinking of, not C/C++ strict aliasing rules.
This is most common where C++ (as the main consumer of these semantics) doesn't define certain semantics, or, the semantics it standardises are just impossible to optimise so nobody really delivers them (ie they don't always work in your C++ programs when you compile them with actual modern C++ programs)
As I understand it an example when it comes to aliasing would be what happens if somebody typed a bunch of ASCII into a prompt and then the program... uudecodes the ASCII to get an integer, and begins just using that as a pointer to rummage around in some data structure.
Is that... OK? Obviously you can't do this in safe Rust, but even unsafe Rust says er, no, that's definitely not allowed (it might actually work but it isn't OK). But the ISO C++ standard says that so long as the ASCII string happened, by some cosmic accident, to be a "valid" address for a pointer after this transformation this works fine even if it now magically aliases a pointer we otherwise had no reason to believe could be aliased.
If we allow this, many optimisation opportunities vanish. So, LLVM doesn't allow it. Our "correct" C++ pointer uudecoding program doesn't work. However, a blanket prohibition blows up real tricks that people actually do, such as hiding flag bits in address values (unlike uudecoding ASCII inputs to make pointers). So LLVM provides behaviour that's not formally standardised anywhere but can be thought of as something akin to PNVI-ae (Provenance Not Via Integers - Address Exposed). If the program "exposes" addresses from pointers, then the compiler assumes it could see those addresses "magically" appear from somewhere else (e.g. a uuencoded string) and so it must choose optimisations accordingly.
One day perhaps C++ will actually document PNVI-ae or some similar scheme in the ISO standard. That day is not today (nor next year, this will not happen in C++ 23) but meanwhile you've got the problem that, in unsafe Rust you actually get whatever arbitrary semantics that were delivered by LLVM. Since they're not just "This is what C++ does" but only "This is how C++ works in Clang" that's even less portable.
As I wrote, this only burns unsafe Rust. If you don't write unsafe and just depend on other people's stuff to use that as necessary (e.g. obviously the standard library is full of unsafety) then it's not your problem to fix this when it breaks. But it sure would be nice if the people writing unsafe code would have more certainty on this sort of topic.
Rust is experimenting with "Strict Provenance" rules to see what happens if, instead, it makes its own provenance rules and dispenses with waiting for C and C++ programmers to actually decide what the rules are in their languages. https://doc.rust-lang.org/std/ptr/index.html#strict-provenan...
Essentially, C and C++ already have a version of provenance for pointers and references, but it isn't identical to the Rust version.
Because some old codebases pull tricks of this kind (messing with pointers, doing casts without going through unions, etc), gcc (and clang) have a flag, -fno-strict-aliasing , to disable some optimizations that would be enabled by taking aliasing/provenance into account.
* Rust doesn't have strict aliasing/tbaa
* Rust does have "restrict" semantics (in C via the keyword and in C++ via common vendor extensions)
And yeah, Rust is considering a provenance model that's different than PNVI-ae.
None of this matters in Rust without unsafe. Pointers technically exist outside unsafe, but you can't do very much with them (you cannot, for example, dereference a pointer). And yes, Rust says that the attempted type pun you describe, which you would need unsafe to create, is Undefined Behaviour.
> you could use unions and get implementation-defined behavior
Type punning with unions is also Undefined Behaviour in C++. The sanctioned way to perform type punning is to memcpy() the data from type A to type B. C++ 20 provides a built-in way to ask for this to be done std::bit_cast
The C++ Standard is OK with this magic trick. Executing this trick in the C++ abstract machine is no problem at all.
However in a real world compiler this is a huge problem - because the optimiser assumes I can't possibly have such a pointer. Where would I get it from? This is where provenance comes into the picture. Where did my pointer come from, that's what provenance means. The C++ standard has nothing to say about provenance but your compiler depends on it.
I've been doing this stuff for a long time, going back to egcs and even before.
Nope, it says nothing whatsoever on the subject of provenance. Periodically people make a run up and try to get this fixed. N2676 is currently in front of WG14 (ie the C standards committee) with the long term hope that if WG14 takes this fix, or something like it, WG21 (C++) could be persuaded to eventually take a similar fix - although lots of people don't like N2676 and want something else (the more vague your "something else" the more popular). If the standard had "a whole lot" to say about provenance you'd be able to quote some of it. but I suggest reading N2676 for an example of what is not yet standard
http://www.open-std.org/jtc1/sc22/wg14/www/docs/n2676.pdf
> has detailed aliasing rules and rules about which pointers are valid
Aliasing rules forbid some type punning shenanigans, but again, and I will repeat myself, that's not what's going on here.
DR260 (Defect Report number 260, about twenty years ago) asks WG14 what their standard says about such magic tricks, and their response basically says well, compilers are allowed to somehow know about provenance, so actually our Standard is correct and this is working as intended. What's working as intended? Well, whatever your compiler actually does.
What a great standard! Note that this response isn't incorporated into subsequent versions of the standard, it's just basically known in the industry, oh yeah, that's DR260, don't worry about it.
> I've been doing this stuff for a long time, going back to egcs and even before.
That's nice, but it's not terribly relevant here, except that it means you remember an era when people didn't even realise this was a problem. It still was a problem, they just didn't know it was a problem yet. And hey, you know about that experience too, because apparently you didn't know this was a problem after 2004 either.
Rust would like not to kick this can down the road for 18 years and counting, which is why Aria Beingessner's experiment is happening.
The front-end compiles your input language (C, C++, Rust, etc) down to a low-level language called an Intermediate Representation. Then the back-end of the compiler optimizes the IR and compiles it into object code. A family of compilers will usually share the back-end.
This kind of split allows for deduplication across different compilers in the same family, and also makes it easier to design a new language without having to fully re-implement everything about the compiler yourself.
In GCC, these are bundled together as a common project afaik. In Clang/LLVM, these are split into the front-end (clang) and the back-end (LLVM).
I knew that Julia uses LLVM. I also knew that Julia has something called IR, but i didn't know what it was. So Julia's IR is probably LLVM's IR...
For some examples, the compiler plugins interface acts on typed IR, so things like Diffractor for automatic differentiation or JET static analysis and EscapeAnalysis.jl act there. Notably, Zygote automatic differentiation acted on the untyped IR. GPUCompiler.jl does some operations on the typed IR then lowers to the LLVM IR level and then intercepts the normal compilation process to choose alternative backends (e.g. compile to .ptx for compilation of native Julia to CUDA). Enzyme.jl uses the Enzyme LLVM-based automatic differentiation, so it does some actions on the typed IR before lowering to LLVM and then injecting the Enzyme LLVM pass before compiling via GPUCompiler.jl (now to the normal CPU backend).