Thoughts on what a next Rust compiler would do
matklad.github.io
matklad.github.io
So if I could choose, I'd invest money into a new ABI/linker that reconciles polymorphic languages with separate compilation without demanding a uniform object model. Not only Rust would benefit from such a linker.
I have no complete solution for how such a thing might work, but I think it's possible because the actual specializations that stem from monomorphization are very few. It might simply be possible to create one function for every possible monomorphization ahead of time. Even if this turns out to be futile, the actual specialization only needs to change very few things in the code, so monomorphization could effectively be handled by a dynamic linker.
This is true in only the simplest case. Already if you have a polymorphic function that sums the elements of an array, you'll be doing function calls to string concatenation if those elements are strings but would really like the vectoriser to produce nice SIMD code when specialised to floats.
Specialisation is something that should happen before the optimiser.
That's true, but your example shows a very important point: These cases are limited and (in current languages) known by the compiler developers. I don't know a language that would allow a user-defined datatype to come with such specific optimizations. So for today's languages one could probably enumerate the special optimizations before specialization and then later only dispatch to the optimized code.
But even if we say that optimization must happen after specialization, I still think that these optimizations are general enough to be put into a dynamic linker. That's what all JIT engines do, after all.
Can you give more specifics on this?
https://github.com/rust-lang/rust/pull/96709#issuecomment-11...
I don't think Rust should tackle any ambitious project to rewrite the compiler while these basic concerns remain.
A brand-new Rust compiler may well never happen, but if it doesn't happen it's because the business value for a complete rewrite isn't there, not because of some mass departure of compiler developers.
Right now, the weak points are mostly library side. Too many 0.x version crates where the API is still in flux.
There also seems to be a general trend towards sophistication and away from simplicity. That’s a tradeoff that makes me personally wary, rather than excited at this point in my life.
Much of that is to be expected, given the age and the unique value props of the language. But the stability to hype ratio seems unattractive overall, even though the language and tooling are great to a large extent.
Rust’s safety is better than C/C++, but memory vulnerabilities are found in cargo packages because of unsafe code.
I do, actually.
> but memory vulnerabilities are found in cargo packages because of unsafe code.
Forbid unsafe in your config and you're good.
Safe rust should be. And C# is a notable example of a GCed language that also has 'unsafe'. Along with most of the GCed languages with FFIs.
It’s pretty crazy how many people will argue this point and downvote comments about it. It’s like a C++ dev claiming they never write memory bugs. I love rust, but the community pretending like memory safety bugs are impossible is as annoying as people claiming memory safety isn’t important.
EDIT: And GC languages “can” call unsafe is not the same as a language that has unsafe in it. It’s laughable if you are claiming there are the same memory bugs in say Elixir as there are in Rust.
If you already know what you want, why pretend there is no solution?
But you skipped the other half of the post, where I explained that the average GCed language either has unsafe code or can easily call unsafe code or both, just like Rust.
As I said in the edit, GC languages are safer than an unsafe language like rust, even if you could find and call unsafe code.
Again, I don’t know why to someone that uses rust I need to make that point. If a C++ dev argued using C++ is fine cause all GC languages can call unsafe code, would you agree? Cause if you do then we might as well end rust now, it’s not adding any safety. All language are unsafe!
- A std library that provides much of the stuff you need to write web servers.
- A few select well established, stable libraries to fill the gaps
- A battle hardened database like SQLite, MySQL or Postgres
- Maybe additional libraries and generators to program against specific, open protocols, standards and over the wire schemas, like json-schema/OpenAPI/GraphQL/Protobuf etc. Which are typically tested against a common test-suite.
- Maybe additional components like message queues and key value stores and object storage in order to scale.
You get the drift. It's not like these things are some random libs and tools that are maybe "unsafe" (whatever that means).
A good thing about Rust is that unsafe is always explicit. And if you're not using the keyword, then satisfying the type checker gives a "if it compiles, it runs" feel, because runtime assertions can be avoided.
But saying "anything is unsafe anyways" is missing the point. The language itself can nudge you towards a certain outcome, but I trust some of the mentioned technologies above, not because of the languages they use, but because they are well designed, stable and full of old battle scars.
Every language ecosystem that I have used has a good subset of its popular libraries implemented in memory-unsafe languages (usually C or C++), and the use of these libraries is typically unavoidable if you need high performance or to interact with the OS. Which is to say, most large projects in GC languages end up pulling in memory-unsafe dependencies.
Rust's advantage is that at least your dependencies are not 100% unsafe code, but have unsafe parts limited to small, easily auditable blocks.
I can't donate my time, but I'd be happy to donate money to anyone taking on this effort.
On that note, I'm looking for more Rust crates and maintainers to donate to. I just started paying Bevy, and I'm looking for more areas of the Rust ecosystem that need it.
This language brings me immense business and personal value, and I'm happy to give back.
cwfitzgerald for wgpu. He took over as the head maintainer after kvark left Mozilla to go to Tesla.
winit maintainers, they're a core ecosystem crate but don't get that much support because windowing isn't flashy.
It could definitely be better, sure.
So, I get your point for production build, that it needs some special optimization that maybe doesn't happen for Go, but I don't need any optimization when doing Dev testing and this could well remain totally unoptimized. Then I would at least expect for an "un-optimized" dev Rust build to be roughly as fast as a production Go build. At least not >10x slower
Or of course if the Go compiler started doing them so it became a lot slower.
One example is “full”monomorphization of generics. But there are lots of such examples.
Usually though I find the Rust compilation speed a non issue, so long as “cargo check” and the IDE analysis is great, I just rarely compile to begin with.
As other commenters have noted, the Go backend is extremely simplistic ("barely an optimizing compiler"). Rustc is also a lot faster in -O0 mode than any optimizing mode. But -O0 is somewhat dumber than Go's compiler. You could imagine an -O1 (or -O0.5) with compilation time vs performance tradeoff similar to Go, but it isn't there today. There's also an attempt to adopt a more simplistic backend (cranelift) into Rustc to improve codegen performance, but so far it hasn't provided much speedup.
Here is an older article that compares compilation times between multiple different libraries, where one of them does spend a significant (and majority of the) time borrow checking: https://wiki.alopex.li/WhereRustcSpendsItsTime
Edit: trait-based OO is a breath of fresh air compared to tree-style class inheritance. Super flexible without having to overthink.
Immutable by default is reassuring, Option/Result are fantastic null/exception replacements.
Enums and match blocks are powerful and gracefully help ensure handling of all cases, have nice syntax, and work well for a systems language.
I programmed Ruby for years before learning Rust, and aspects of Rust's library and language inspiration from Ruby are obvious to me: the closure syntax closely resembles Ruby's blocks, for example, and much of the core iterator APIs match their Ruby counterparts. It's a very different language overall, of course!
(Then of course there's the part where many early Rust contributors came from the Ruby community.)
If you think strong typing is "ceremony" then you probably spend most of your time writing (and documenting) code and not reading, refactoring, or collaborating on it.
The time people spend writing unnecessary, buggy unit tests (that static analysis can do in better languages) is far greater than the time to just use the type system.
Most languages don't even force you to be explicit anymore. They infer the types, so you don't even write extra code. You just get better errors and speed.
JS/TS just requires fewer decisions per line of code. As a pithy example, in javascript I don’t have to decide whether I want a String / &str / SmartString / Rc<String>, etc. It’s just string. When I pass a callback I don’t have to decide between accepting a closure or FnOnce/Fn/FnMut. Or decide whether to accept a function pointer or take a generic parameter. In javascript I never have to think about the lifetime of the callee’s stack.
I tried to implement my own server-sent events style protocol in rust a couple years ago. I spent 2 weeks trying a bunch of different approaches but I eventually gave up - I just couldn’t get it working. (Mind you, async was pretty new then - I don’t think it was ready). I moved to javascript and had the whole thing working correctly in about a day and ~50 lines of simple code.
I think the resulting program is much better when I write it in rust than what I get when I build on top of nodejs. Rust programs can be orders of magnitude faster and I have a shockingly low defect rate in shipped rust code. But there’s a trade off. Rust programs take more effort to write.
Rust is a brilliant systems programming language. I’d pick it over C or C++ any day of the week. But not all problems are systems programming problems.
Never mind that some of these choices you can just avoid by thinking in general terms and deferring optimization for later. Just passing String and FnMut everywhere will be a lot more efficient than whatever the JS translates to on the machine.
> Just passing String and FnMut everywhere will be a lot more efficient than whatever the JS translates to on the machine.
I can’t comment on the performance of FnMut, but I’ve seen rust code run slower than the equivalent javascript because the rust code in question allocates everywhere without thinking about it (via Vec, String and Box).
In my mind, javascript and friends are good languages to build things fast. Rust is a good language to build things right. For plenty of software (eg websites), good enough is good enough.
Absolutely. Thank you for speaking truth.
Lifetime annotations help how you use and exchange certain data, but you can get away without them unless you need them. For shared data structures (when synchronization primitives are too much) and highly performant code.
You can write high level Java, Ruby, and Python code in Rust if you want to. Writing Actix or Axum web handlers feels no different than Golang or Python/Flask.
let mut v = Vec::new();
v.push(m_foo);
In fact it's so nice that you often rewrite code to keep clear what's actually happening, even though you could keep types largely implicit.Here's another example:
let f: Vec<i32> = vec![1, 2];
let f2: HashSet<f32> = f.iter().map(|n|n.into()).collect();I'm not sure what it is - there just always seem to be random symbols that dont seem to follow conventions of other languages. Guess I am use to C-style (including C++, C#, JS) syntax and python. Its possible that there are just new concepts that aren't really represented in the languages I use too.
Again, I don't necessarily think its bad. Just what I don't find intuitive.
What you're actually seeing is a trap a lot of programmers fall into where if something is unfamiliar then it's treated as if it's incorrect.
I said that I don't think it's bad or wrong.
One of the hopes is that it's usually not necessary to write those lifetimes explicitly in most programs, unless you're doing something unusual.
If you could go back and change it, what would you have used for the lifetime syntax?
&'a X
looks less like Perl / line noise than this:
&@a X
. In any case, not something we'd be likely to change at this point, but I agree that ' is effectively "arbitrary unique symbol with no evocative meaning".
(Arguably the same is true for things like & for "address of", but that has a long history in many languages which makes it more intuitive for current developers.)
&x\a
The current tick is much easier on the eyes.Other than 'static, a Rust beginner is unlikely to run into any lifetime with a meaningful name for quite a while. If when learning a new language you only ever saw variables named a, x, t, s, v and m it's not a big jump to guess that p, z and b would be allowed as well, but would you assume that in fact the language allows big_set and old_corp_logo as identifiers and it just forgot to mention that ? I'm not sure I would.
The Rust standard library does have lifetimes with reasonable names, for example std::thread::Scope needs two lifetimes and it names them 'scope and 'env which, although brief, are clearly not single letters. But most of the library and documentation doesn't bother with meaningful names, since it doesn't have anything worth saying about the lifetime, e.g. the signature of str::ends_with:
pub fn ends_with<'a, P>(&'a self, pat: P) -> bool
where P: Pattern<'a>,
<P as Pattern<'a>>::Searcher: ReverseSearcher<'a>I think it'd be a really good idea for the standard library to change many of its lifetimes to be descriptive, to encourage others to do the same. With my libs-api team hat on, I'd happily merge PRs that do that. (In small batches, please, not the whole library at once.)
It's the same syntax in Lisp, Haskell, Ocaml, Ada and VHDL.
I've put together a playground at https://play.rust-lang.org/?version=stable&mode=debug&editio.... It's not just dereferencing invalid */& that's illegal in Rust, but constructing and dereferencing valid &/&mut in ways that don't respect tree-shaped mutability. Rust's pointer aliasing rules invalidate otherwise-correct code, placing roadblocks in the way of writing correct code. There's so much creation of &mut (which invalidates aliasing pointers for the duration of the &mut, and invalidates sibling &mut and all pointers constructed from them), that's so implicit I don't know what's legal and what's not by auditing code. (Box<T> used to also invalidate aliasing pointers, but this may be changed. The current plan for enabling self-reference is Pin<&mut T>, but the exact semantics for how and when putting a &mut T in a wrapper struct makes it not invalidate self-reference and incoming pointers, is still not specified.)
(I've elaborated further at https://news.ycombinator.com/item?id=33658253.)
But it leaves that work to the developer, thousands of hours of testing and reviewing to ensure no little corner case is missed.
I think the argument that unsafe is used making the language have “holes” is somewhat misdirected. When I see a rust implementation of a doubly linked list with unsafe I know exactly where I’m on my own (a few lines) and where the compiler does the job for me (the rest of it). It’s not as if that means safety or flexibility is out the window. It means “do the manual safety review on these two lines”.
That's not true, because you can wrap unsafe code in safe interfaces and use it from safe code.
With C, you can only do these kinds of safety guarantees with a lot of discipline. With Rust, you can offload much of your discipline to the compiler.
For example you might build an abstraction named ReadWriteMutex that guarantees safe access at runtime.
I don't see a scenario where this would impede performances. The invariants to validate are the same. But rust would guarantee that it is memory safe.
If you do need to do something this exotic, the fundamentals don't change: you still need to take care of memory lifetimes, you still need to ensure thread safety, etc. I really don't think it's as all-or-nothing as you believe it to be.
I hear what you’re saying, but that hasn’t been my experience with rust. I love C, but I’ve been writing rust for the last couple of years. And rust has stolen my heart.
You’re right that unsafe rust is a bit less ergonomic than just writing C. Sometimes I miss void pointers. I definitely miss how fast C compiles. But most rust isn’t unsafe rust. Even deep in my custom in-memory btree implementation I think more than half of my methods are safe. And safe rust is a fabulous language when you can use it. There’s all these bugs you just can’t write.
Early on with rust I had this magical experience. My program segfaulted in one of my tests because of memory corruption. It was “spooky action at a distance” where the segfault happened well after the buggy code was executed. Bugs like this are awful to track down in C because the bug could be literally anywhere. But when I looked at the trace, there was only one unsafe function which could be causing the problem. (Since I could rule out all my safe code). Sure enough, 10 minutes later I had a fix.
I have a skip list for large strings that I ported from C to rust. I still have no idea why, but the rust code ran about 20% faster out of the box than the original optimized C code. And yet, the code is significantly smaller and easier to work with. I’ve added a few more optimizations since then - it’s about 10x faster than the C code now. The new optimizations are all from tricks I never got around to adding in C because I was afraid I might break something.
And then there’s the things rust does well that aren’t in C at all. Cargo is incredible. Monomorphization is the right tool for a lot of problems, like custom collections. (Higher performance and types? Yes please!) Parametric enums & match statements are so much better to work with than C enums and unions. Then there's rust's std, which is fantastic. I use Option, Result, Vec, BtreeMap, PriorityQueue, the OS-agnostic filesystem API, and so on daily.
So yeah, I hear you about C being lovely. But I think rust is even better. Rust is a bit of an ordeal to learn but I'm really glad I learned it. Its a joy to use.
But the compiler is really slow. Even an incremental build on a mid size project is never below 10s, whereas on a similar size project, Golang will take me less than 1 second to build incrementally. Also, I find cross-compilation (from/to major platforms) really challenging on Rust, as opposed to Golang where it's for my use cases super straightforward
What I want is a Rust, but with almost everything stripped away. Complexity similar to Go.
ocaml has potential but I strongly dislike the numeric operators, and I'm not a huge fan of the syntax in general.
OCaml with python like syntax would be damn near perfect. I could even overlook the operator thing.
I also want a more lightweight syntax but beggars can't be choosers.
It's like saying "I would like to have a non-volatile storage device that is as long-lasting as microfilm, as fast as an SSD and as cheap as a HDD." At some point, you have to renege on at least one of your constraints.
This is not to deny that Rust also has other sources of complexity. But, most of Rust’s complexity can be traced to its fundamental goals of safety and performance. If you’re willing to sacrifice one of those, there are much simpler and easier languages to use.
it's true golang is smoother, but the language is a lot messier (idioms and typechecking).
Can't conclude anything yet, but it makes me think back of rust more often than others.
I do 100% agree on your commends on cross compiling though. I also don't like how the target for linux has "unknown" in it.
For me the biggest hurdle to adopting Rust is that rust-analyser turns my laptop's fans up to 11. I don't know whose problem that is, but it doesn't happen with other languages, and my black box impression of rust's compilation infrastructure is that it's fucked up.
It’s just tricky because writing a new rust compiler is a monster amount of work because of how complex Rust is. And rearchitecting the existing compiler would also be a massive undertaking because there’s so much existing code, and it’s spread over rustc (written in rust) and llvm (in C++).
A faster compiler probably wouldn’t bother with all of llvm’s intermediate representations. Refactoring across a code boundary that spans two languages (and two projects) sounds utterly exhausting.
Rust is the first language in its category. It won't be the best language we ever make in that category.
Since a major thrust of the linked article is barriers to rust adoption it seemed relevant.
Correct.
I think it’s a symptom of there not being enough places where adherents of different programming ecosystems can argue back and forth about our choices. I think there’s an unmet need in our community to have those conversations, and there’s probably something healthy in that.
It definitely distracts from the specifics of the article though. The corresponding discussion on reddit is much more technical:
https://reddit.com/r/rust/comments/10ld2vn/blog_post_next_ru...
Can you give some examples of this? If there are ways we could simplify the language without hurting existing use cases, we should consider doing so.
The good thing about the default hello world example is that it introduces you to both language features without being overwhelming (IMO).
You can read more about format strings here: https://doc.rust-lang.org/std/fmt/
It would be more impressive if Rust had an actual type checked variadic print format function like C++ 23's std::println - nobody should be under any false impression about how impressive that feature is, but the type safe macro is effective, that'll do pig.
AFAIK C++ does format string parsing at runtime (as you need to provide a parser function), but maybe it can be made constexpr? I've not written C++ for over a year now so my fluency is fleeting a bit
IMO the market share for this profile is indeed small. Languages that trade off a part of that security for simplicity may have considerably more market share in the future, although on the other hand, the public may instead prefer a "fast-enough and dead-simple" language.
Some examples:
It defaulted to the fully backwards compatible version (vs 2021) which threw errors as I went through some recent example code.
(I think) I had to add a few lines to my cargo.toml so the compiler would not rebuild bevy every time I recompiled (when I only changed 1 line in my program).
"cargo init" and "cargo new" default to the 2021 edition, and have ever since it was stabilized: https://github.com/rust-lang/cargo/pull/9800
Either you accidentally installed a version of cargo from before the 2021 edition was stabilized, or you ran "cargo new --edition <something>", or you started by cloning an out of date project of some sort, in which case it's not really an issue with "defaults".
> smart defaults for the compiler. I was surprised how many lines I had to add to my cargo.toml
Normally this would be a pointlessly pedantic point, but cargo is not the compiler. This thread, the linked title blog post, they are about the rust compiler, not cargo. There's a close relationship, but cargo's defaults aren't necessarily related to what the "next rust compiler" might do.
I started my project without cargo at first and tried to start writing code without a cargo.toml. I was surprised that cairo didn't default to 2021 until I specified it in the .toml file. Good point that cargo init/new would have solved this!
I guess my point about the compiler was that it seems to rely on cargo.toml for many 'optimizations' that I would expect to be defaults. (Examples include the two i mentioned above).
But I'm new to the language and understand that most people will just use `cargo init` and google a few other common cargo.toml settings to improve compile times.
However, I don't think the solution is a better default, but rather the solution is time-travel.
Defaulting to the latest edition would mean that any rust library that predates editions would likely break when you imported it (since it would default to an edition that didn't exist when it was written, and editions are allowed to make breaking changes of that sort).
The thing that would fix your issue would be time-traveling back in time to when cargo was created, and making edition a required field of all cargo.toml files that results in an error until you add one. That would have saved you from any trouble.
Rust could also do a python2 -> python3 like transition, where crates from the old "edition not required" world can't be imported anymore at all, but that seems like a very small thing to cause so much ecosystem pain over.