Glad to see this progress.
Glad to see this progress.
Rust’s unique feature itself fundamentally depends on extensive static analysis. It’s not a design choice, it is pretty much what Rust is - a low-level language without a GC that is still memory safe. The price for that is hefty compile times.
Check out https://github.com/lqd/rustc-benchmarking-data/tree/main/res... and the other benchmarks in that repository for some data on how real world crates compilation times are spent. You'll find that backend code generation and optimization dominate most crates compile times. There are a few exceptions: particularly macro heavy crates, a couple crates with deeply nested types that hit some quadratic behavior in the compiler. But overall, the backend is still the largest piece.
I think Zig is correct in having different modes.
What you would need to go faster would be not only a non-monomorphizing compiler but also boxed types. That would be a very different language, one higher-level than even Go (which monomorphizes generics).
https://github.com/golang/proposal/blob/master/design/generi...
One of the explicit goals, by Go's creators, was fast build times. I still remember Rob Pike introducing Go during an all-hands at Google, where he talked about the very long build times for C++ and Java in Google's monorepo, and then showed some promising demos. (Most of us rolled our eyes at it then, because it was just a "hello world", but it's quite impressive how the language has evolved and remained true to its goals.)
> - it is just a plain language where the compiler can just spit out vaguely optimized code, and call it a day.
It's a simple language, but I wouldn't call it plain, nor characterize the optimizers that way.
Also, as can be seen, go is not a well-designed language, having language warts we knew for 50 years. I would take the creators’ claims with a huge grain of salt.
why do you think so? Their goal was to be much faster, but I can't find much benchmarks..
Unless you meant that Java's AOT compilation is faster than Go's?
Also, there are single-pass compilers that produce machine code, they are not fundamentally slower than a byte code generator. Of course extensive optimizations will be more expensive.
1. The compiler could ship binary artifacts, which would avoid all compilation of build scripts/ proc macros, and allow those to be compiled with performance optimizations enabled. This would be huge on its own.
2. Cranelift could potentially improve backend codegen compile times significantly as well.
3. Link times are still suboptimal, mold is promising here.
We can definitely still get significant wins out of the compiler.
Pretty sure compile times can get cut in half (or better) with those changes.
Also, mold was designed as an alternative to gold / lld, therefore it would require to be open-source and free on their main platform: linux.
It's AGPL on Linux now, and they sell commercial licenses for companies that won't touch that license, and they were contemplating earlier making mold only available under a non-free source available license like BSL, so there's no "requirement" as such that it be free and open source, even on Linux.
Do you have any data on this ? Maybe that's industry dependent, but I hardly know any Windows (not even talking about macOS, that's almost nil) developers outside of video games and web dev. 100% of Rust devs I know use Linux, to keep on the subject.
This is not a knock on Rust—I doubt it’s possible to do what Rust does—including zero overhead abstractions—in a fast compiling language. Go certainly pays a performance penalty with things like boxed generics.
But sure, twice as fast isn't fast, it's just faster. My point is that we're not at the point of serious diminishing returns, there's tons of stuff left to do.
It is a knock on Rust. The circumstances of Rust's state of existence in 2023, as a language created in this millennium but not in the last decade, are absurd.
> I doubt it’s possible to do what Rust does—including zero overhead abstractions—in a fast compiling language
People packaging releases for software written in Rust and others who are passive consumers and finding themselves downloading some project repo to compile from source for whatever reason (e.g. because the creators don't do binary releases themselves) don't need the Rustlang toolchain to do the things that active contributors to a given project (who want type system diagnostics, etc.) need from it.
I'd call this oversight a massive lack of imagination on the part of TPTB, but that would be wrong, because there is no need to imagine the differences between these use cases. They exist. An adequate toolchain for dealing with projects written in Rust—despite the deliberate decisions made during language design that led to these problems—does not.
Sorry, could you be more vague?
Anyway, your post seems to be about packaging software? Or something? Confusing since that has nothing to do with the language...
Yes, I can, since "2023", "this millennium", and "in the last decade" are all concrete, well-defined things.
I've been told that the Cranelift team (at least for the time being) doesn't have the intention on focusing on the optimizer to a degree where it would be competitive with LLVM's optimizers (which would also be a huge effort). So if you want faster compile times you would have to take significant performance hits (which for a lot of code compiled in CI is not a trade-off that people are willing to take).
(1) and (3) are still very significant, fwiw.
That being said, I definitely find the trade-off worth it. Though, I've never been the kind of programmer that desires the constant iteration and feedback of something like "REPL driven development".
Rust has a massive advantage, which is having a 'sanctioned' package manager and built-time capabilities. A huge part of Rust's slowdown is due to:
a) Having to compile build scripts
b) Those build scripts being built without optimizations (100s of times slower at runtime)
If cargo + crates.io supports pre-built dependencies that is a massive optimization.
This isn't theoretical or optimistic, it's just a fact - we can already see this by compiling build and proc macro crates with optimizations, it's just not the default and they still have to be compiled once. IF you remove that compilation time, again, it's not theoretical, it's turning N time spent on those deps into 0 time spent.
There is easily a 200% performance win available, just from the known optimizations that are on the table.
I'm hopeful something like watt (https://github.com/dtolnay/watt) will land in Cargo that'll allow us to ship pre-compiled wasm blobs for proc-macros so we can just have sandboxed binaries.
When you export a generic function in C++, every file that pulls it in has to re-parse it, and every instantiation has to re-type-check it. C++20 modules should help with the first part, but they can't help with the second (and neither can concepts). Further, separate translation units can wind up duplicating the same instantiations, which the linker has to deduplicate.
When you export a generic function in Rust, by the time it gets pulled in somewhere else, it takes the form of pre-parsed, pre-type-checked MIR. It can also be pre-optimized, so type-independent optimization work is shared between instantiations. The compiler can also tell, before instantiation, which type parameters a function does not actually depend on, and essentially erase them ("polymorphization"). Further, Rust's compilation model reduces the redundant duplicate instantiations C++ does, both by using larger translation units and by automatically sharing any instantiations in dependencies with their dependents (though you can do this by hand in C++).
(Incidentally, these differences also apply to inline functions- in C++ you wind up putting their definitions in headers and recompiling them from scratch over and over; in Rust they are shared MIR form.)
Rust makes the problems easier to fix, IMHO. So, maybe even with same (or slightly longer) compile times, you'll hopefully have faster time to delivery.
In fact, in my experience, Rust has faster time to delivery than any other language I've used. It takes forever to compile, but I have so many fewer runtime bugs that have to be caught (hopefully) by testing, that it still comes out ahead, overall (again, for me and my various projects).
I also find write-time to not be as slow as others complain about, except when it comes to async/futures where it is, indeed, pretty rough. But, if I sit and think about how many times I have to flip back and forth between my code and some library code to try and guess what exceptions it may or may not throw in other languages or whether something could be null or not, I find that the dev times aren't so much better in these other languages as people sometimes claim.
Sure, if you're a fulltime JavaScript dev with 10 years of experience, you might remember things like that calling the Array constructor with 0 or >1 arguments creates an array with those values, but if you call it with exactly 1 number, it will create an empty array with that capacity. But, since I have to switch between many languages regularly, my time to delivery is significantly reduced by nonsense like that. Likewise, it's reduced by NPEs in Java, double-frees in C++, Kotlin's inane idea to use exceptions for errors and coroutine control-flow, etc, etc.
The fact that my only complaint is that compile times are slower than I'd like should be seen as high praise.
This is a HN thread about a blog post about how compile times have become dramatically better thanks to newly introduced parallelism in an area that was completely single threaded.
> However, at this point the compiler has been heavily optimized and new improvements are hard to find. There is no low-hanging fruit remaining. But there is one piece of large but high-hanging fruit: parallelism.
From discussions I've seen, there's not much high-hanging fruit left either, short of rewriting the entire compiler for better incremental compilation.
Incremental matters more than clean build times because (A) you're likely to do a lot more of them (B) they break developer flow more than waiting on CI does (C) at least in theory, you can always add more cores to your CI and get reasonable speedups, less so for incremental.
Why not? If I add a new struct with `#[derive(serde::Serialize)]` I'll benefit from serde being compiled with optimizations.
> they break developer flow more than waiting on CI does
Eh, depends.
It might not get 10x better, but 3x isn't outside the realm of possibility. Just swapping the LLVM backend for cranelift can cut compile times in half.
The low-hanging fruit is gone but there are lots of hard but likely-significant improvements left on the table.
Also Java style of tiny classes and tiny files.
It’s an issue on both the implementation side and the thing-being-implemented
I would have thought Rust would be better on both fronts.
How many lines are the Rust codebases and their dependencies?
To balance / explain my point:
- my day work is Java / Kotlin with Gradle. Now, we can talk about glacial compilation times
- on my open source Rust project, we try to minimise dependencies, don't use macros (apart derive[Debug, Clone] etc...), and have a very moderate generics usage
If you take the time to `cargo build` my project, I'll be happy to have feedbacks on compilation times
I just did a clean build `cargo build`, 19 minutes 44 seconds.
I added 1 line (`dbg!("foo")`) and it took 14.76s
- cloc shows that there are ~70,000 lines of Rust
- with `cargo tree`, I see that the project depends on ~600 crates
In my toy project:
- cloc shows that there are ~40,000 lines of Rust
- with `cargo tree`, I see ~40 crates
I don't know the scope of grapl, but 600 (transitives) crates seems a lot to me. Maybe that explains why this particular build is so long. I haven't managed to build it (seems to have prerequisites on proto buffer stuff).
Naturally more crates means more time on compilation. Grapl is a pretty large project, lots of services that do different things, so it isn't too surprising that it has a lot of dependencies relative to what I assume is a more tightly scoped project.
For example, Grapl talks to multiple different databases, AWS services, speaks HTTP + JSON and gRPC (with protobuf), has a cli, etc etc.
Optimized: 1m 25s
cargo build 42.55s user 4.35s system 748% cpu 6.269 total
cargo build --release 99.38s user 4.13s system 267% cpu 38.660 total
(I can really recommend this box for Rust. I got it in August, should have waited for M3...)
Kudos for making hurl, it looks super-userful!
From a quick tokei, looks like 49kloc of Rust.
When using local disks, a 10k LoC project takes ~5 seconds to build
When using network FS, the same project takes ~23 seconds to build
strip = "debuginfo"
This was a ~20x reduction in size, 279M -> 19M. If you have slow storage it could really help performance (don't recall, personally).FWIW it does break debuggers since you won't have the symbols. Just comment it out when you need a build that has all of that info.
Don't hesitate to be brutally honest — nothing can shock me as my go-to compiler is GHC.
Here's the code: https://github.com/grapl-security/grapl/tree/main/src/rust
I seem to remember that caching was really, really bad.
Finished dev [optimized + debuginfo] target(s) in 16m 58s
Neat.What improvements are you expecting?
async-stripe takes over two minutes to build due to codegen. We're considering switching to dolladollabills.
Our core API server takes a minute to build, and we have about a dozen services and command line apps, a bunch of little shared library crates, two desktop apps, and a Bevy app.
Our Github actions docker build takes ~10 minutes if you don't include the tests, but we're starting to shave off more time. (Our monorepo is 105589 Rust LOC total)
I thought this was a joke. That has to be one of the best library names I've ever seen.
Ooohh interesting. We also use async-stripe, definitely going to have to check out dolladollabills though. Also in the Rust monorepo camp, our proof release takes ~5 mins from clean, tests are about 6 mins. We’ve invested a bit of effort getting our build-time down: don’t build in a docker container, we just copy the final artefact in-this wiped the most time of our builds. More parallel codegen units too.