Improving Rust compile times to enable adoption of memory safety
memorysafety.org
memorysafety.org
https://blog.rust-lang.org/inside-rust/2023/01/30/cargo-spar...
Before:
Download all metadata, Download xyz package
After:
Downolad xyz's metadata, Download xyz
They already know you are using xyz.
Edit: A thought occurs to me. Cargo downloads metadata from crates.io but clones the package repo from GitHub/etc. So unless I'm missing something, downloading specific metadata instead of all metadata allows for crates.io to track your specific packages in addition to GitHub.
crates.io crates are tarballs stored in S3. The tarball downloads also go through a download-counting service, which is how you get download stats for all crates (it's not a tracker in the Google-is-watching-you sense, but just an integer increment in Postgres).
Use https://lib.rs/cargo-crev or source view on docs.rs to see the actual source code that has been uploaded by Cargo.
If knowing you use a crate is too much, then running your own registry with a mirror of packages seems like all you could do.
If you use GitHub as like a storage server and totally externalize the costs of the package index onto them, then it's workable for free. But if you're running your own servers then it's a whole different ballgame.
"Location of where to place all generated artifacts, relative to the current working directory."
The problem is that if two workspace build a dependency with sightly different features or flags, it will always be rebuild when changing workspaces
$ git clone https://github.com/BurntSushi/ripgrep
$ cd ripgrep
$ git checkout 0.8.0
$ time cargo +1.20.0 build --release
real 34.367
user 1:07.36
sys 1.568
maxmem 520 MB
faults 1575
$ time cargo +1.67.0 build --release
[... snip sooooo many warnings, lol ...]
real 7.761
user 1:32.29
sys 4.489
maxmem 609 MB
faults 7503
As kryps pointed out on reddit, I believe at some point there was a change to add/improve compilation times by making more effective use of parallelism. So forcing the build to use a single thread produces more sobering results, but still a huge win: $ time cargo +1.20.0 build -j1 --release
real 1:03.11
user 1:01.90
sys 1.156
maxmem 518 MB
faults 0
$ time cargo +1.67.0 build -j1 --release
real 46.112
user 44.259
sys 1.930
maxmem 344 MB
faults 0
(My CPU is a i9-12900K.)These are from-scratch release builds, which probably matter less than incremental builds. But they still matter. This is just one barometer of many.
[1]: https://old.reddit.com/r/rust/comments/10s5nkq/improving_rus...
Total CPU time in user mode will normally increase when you add more threads, unless you're getting perfect or better-than-perfect scaling.
If I had to guess as to what's happening, it's that there's some thread pool, and at some point, near the end of compilation, only one or two of those threads is busy doing anything while the other threads are sitting and idling. Now whether and how that "idling" gets interpreted as "CPU being actively used in user mode" isn't quite clear to me. (It may not, in which case, my guess is bunk.)
Perhaps someone more familiar with what 'user' time actually means and how it interplays with multi-threaded programs will be able to chime in.
(I do not think faults have anything to do with it. The number of faults reported here is quite small, and if I re-run the build, the number can change quite a bit---including going to zero---and the overall time remains unaffected.)
Extremely parallel programs can improve on this, but it's perfectly normal to see 2x overhead for fine-grained parallelism.
But maybe for Cargo, parallelism is more fine grained than I think it is. Perhaps because of codegen-units. And similarly for ripgrep, if it's searching a lot of tiny files, that might result in fine grained parallelism in practice.
Which is fine; it’s still faster overall. Disable SMT and you’ll see much lower overhead, but higher time spent overall.
That's really it. That's the entire explanation. It's useful if and only if there are unused resources behind it, due to pipeline stalls or because the siblings are doing different things. It's virtually impossible to fully utilize a CPU core with a single thread; having two threads therefore boosts performance, but only to the degree that the first thread is incapable of using the whole thing.
That's why the speedup is around 20%, not 100%.
edit: misparsed, like corrected below, my bad.
I suspect the answer is: Perfect scaling doesn't happen on real CPUs.
Turboboost lets a single thread go to higher frequencies than a fully loaded CPU. So you would expect "sum of user times" to increase even if "sum of user clock cycles" is scaling perfectly.
Hyperthreading is the next issue: multiple threads are not running independently, but might be fighting for resources on a single CPU core.
In a pure number-crunching algorithm limited by functional units, this means using $(nproc) threads instead of 1 thread should be expected to more than double the user time based on these two first points alone!
Compilers of course are rarely limited by functional units: they do a decent bit of pointer-chasing, branching, etc. and are stalled a good bit of time. (While OS-level blocking doesn't count as user time; the OS isn't aware of these CPU-level stalls, so these count as user time!) This is what makes hyperthreading actually helpful.
But compilers also tend to be memory/cache-limited. L1 is shared between the hyperthreads, and other caches are shared between multiple/all cores. This means running multiple threads compiling different parts of the program in parallel means each thread of computation gets to work with a smaller portion of the cache -- the effective cache size is decreasing. That's another reason for the user time to go up.
And once you have a significant number of cache misses from a bunch of cores, you might be limited on memory bandwidth. At that point, also putting the last few remaining idle cores to work will not be able to speed up the real-time runtime anymore -- but it will make "user time" tick up faster.
In particularly unlucky combinations of working set size vs. cache size, adding another thread (bringing along another working set) may even increase the real time. Putting more cores to work isn't always good!
That said, compilers are more limited by memory/cache latency than bandwidth, so adding cores is usually pretty good. But it's not perfect scaling even if the compiler has "perfect parallellism" without any locks.
Ah yes, this is a good one! I did not account for this. Mental model updated.
Your other points are good too. I considered some of them as well, but maybe not enough in the context of competition making many things just a bit slower. Makes sense.
When you max out parallelism, you're using 1) hardware threads which "split" a physical core and (ideally) each run at a bit more than half the CPU's single-thread speed, and 2) the small "efficiency" cores on newer Intel and Apple chips. Also, single-threaded runs can feed a ton of watts to the one active core since it doesn't have to share much power/cooling budget with the others, letting it run at a higher clock rate.
All these tricks improve the throughput, or you wouldn't see that wall-time reduction and chipmakers wouldn't want to ship them, but they do increase how long it takes each thread to get a unit of work done in a very multithreaded context, which contributes to the total CPU time being higher than it is in a single-threaded run.
The HT cores aren't real CPU cores. They're just an opportunistic reuse of hardware cores when another thread is waiting for RAM (RAM is relatively so slow that they're waiting a lot, for a long time).
So code on the HT "core" doesn't run all the time, only when other thread is blocked. But the time HT threads wait for their opportunity turn is included in wall-clock time, and makes them look slow.
The end result was that doing WebSphere development actually got slower, because of their virtual nature and everything else on the CPU being shared.
So I ended up disabling it again to get the original performance back.
I assume OSes have since then developed proper support for scheduling and pre-empting hyperthreading. Also the gap between RAM and CPU speed only got worse, and CPUs have grown more various internal compute units, so there's even more idle hardware to throw HT threads at.
Perhaps they were on old builds or some massive projects?
ripgrep is probably on the smallish side. It's not hard to get a lot bigger than that and have those incremental times also get correspondingly bigger.
And complaining about compile times doesn't mean compile times haven't improved.
My personal project takes seconds to compile, but fair enough it's small, but even bigger projects like a game in Bevy don't take that much to compile. Minute or two tops. About 30 seconds when incremental.
People complained of 10x slower perf. Essentially 15min build times.
Fact that older versions might be slower to compile fills another part of the puzzle.
That and fact I have a 24 hyper thread monster of CPU.
I work on a large'ish C++ project and incremental is generally 1-2 seconds.
Incremental must work in release builds(someone else said it only works in debug for Rust), although it is fine to disable link time optimizations as those are obviously kinda slow.
I don't recall exact numbers. But bevy can pull a lot of depenencies. Enough for `target` directory to rival NPM worst offenders (e.g. ~1GB).
I truly just do not see what is difficult to understand here.
Also, people have vastly different work flows. Some people tend to slowly write a lot of code and compile rarely. Maybe they tend to have runtime tools to tweak things. Otherwise like to iterate really fast. Try a code change, see if the UI looks better or things run faster, and when you work like this even a compile time of 3 seconds can be a little bit annoying, and 30 seconds maddening.
Notably, serde can drive up compile times a lot, which is why miniserde still exists and gets some use.
diesel = { version = "*", features = ["128-column-tables"], ... }
> How many LoC there is in ripgrep? 46sec to build a grep like tool with a powerful CPU seems crazy.
I wrote out an answer before I knew the comment was deleted, so... I'll just post it as a reply to myself...
-----
Well it takes 46 seconds with only a single thread. It takes ~7 seconds with many threads. In the 0.8.0 checkout, if I run `cargo vendor` and then tokei, I get:
$ tokei -trust src/ vendor/
===============================================================================
Language Files Lines Code Comments Blanks
===============================================================================
Rust 765 299692 276218 10274 13200
|- Markdown 387 21647 2902 14886 3859
(Total) 321339 279120 25160 17059
===============================================================================
Total 765 299692 276218 10274 13200
===============================================================================
So that's about a quarter million lines. But this is very likely to be a poor representation of actual complexity. If I had to guess, I'd say the vast majority of those lines are some kind of auto-generated thing. (Like Unicode tables.) That count also includes tests. Just by excluding winapi, for example, the count goes down to ~150,000.If you only look at the code in the ripgrep repo (in the 0.8.0 checkout), then you get something like ~13K:
$ tokei -trust src globset grep ignore termcolor wincolor
===============================================================================
Language Files Lines Code Comments Blanks
===============================================================================
Rust 34 15484 13205 780 1499
|- Markdown 30 2300 6 1905 389
(Total) 17784 13211 2685 1888
===============================================================================
Total 34 15484 13205 780 1499
===============================================================================
It's probably also fair to count the regex engine too (version 0.2.6): $ tokei -trust src regex-syntax
===============================================================================
Language Files Lines Code Comments Blanks
===============================================================================
Rust 29 22745 18873 2225 1647
|- Markdown 23 3250 285 2399 566
(Total) 25995 19158 4624 2213
===============================================================================
Total 29 22745 18873 2225 1647
===============================================================================
Where about 5K of that are Unicode tables.So I don't know. Answering questions like this is actually a little tricky, and presumably you're looking for a barometer of how big the project is.
For comparison, GNU grep takes about 17s single threaded to build from scratch from its tarball:
$ time (./configure --prefix=/usr && make -j1)
real 17.639
user 9.948
sys 2.418
maxmem 77 MB
faults 31
Using `-j16` decreases the time to 14s, which is actually slower than a from scratch ripgrep 0.8.0 build. Primarily do to what appears to be a single threaded configure script for GNU grep.So I dunno what seems crazy to you here honestly. It's also worth pointing out that ripgrep has quite a bit more functionality than something like GNU grep, and that functionality comes with a fair bit of code. (Gitignore matching, transcoding and Unicode come to mind.)
Steam hardware survey, Jan 2017 [1] vs Jan 2023, "Physical CPUs (Windows)"
2017 2023
1 CPU 1.9% 0.2%
2 CPUs 45.8% 9.6%
3 CPUs 2.6% 0.4%
4 CPUs 47.8% 29.6%
6 CPUs 1.4% 33.0%
8 CPUs 0.2% 18.8%
More 0.3% 8.4%
[1] https://web.archive.org/web/20170225152808/https://store.ste...I suppose it's not the worst problem to have. Makes me realize how spoiled I got after multiple-core computers became the norm.
[0]: https://www.gnu.org/software/make/manual/html_node/Job-Slots...
[2]: https://github.com/rust-lang/cargo/issues/1744
[2]: https://github.com/rust-lang/rust/pull/42682
`nice cargo build` will run all threads at low priority, but this is generally a good idea if you want to prioritize interactive processes while running a build in the background.
See https://github.com/rust-lang/cargo/blob/master/CHANGELOG.md#...
Make it a shell script like `takeiteasy`, and run `takeiteasy cargo ...`
dude cargo ...
has a nice flow to it.SMT is a throughput thing, and I honestly turn it off on my workstation for that reason. It's great for cloud providers that want to charge you for a "vCPU" that can't use all of that core's features. Not amazing for a workstation where you want to chill out on YouTube while something CPU intensive happens in the background. (For a bazel C++ build, having SMT on, on a Threadripper 3970X, does increase performance by 15%. But at the cost of using ~100GB of RAM at peak! I have 128GB, so no big deal, but SMT can be pretty expensive. It's probably not worth it for most workloads. 32 cores builds my Go projects quickly enough, and if I have to build C++ code, well, I wait. ;)
It's beyond frustrating that any "i+=1" change requires relinking a 50mb binary from scratch and rebuilding a good chunk of the Win32 crate for good measure. Until such enterprise features become available, high developer productivity in Rust remains elusive.
I don't think it's enabled by default in release builds (because it might sacrifice perf too much?) and it doesn't make linking incremental.
Making the entire pipeline incremental, including release builds, probably requires some very fundamental changes to how our compilers function. I think Cranelift is making inroads in this direction by caching the results of compiling individual functions, but I know very little about it and might even be describing it incorrectly here in this comment.
It’s especially hard to solve this with a language like rust, but I agree!
I’ve long wanted to experiment with a compiler architecture which could do fully incremental compilation, maybe down the function in granularity. In the linked (debug) executable, use a malloc style library to manage disk space. When a function changes, recompile it, free the old copy in the binary, allocate space for the new function and update jump addresses. You’d need to cache a whole lot of the compiler’s context between invocations - but honestly that should be doable with a little database like LMDB. Or alternately, we could run our compiler in “interactive mode”, and leave all the type information and everything else resident in memory between compilation runs. When the compiler notices some functions are changed, it flushes the old function definitions, compiles the new functions and updates everything just like when the DOM updates and needs to recompute layout and styles.
A well optimized incremental compiler should be able to do a “i += 1” line change faster than my monitor’s refresh rate. It’s crazy we still design compilers to do a mountain of processing work, generate a huge amount of state and then when they’re done throw all that work out. Next time we run the compiler, we redo all of that work again. And the work is all almost identical.
Unfortunately this would be a particularly difficult change to make in the rust compiler. Might want to experiment with a simpler language first to figure out the architecture and the fully incremental linker. It would be a super fun project though!
https://www.youtube.com/watch?v=yLZwLSzkH3E
VC++ has similar kind of support nowadays.
Are you really running your test suite for every "i+=1" change on other languages?
You don't have to run your testsuite for a small bugfix (that's what CI is for), but you DO need to restart, reset the testcase that triggers the code you are interested in, and step through it again. Rinse and repeat for 20 or so times, with various data etc. - at least that's my debug-heavy workflow. If any trivial recompile takes a minute or so, that's a frustrating time spent waiting as opposed to using something like a dynamic language to accomplish the same task.
So you would instinctively avoid Rust for any task that can be accomplished with Python or JS, a real shame since it's very close to being an universal language.
GET OFF MY LAWN
For example Java has a perfectly serviceable TLS stack written entirely in a memory safe language. Although you could try to make OpenSSL memory safe by rewriting it in Rust - which realistically means yet another fork not many people use - you could also do the same thing by implementing the OpenSSL API on top of JSSE and Bouncy Castle. The GraalVM native image project allows you to export Java symbols as C APIs and to compile libraries to standalone native code, so this is technically feasible now.
There's also some other approaches. GraalVM can also run many C/C++ programs in a way that makes them automatically memory safe, by JIT compiling LLVM bitcode and replacing allocation/free calls with garbage collected allocations. Pointer dereferences are also replaced with safe member accesses. It works as long as the C is fairly strictly C compliant and doesn't rely on undefined behavior. This functionality is unfortunately an enterprise feature but the core LLVM execution engine is open source, so if you're at the level of major upgrades to Rust you could also reimplement the memory safety aspect on top of the open source code. Then again you can compile the result down to a shared native library that doesn't rely on any external JVM.
Don't get me wrong, I'm not saying don't improve Rust compile times. Faster Rust compiles would be great. I'm just pointing out that, well, it's not the only memory safe language in the world, and actually using a GC isn't a major problem these days for many real world tasks that are still done with C.
Or just write a better crypto stack without the many legacy constraints holding OpenSSL back. Rustls (https://github.com/rustls/rustls) does that. It has also been audited and found to be excellent - report (https://github.com/rustls/rustls/blob/main/audit/TLS-01-repo...).
You're suggesting writing this stack in a GC language. That's possible, except most people looking for an OpenSSL solution probably won't be willing to take the hit of slower run time perf and possible GC pauses (even if these might be small in practice). Also, these are hypothetical for now. Rustls exists today.
Point is that the code already exists - it's not hypothetical - and has done for a long time. It is far easier to write bindings from an existing C API to a managed implementation than write, audit and maintain a whole new stack from scratch. There are also many other cases where apps could feasibly be replaced with code written in managed languages and then invoked from C or C++.
Anything written in C/C++ can certainly tolerate pauses when calling into third party libraries because malloc/free can pause for long periods, libraries are allowed to do IO without even documenting that fact etc.
I think it's fair to be concerned that rewrite-it-in-rust is becoming a myopic obsession for security people. That's one way to improve memory safety but by no means the only one. There are so many cases where you don't need to do that and you'll get results faster by not doing so, but it's not being considered for handwavy reasons.
I’d agree, if rustls wasn’t already written, audited and maintained. And there are other examples as well. The internationalisation libraries Icu4c and Icu4j exist, but the multi-language, cross-platform library Icu4x is written in Rust. Read the announcement post on the Unicode blog (http://blog.unicode.org/2022/09/announcing-icu4x-10.html?m=1) - security is only one of the reasons they chose to write it in Rust. Binary size, memory usage, high performance. Also compiles to wasm.
Your comment implies that people rewrite in Rust for security alone. But there are so many other benefits to doing so.
It’s much more fun to rewrite something in a new language than maintain bindings to some external language. You could wrap a Java library with a rust crate, but it would depend on Java and rust both being installed and sane on every operating system. Maintaining something like that would be painful. Users would constantly run into problems with Java not being installed correctly on macos, or an old version of Java on Debian breaking your crate in weird ways. It’s much more pleasant to just have a rust crate that runs everywhere rust runs, where all of the dependencies are installed with cargo.
Golang users would?
That aside excellent points about rust tls, and libssl legacy cruft.
To include a library written in another language as a shared lib, it needs to be C, C++ or Rust.
The other thing to consider is that in many applications, nearly every single bit of i/o will flow through a buffer and cryptographic function to encrypt/decrypt/validate it. This is the place where squeezing out every ounce of performance is critical. A JIT + GC might cost a lot more money than memory safety bugs + AOT optimized compilation.
Using a GC doesn't mean never reusing buffers, and Java has intrinsics for hardware accelerated cryptography for a long time. There's no reason performance has to be less, especially if you're willing to fund full time research projects to optimize it.
The belief that performance is more important than everything else is exactly how we ended up with pervasive memory safety vulns to begin with. Rust doesn't make it free, as you pay in developer hours.
Rust interops much better with C than most languages, including Java, and offers a smoother transition path to increased safety.
This fallacy is why these languages are still in use. Time after time, designers of all the safer languages were deciding that GC makes everything so much easier, and it's perfectly fine for overwhelming majority of programs. This is correct and rational if the goal is to get many people use the language, but a total self-own if the goal is to replace C and C++ entirely.
This dodging of the low-level memory management problem was consistently avoiding exactly the types of programs that people felt they had to use C or C++ for. The easy majority of programs that don't need C is already well served, but the tough cases that needed C were left uncontested.
From Rust's first introduction:
http://venge.net/graydon/talks/intro-talk-2.pdf
> Go seems to be barking up a different tree?
> Everyone is dodging the niche I'm interested in
I don't understand this. The vast majority (I would guess 95%+) of people using Rust have CPUs with AVX2 or NEON. Why is that a good reason? Why can't there be a fast path and slow path as a failover?
Some C and C++ compilers offer this and it requires some infrastructure to make it happen (simd attribute in GCC), or explicitly loading different kinds of dynamic libraries.
Although maybe GL 4.1 would do it nowadays.
The default configuration for the "x86-64" target is really meant to run on every x86-64 CPU. I know at lest Debian has the same policy. SSE2 is as old as x86-64 itself, so it can be assumed to be available in every x86-64 CPU, but nothing else can.
There has been some movement to define pseudo-targets like x86-64-v1, x86-64-v2, x86-64-v3 for higher baselines.
Without that you can't really do distributed and cached compilation 100% reliably.
Naturally somehow has to spend time analysing how their way maps into Rust compilation story.
For monomorphized code the compiler just needs a mode where it automatically does what the Momo crate does.
For proc macros the Watt crate (precompiled WASM macros) will make a big difference. It just needs official sanction and integration.
Anyway yeah those are totally separate problems to caching and distributed builds.
I wish there were some efforts at dramatically different approaches like this because there’s all this work going into compilation but it’s unlikely to make the development cycle twice as fast in most cases.
It’s not really a mix, you do one or the other although certainly ghci for local dev and compilation for prod.
I can see how someone would come to rust, type `cargo run`, wait 3-5 minutes while cargo downloads all the dependencies and compiles them along with the main package, and then say, "well that took awhile it kinda sucks". But if they change a few lines in the actual project and compile again it would be near instant.
The fair comparison would be something akin to deleting your node or go modules and running a cold build. I am slightly suspicious, not in a deliberate foul play way but more in a messy semantics and ad-hoc anecdotes way, that many of these compile time discrepancies probably boil down more to differences in how the cargo tooling handles dependencies and what it decides to include in the compile phase, where it decides to store caches and what that means for `clean`, etc. compared to similar package management tooling from other languages, than it does to "rustc is slow". But I could be wrong.
If it’s a big project and the lines you are changing are in something that is being used many other places then the rebuild will still take a little while. (30 seconds or a minute, or more, depending on the size of the project.)
Likewise, if you work on things in different branches you may need to wait more when you switch branch and work on something there.
Also if you switch between Rust versions you need to wait a while when you rebuild your project.
I love Rust, and I welcome everything that is being done to bring the compile times down further!
I'm working on a largeish modern java project using gradle, and this sounds great... Every time I start my server it takes 40 seconds just for gradle to find out that all the sub projects are up to date, nothing has been changed and no compilation is necessary...
Neither are especially fast though.
Having written Rust professionally for a number of years, this didn't happen too much. Where it did it was stuff like "yeah you need to Box the thing today", which... did not matter, we just did that and moved on.
> It can ever so slightly give one the impression that the rust community has decided that the language is mature and the only thing missing is faster compile times.
That is generally my feeling about Rust. There are a few areas where I'd like to see things get wrapped up (async traits, which are being actively worked on) but otherwise everything feels like a bonus. In terms of things that made Rust difficult to use, yeah, compile times were probably the number one.
- It's impossible to describe a type that "implements trait A and may or may not implement trait B"
- It's impossible to be generic over a trait (not a type that implements a trait, the trait itself)
How is this different from just describing a type that only "implements trait A" ?
So, specialization? Or something else? I haven't found a need for specialization. I remember when I came from C++ I had a hard time adjusting to "no specialization, no variadics" but idk I haven't missed it in years.
> - It's impossible to be generic over a trait (not a type that implements a trait, the trait itself)
Not sure I understand.
Basically yes. But that works with dynamic dispatch (trait objects) as well as static dispatch (generics).
> Not sure I understand.
A specific pattern I'd like to be able to represent is:
trait AlgorithmAInputData {
...
}
trait AlgorithmA {
trait InputData = AlgorithmAInputData;
...
}
trait DataStorage<trait AlgorithmA> {
type InputData : Algorithm::InputData;
fn get_input_data() -> InputData;
}
fn compute_algorithm_a<Storage: DataStorage<AlgorithmA>>() {
...
} let block = || {
let my_a = a.clone();
let my_b = b.clone();
let my_c = c.clone();
async move {
// use my_a, my_b, my_c
let value = ...
Ok<success::Type, error::Type>(value)
}
}
And you can't use `if let ... && let ...` (two lets for one if) because it doesn't desugar correctly.And error handling and backtraces are a beautiful mess. Your signatures look like `Result<..., Box<dyn std::error::Error>>` unless you use `anyhow::Result` but then half the stuff implements std::error::Error but not Into<anyhow::Error> and you can't add the silly trait impl because of language limitations so you have to map_err everywhere.
It's not just "oh throw a box around it and you're good". It's ideas that were introduced to the language when there was lots of steam ultimately not making it to a fully polished state (maybe Moz layoffs are partly to blame IDK). Anyway I love Rust and we use it in production and have been for years, but I think there's still quite a bit to polish.
I am clearly missing some context because that code is needlessly complex. Are you just trying to show that you have to clone values that you hold a reference to if you want to move them? Because yes, you do. But your example also needlessly borrows them in the outer closure.
Is not the case, is that the features are now good enough and compile times is the one major, big, sore point.
So, if you compare Rust to X you can make a very good case until you hit:
"... wait, Rust is THAT SLOW TO COMPILE?"
":(. Yes"
FWIW cold builds (i.e., in docker with no cache) of cargo are much slower than go, hanging for a long time on refreshing cargo.io indexes. I don't know exactly what that is doing but I have a feeling it is implemented in a monolithic way rather than on-demand. Rust has had plenty of time to make this better but it is still very slow for cold cargo builds, often spending minutes refreshing the crates index. But Go misses easy optimizations like creating strings from a byte slice.
So it is what it is - Go makes explicit promises of fast compile times. Thanks to that, build scripts in go are pretty fast. Any language that doesn't make that explicit might be slow to compile and might run fast - that's totally fine and I would rather have two languages optimized to each case than one mediocre language.
> We were able to benchmark bjorn3's cranelift codegen backend on full crates as well as on the build dependencies specifically (since they're also built for cargo check builds, and are always built without optimizations): there were no issues, and it performed impressively. It's well on its way to becoming a viable alternative to the LLVM backend for debug builds.
And the Cranelift codegen backend itself is also clear about it not being ready yet: https://github.com/bjorn3/rustc_codegen_cranelift
(To be clear, I am super excited about using Cranelift for debug builds. I just want to clarify that it isn't actually used by default yet.)
1. Clean builds can happen more often than some may think. CI/CD pipelines can end up with a lot of clean builds - especially if you use ephemeral instances (to save money), but even if you don't it's very likely.
Even locally it can happen sometimes. For example, we used Docker to run builds. For various reasons the cache could get blown. Also, sometimes weird systemy things happen and 'cargo clean' fixes it, but you have to recompile from scratch. This can take 10+ minutes on a decent sized codebase.
2. On a large codebase even small changes can lead to long recompile times, especially if you want to run tests - cargo check won't be enough, you need to build.
The incremental build performance seems to be really dependent on single-thread performance. An incremental build on a 2014ish Haswell e5-2660v3 xeon takes ~30s.
`cargo test` and default `cargo build` use the same profile, so presumably the first number is referring to `cargo build --release`. Release builds deliberately forego compilation speed in favor of optimization. In practice, most of my development involves `cargo check`, which is much faster than `cargo build`.
Both numbers are for debug builds. I don't know why `cargo test` is faster but I appreciate it.
Incremental release builds with `cargo build --release` are even slower taking ~35s on the 1240p.
Also it is quite irritating sometimes seeing the same crate being compiled multiple times as it gets referenced from other crates.
Ideally Rust could use a dumb compilation mode (or interpreter) for change-compile-debug cycles, and proper compilation for release, e.g. Haskell and OCaml offer such capabilities on their toolchains.
HTML+Js projects used to be testable with the load of a web page and that community has opt-in to long build times.
Most people are so far away from flow-state that they can't even imagine another way of being.
My machine is 7yo. People tell me to buy a new one just for compiling a Rust project... That's ecologically questionable.
Package management, one of Rust’s biggest strengths, is one of its biggest weaknesses here. It’s so easy to pull in another crate to do almost anything you want. How many of them are well-written, optimized, trustworthy, etc.? My guess is, not that many. That leads to applications that use them being bloated and inefficient. Hopefully, as the ecosystem matures, people will pay better attention to this.
Rust has a culture of splitting dependencies into small packages. This helps pull in only focused, tailored functionality that you need rather than depending on multi-purpose large monoliths. Ahead-of-time compilation + generics + LTO means there's no extra overhead to using code from 3rd party dependency vs your own (unlike interpreted or VM languages where loading code costs, or C with dynamic libraries where you depend on the whole library no matter how little you use from it).
I assume people scarred by low-quailty dependencies have been burned by npm. Unlike JS, Rust has a strong type system, with rules that make it hard to cut corners and break things. Rust also ships with a good linter, built-in unit testing, and standard documentation generator. These features raise the quality of average code.
Use of dependencies can improve efficiency of the whole application. Shared dependencies-of-dependencies increase code reuse, instead of each library rolling its own NIH basics like loggers or base64 decode, you can have one shared copy.
You can also easily use very optimized implementations of common tasks like JSON, hashmaps, regexes, cryptography, or channels. Rust has some world-class crates for these tasks.
The extract function tool is very buggy. As I spend a lot of time refactoring, maybe putting time in those tools would have a better ROI than so much work into making the compiler faster.
By the way AST manipulation is easy, the really hard part of refactoring (that I had a lot of problem with) is creating the lifetime annotations, which requires a deep understanding of the type system.
I was trying to learn some type theory and read papers to understand how Rust's life times work, but only found long research papers that don't even do the same thing as Rust.
I haven't found any documentation that documents exactly when a function call is accepted by the lifetime checker (borrow checking is easy).
there's this part though:
// NOTE: `'a: {` and `&'b x` is not valid syntax!
I hate that I can't introduce new lifetime inside a function, it would make refactoring so much easier. Right now I have to try to refactor, see if the compiler accepts it or not, then revert the change.
Sometimes desugaring would be a great feature in itself, sugaring makes interactions between functions much harder to understand.
Or a mode where it compile automatically every time you change a line? (With absolutely no optimization like inlining stuff etc to make it fast) kind of like just compiling the new line from Rust to its ASM equivalent and adding that to the rest of the compiles code. Like a big fatjar type of way, if that make sense.
If the concern is that I could change something in a crate, then could a checksum be created on the first compilation, then checked on future compilations, and if it matches then the crate doesn't need to be recompiled.
> If the concern is that I could change something in a crate
It's possible for a change in one crate to require recompiling its dependencies and transitive dependencies, due to conditional compilation (aka, "features' [0]). Basically you can't know which thing to compile until it's referenced by a dependent and provided a feature set.
That said, many crates don't have features and have a default feature set, but the number of variants to precompile is still quite large.
[0] https://doc.rust-lang.org/cargo/reference/features.html
Note that C and C++ have the exact same problem, but it's mitigated by people never giving a shit about locking dependencies and living with the horrible bugs that result from it.
- Feature flags
- Platform conditionals
- The specific rust version being used
- (unsure on this) the above for all dependencies of what is being pre-compiled
There is also the impediments of designing / agreeing on a security model (do you trust the author like PyPI, trust a central build authority, etc) and then funding the continued hosting.
Compiling on demand like in sccache is likely the best route for not over-building and being able to evict unused items.
The Rust dev team can't win!
Since it is impossible to mentally model them as the number of humans they are, I find it helpful to model them as at least a few very distinct individuals, or sometimes just as an amorphous philosophical gas that will expand to fill all available comments, where the only question is really with what distribution rather than whether a given point will be occupied.
I see all of these complaints — both the volume, and the count — as great indicators of Rust's total health.
Part of the problem is that news aggregates reward people who comment early, and the earliest comments are the kneejerk reactions where you braindump thoughts you've had brewing but don't have anywhere to put. (Probably without actually clicking through.)
Still, I believe what makes HN unusually nice is just the stellar moderation. It is definitely imperfect, but it creates a nice atmosphere that I think ultimately does encourage people to try to be civil, even though places like these definitely have a tendency to bring out the worst in people. Having a deft touch with moderation is very hard nowadays, especially with increasingly difficult demands put against moderators and absolutely every single possible subject matter turning into a miniature culture war (how in the hell do you turn the discussion of gas ranges vs electric ranges into a culture war?!) and the unmoderated hellscapes of the Internet wrongly painting all lightweight moderation with a black mark.
I definitely fear for the future of communities like HN, because the pressure from increasingly vile malicious actors as well as the counter-active pressure from others to moderate harder, stronger, faster will eventually break the sustainability of this sort of community. When I first joined HN, a lot of communities on the Internet felt like this. Now, I know of very few.
I find Rust extremely difficult and slow to work with. It takes me a long, long time to get anything working at all. Easily 5x more than any other language I've used (and that's many, and a good handful professionally), and I've been learning Rust for over a year.
Figures are hard to come by of course, but anecdotally I'm the only person in my circle who's continued using Rust. All the others have dropped out because they just find it too hard to get anything done in. Not sure if this is true, but I heard on a podcast the other day that one of the big surveys showed Rust to have the biggest learning drop-out rate of any mainstream programming language. That wouldn't surprise me, and comports well with Rust being the 'most loved' (people tend to love skills they have gained with much effort!).
I've just scrubbed back through - the podcast was https://syntax.fm/show/571/supper-club-rust-in-action-with-t..., which follows the pleasing practise of providing chapters (yay). It was Tim McNamara speaking from about 12'40": what he actually said was that about half of Rust learners who do drop out fall away because of the purported difficulty of the language. Very different from my faux summary so apologies for the grievous misrepresentation.
I still find Rust as hard to use as others I've spoken to do though. I'm kind of keeping at it for reasons specific to projects I have in mind, as well as a certain dense stubbornness.
Is there anything there someone like me (ie. having difficulty with Rust) might usefully contribute to? I'm a little overwhelmed right now trying to keep a roof over my head, but am compiling a list of things I might like to help with when the current storm has passed.
Personally I think contributions work best when you're trying to solve a pain that you personally have or at least have some sort of connection to, so I'd encourage you to consider what/how/why you struggled to learn, and then try to fix that. I know it's vague, but it's the best I've got right now!