I Improved My Rust Compile Times
benw.is
benw.is
Yes, hot module reloading is nice and quick. But it reloads with errors included.
The primary benefit of the rust compiler, at least from my point of view, is telling me what's wrong, where. It's a worthwhile sacrifice for a few seconds of my time, when at the end of addressing all of the obvious errors, I have something that works. I find this miles better than HMR.
With that said, Elm is super fast to compile too, which is very nice.
If only devs with beefy machines are able to use Rust properly, that hinders adoption.
While I'm not a developer of any sort, you couldn't be more on the nose here. I waited and waited for a Surface Laptop that has 32gb on ebay. When it came, price skyrocket to almost double the 16gb model. Twice more it happened and I literally gave up and bought the 16gb. I don't know what the retail premium from Microsoft is, but the reseller premium is truly insanity. Surely, Microsoft spent the time engineering it to fit 32gb and could reap more producing more of them.
It's an absolute beast, build quality is almost on par with Apple (screen is better than what Apple offered in 2022, actually) and the 32GB RAM base model from this year can be had for less than 1,000€.
I dunno where the "destruction of your wallet" threshold lies for you but for me it's way higher.
The 2022 version also used Ryzen APUs and there is excellent Linux support. It's my daily driver for Rust development (I run Pop!_OS on it).
Point not comprehended.
There is a bit of a qualitative distinction between the Lenovo backdoor stuff and that. But that’s just my personal threat model
Stop buying macbooks if you want lots of RAM and performance for a low price.
Incredible! Rust developers have invented "IDEs"!
And so won't Rust, a runtime error is a runtime error after all.
But even Typescript has pretty good control flow analysis which catches a lot of (what would be) JS runtime errors while typing when properly integrated with an IDE.
It's not like Rust invented static typing, it just has a very strong type system (which can quickly become an annoyance though, that's why there's a wide range of 'strictness' among statically typed languages, because while static typing is generally a good thing, a 'too strong' type system might not be.
Apparently they invented static linking, pattern matching, structural typing, unsafe keyword, and plenty of other stuff already present on compiler toolchains and type systems that predate them...
I wonder if we will ever sort out the issue of developers caring about programing language history, or operating systems for that matter.
I mean in the wider sense, not the geeky community that already does that as a hobby.
TFA talks about reload times of 20-30s before their optimization, on a powerful computer. It’s certainly not “a few seconds” especially for people on less powerful hardware. That’s a ridiculous wait time for, say, tweaking a CSS class.
I find this comment highly suspect.
With rust, it's very refreshing in that it's idiot proof, which JavaScript is not. CSS is not relevant, since it is not actual programming.
I don’t need a lecture on the history of JavaScript, thank you. Especially since JavaScript/TypeScript has come such a long way that JavaScript the Good Parts is barely relevant anymore.
> Rust… idiot proof
That’s such a… interesting statement that I’m not going to respond further.
Another, very real "programming" situation could be the following: correctly parsing 3D mesh information in a multithreaded fashion and pass that to your preferred API of choice. Is one going to get all edge cases right the first time? Is your program's codebase large enough? Eventually it becomes kinda yucky having to wait.
The compiler may know what is programmatically correct, not what your intention is. I think that's what the other person was trying to argue.
I'm not too concerned about compile times myself, I'm used to them given I mostly dabble with compiled languages (I do a lot of CSS though!), but there's a argument to be made about iteration time and being fast in solving a problem.
So no, 72GB or 128GB makes no difference.
The thing about rust is that the overhead is all upfront.
Not to be sacrilegious, but bleeding at the move or &wtf borrow boundary is a price one pays to have confidence, subsequently.
Your cms/customer acronym probably doesn’t need it, but you’ll love the —fast option in some stacks that merely means “compile the rust extension.”
Plus… it’s a gainful hobby. I’m fine if other people have hobbies that don’t resemble x86 in 1998 on slackware ;)
The elephant in the room is monomorphisation and macro expansion. Those generate a ton of code that LLVM has to process. In return you get much faster programs at runtime.
> [package.metadata.leptos]
> separate-front-target-dir = true
This option is deprecated since cargo-leptos 0.2.3 (now it's enabled unconditionally), it's going to be removed in 0.3
https://github.com/leptos-rs/cargo-leptos/pull/216
https://github.com/leptos-rs/cargo-leptos/issues/217
https://github.com/leptos-rs/cargo-leptos/commit/b0c19a87cff...
I’ve found in 10+ years of software development that speed of iteration cycle is highly correlated with productivity. Compile times is not the only input into this cycle time, but it’s a big one, and importantly, it’s within the control of the language tooling itself to solve. The human idle attention time of 1-2 seconds should be the gold standard to strive for, even if not always achievable.
There seems to be quite a bit of cope around Rust build times in the community, which was natural a few years ago (a lot of people used to “blame” llvm, but it doesn’t seem to be as big of a culprit) but things are different now, no? Given the maturity, growing ecosystem and corporate investment, I would expect incremental build speedup to be prioritized, and steadily improving. But clearly it isn’t moving very fast in that direction. So why not?
I agree, though in practice with many Rust programs that iteration cycle is actually not "edit, compile, run", it's "edit, save, wait for rust-analyzer to update" which is generally much quicker (even if it is sometimes still a little slow).
In most cases, Rust's very strong type system means you spend a lot less time actually running your program.
There are definitely exceptions though. E.g. if you're making a web app and editing CSS or whatever then the Rust compiler isn't going to tell you if you've got the colour wrong.
> a lot of people used to “blame” llvm, but it doesn’t seem to be as big of a culprit
It still is; there was a post here really recently where someone broke down the time spent in various phases. LLVM is still the vast majority, though apparently that is partly due to Rust generating a lot of work for it.
> I would expect incremental build speedup to be prioritized, and steadily improving.
It is.
> But clearly it isn’t moving very fast in that direction. So why not?
Because Rust's unit of compilation (a crate) is very large. In C/C++ it's a single file.
This is one area that I think Dart did a good job on. They support both a JIT (I am not sure if it’s tiered) and AOT modes of compiling, the JIT targeted for development and AOT for production. This is the ultimate extreme of optimizing for both use cases, obviously it’s a ton of work, but I would love more languages to be like this. (I believe the JIT mode works directly on the source so there isn't an intermediate compilation step like Java, but it’s been a few years since I’ve used Dart)
Have all variants available and allow the devs to pick the right one for their current workflow.
Java, .NET, Eiffel, Common Lisp, Scheme, C++ (via stuff like ROOT and Live++, VC++ hot reload isn't that full proof), OCaml, Haskel, are all examples of these mixed tooling approach.
It is also generally very easy to break apart any file or decouple any header files using various techniques that don't affect your project layout or require any mental maintenance once done. In contrast, with Rust, you have to essentially refactor what your project even "means" to pull that off as you now need to be juggling a ton of microcrates (and then it seems like there is some inherent crate overhead so this becomes a painful tradeoff plane you need to optimize; compilation units have to become degenerated small before that matters in C++).
And, even then, the remaining slowest file in my insanely complex C++ project does tend to take a ridiculously long time to build (I have a 30-60 second long compile for one small component) and I have run out of low-hanging fruit; but, I honestly hardly ever actually have to build it: the coupling of header files is massively overstated as when you are in a tight build/update cycle you are generally modifying a single function in a source code file, not changing some header file of basic data types that will affect the code generation of literally the entire rest of the project.
Like, I get it: thinking about code as a merged set of linear forward passes is sometimes annoying, but it doesn't just enable massive gains in compile time... it also encourages more legitimate coupling between components as circular dependencies require extra code and upkeep for prototypes and that's good because circular dependencies are bad. Rust thereby is doing something that frankly feels super awkward: it is elevating behavior that frankly makes more sense only within a single file or at best a single object to a project organizational unit, which suddenly makes this all seem much more complicated to handle correctly.
That just sounds wrong to me TBH. IME it never pays off to 'front-load' too much, you might get a working implementation of your original idea without ever running the code, but that original idea will almost certainly be "wrong" once you see it in action and you just wasted a lot of time agonizing over your type structure which turned out to implement a quickly outdated idea.
Nothing beats quick iteration on a running program when it comes to productivity.
> Because Rust's unit of compilation (a crate) is very large. In C/C++ it's a single file.
If the a C++ file includes any C++ stdlib header it will also also become very large (tens of thousands of lines of gnarly template code for common headers like <vector> or <string>). If the unit of compilation is a crate in Rust it should actually be faster to build than typical C++ projects which often put the implementation for one class into one source file (and each of those source files is expanded to tens of thousands of lines of code after includes are resolved).
C mostly doesn't suffer from this problem though, since C header typically are only a couple dozen or hundred lines of API declarations.
> C mostly doesn't suffer from this problem though, since C header typically are only a couple dozen or hundred lines of API declarations.
But then, there's no reason why Rust couldn't be like C. The issue comes from the dependencies and libraries bundled, which is not an issue per se of the language.
Rust-analyzer is struggling significantly on larger repos, to the point of being barely usable.
What do you mean exactly here?
So the smallest chunk of work that can be parallelised is much bigger in Rust than in C++. (Unless you are writing atypically small Rust crates and atypically enormous C++ files.)
I ask becasue I wonder what's the real culprit here and whether it can be solved or if it's due to the language.
While staying in the same ecosystem you can split your application into smaller crates, which many are doing. It's just not that ergonomic.
> There are definitely exceptions though.
I mean.. I know what you mean but this is kinda what I meant with cope. Putting “running code” aside as an edge case seems like a dangerous moving-the-goalpost kind of Rustism that kinda grew out of the “Rust is so safe that if it compiles it’s correct”-meme. Once we’re doing serious and boring good old development, it’s a redeeming quality at best, imo. Not to mention that rust analyzer also suffers slowdowns in similar ways.
> LLVM is still the vast majority
That’s interesting. So why the meager numbers from cranelift then, judging by the post? Granted, I’m not doing Rust day-to-day anymore so I’ve been ootl.
> Because Rust's unit of compilation (a crate) is very large. In C/C++ it's a single file.
Isn’t this the same for Go? Sure, it’s a simpler language but we’re talking many, many times on similar sized projects (wish I could link to something but unfortunately anecdotally only). Is it monomorphization, macros, or just everything combined?
People are seeing 20-30% improvements thanks to Cranelift. Not sure why you think that’s meagre.
> edge case
It’s not an edge case but it is valid to say that the pattern of development is different. You’re right that people need feedback that what they’re doing is right within 1-2 seconds. They are getting that thanks to the IDE. They need their tests to compile and run within a few seconds, and they’re getting close to that.
I’m not disagreeing with you. In an ideal world all builds and tests would run within 2 seconds, like you’re suggesting. And yet, people are managing to ship software successfully, despite it taking more than 2 seconds. That suggests that maybe this 2 second thing isn’t a requirement and more of a nice to have. And that maybe, there is something to people relying on IDE feedback while coding.
I don’t know if it’s unpopular, but it may be inaccurate. OP claims that they improved by 75%, so you probably meant “25% of a lot is still a lot”.
> I would expect incremental build speedup to be prioritized, and steadily improving. But clearly it isn’t moving very fast in that direction. So why not?
Is this assessment based on something?
For what it’s worth, people have been working on it for years, with varying approaches and levels of success. The recent spate of articles you’ve seen is precisely because these approaches are bearing fruit. Several folks have worked on making a previously single threaded frontend multithreaded, others have added an option to swap the LLVM backend out for a much faster one (cranelift), and yet others are working on moving to a faster linker by default.
The net gain from all of these once they land could be as much as 30–50%, depending on the project, mode of compilation (clean or incremental) and hardware (single/multiple cores). That is a pretty fantastic improvement, considering the whole year improvements in the last 4 years have been 7%, 17%, 13% and 15%, for an overall speed gain of 37%.
If all of this doesn’t sound impressive, I’d encourage you to read nnethercote’s annual reports which explain how much effort went into making compile times faster. 2023’s report - https://nnethercote.github.io/2024/03/06/how-to-speed-up-the...
Ultimately, all of these improvements might not be enough for the people who need HMR to be productive. That’s possible. But the idea isn’t to get to a point where every line of code written in the world is Rust. There’s plenty of Rust development happening today. When the compile times drop as hoped, more people join and that grows the ecosystem.
Yeah exactly. This was afaik a fairy small sized application (where the majority was probably framework code) which went from 20s to 5s for incremental compilation on a very beefy desktop rig. Those absolute numbers are important to grok the situation.
> That is a pretty fantastic improvement, considering the whole year improvements in the last 4 years have been 7%, 17%, 13% and 15%, for an overall speed gain of 37%.
> If all of this doesn’t sound impressive, I’d encourage you to read nnethercote’s annual reports which explain how much effort went into making compile times faster
It’s not at all that it’s unimpressive. It’s abundantly clear that extremely bright minds are putting a ton of effort into it, consistently and over time. I’ve worked personally with some of these people and they are insanely talented and knowledgeable… Which is precisely why I’m concerned that we’re not seeing the big boosts that comes with picking low hanging fruit, perhaps there isn’t any left. Optimization tends to yield lower results over time (although the recent advances are great to see and I wish it continues). From the outside though, it appears like either inherent constraints that limit how far you can go, or at least initial compiler design reasons that would be cost-prohibitive to change.
As the link I shared says
> For the first time ever, I’m writing one of these posts without having made any improvements to compile speed myself. I have always used a profile-driven optimization strategy, and the profiles you get when you measure rustc these days are incredibly flat. It’s hard to find improvements when the hottest functions only account for 1% or 2% of execution time. Because of this I have been working on things unrelated to compile speed.
> That doesn’t mean there are no speed improvements left to be made, as the previous section shows. But they are much harder to find, and often require domain-specific insights that are hard to get when fishing around with a general-purpose profiler. And there is always other useful work to be done.
---
> From the outside though, it appears like either inherent constraints that limit how far you can go, or at least initial compiler design reasons that would be cost-prohibitive to change.
Also correct. Rust prioritised faster run time over fast compilation times. For example, monomorphization is something used pervasively in most Rust code. Fast at run time, but slow to compile. The fact that the entire crate needs to be compiled if a line is changed is yet another issue.
Everything indicates that even if Rust improves, it will never reach that 2 second recompile for large projects. But fortunately the experience of the last few years suggests that it's still possible to succeed despite (or maybe because of) prioritising runtime over compile time.
Thanks anyway for the work on sold.
Essentially, yes, but that’s because Apple’s new linker is so much faster than it was. You can use it instead.
That's the least I can do to show my appreciation for your incredible work; I can't thank you enough brother!
I also wish to see that chibicc compiler plus the book completed. I want to learn everything around such topics.
If you use a lot of VMs, it comes in handy. If you do a lot of heavyweight compilation in parallel, same thing. But at least on Linux, popping 3 VS Code projects and 4 more browsers doesn't even hit 50% utilization.
But the 7950X is grossly underutilised, I don't know what to do with it.
MSAN: https://github.com/google/sanitizers/wiki/MemorySanitizerI recently upgraded to 64GB to run quantized LLM models, but when I'm not using an LLM I rarely use more than 16GB, much less 32GB.
Never have the same problem on Windows even though I only have 16GB there. Feels like Linux is just really bad at memory management. That may be skewing the Steam survey since almost everyone there will be using Windows.
On Linux I wish I had 64GB.
I believe Fedora comes ootb with a lot of things to make ram management quite good. Using another distro these days, but Fedora impressed me on that front.
(edit: Try to use zram and zswap, check if your i/o scheduler is correct for your hardware, and increase your swapiness value. These should make your system behave better. For other performance improvements, use ananicy and irqbalance.)
Even if it is possible to fix this, it's still a horrible bug because you absolutely should not have to blindly tweak settings to stop a desktop computer from hard-rebooting when it runs out of RAM.
I think this is a really fundamental problem on Linux though - it would require enormous changes and unprecedented cooperation between different groups for a proper fix. Much easier to say "you're holding it wrong".
If it's the latter, that's because Linux deliberately uses as much RAM as possible for caches and such, which means "free" should be near zero most of the time. The number that actually matters on Linux is "available" memory, which is the amount of memory that is currently free or used for optional caches and could be reallocated for a new required chunk of RAM should an application need it. Here [0] is a decent in-depth explanation.
If it's the former, there's something very wrong with your installation, because every Linux system I have makes much better usage of memory than Windows, even running a heavyweight DE like KDE. Even my machine that dual-boots Linux and Windows performs noticeably better on all metrics (including memory usage) when in Linux.
So no, 72GB or 128GB makes no difference.
Rust builds, in comparison, don't actually reach that far, mainly because it is proportional to the maximum number of parallel jobs. In my experience, 48 GB was more than plenty for 8 parallel jobs if I wasn't doing anything else.
But in this article the 7900x performed worse than the 5950x, is it because of the core count, or are there some other factors?
The performance looks good, but are there any plans for any other improvements such as LTO?
Is there a project that publishes benchmarks compiling projects across different CPUs? I know passmark is kinda the go-to but this would be cool.
Someday I'll feel the need to upgrade from a 2700x...
Software development: the art of redistributing aggregate lifecycle pain; who bears what, when and how much.
That's amusingly ironic, since it doesn't originate in financial markets anyway!
Just looking at the blog this seems like a fully static site that should be able to be built in milliseconds. Some input from the developer here on what takes all this time would be interesting.