This isn’t the way to speed up Rust compile times
xeiaso.net
xeiaso.net
Off topic: I don't understand why the article repeatedly says the Rust compiler or procedural macros are already fast, even "plenty fast". Aren't they slower or about as slow as C++, which is notorious for being frustratingly slow, especially for local, non-distributed builds?
"Plenty fast" would mean fast enough for some simple use-case. e.g., "1G ethernet is plenty fast for a home network, but data centers can definitely benefit from faster links". i.e., fast enough you won't notice anything faster when doing your average user's daily activities, but certainly not fast enough for all users or activities.
The Rust compiler is then in no way "plenty fast". It has many benefits, and lots of hard work has gone into it, even optimizing it, but everyone would notice and benefit from it being any amount faster.
Why?
I mean, OS distributions making available whole catalogs of prebuilt binaries is a time-tested solution solution that's in place for decades. What leads you to believe that this is suddenly undesirable?
I've clarified it to "upstream-provided 3rd party" in the original post.
Binaries are relatively inscrutable, and there are massive differences in ease of doing something nefarious or careless with OS binaries, and upstream checked in source code, and upstream checked in binaries.
That's just one or two reasons why I think this approach is not at all commonplace in open settings, and is generally limited to closed and controlled settings like corporate environments.
A more typical and reasonable approach would be having a dependency on such a binary that you get however you like: OS, build from upstream source, upstream binary, etc., and where the default is not the latter.
And now if you allow me to get a bit off-topic again: to me, having to "innovate" like this is a major sign that this project knows it is very slow, and knows it's unlikely to ever not be very slow.
That tells you more about crates.io than consuming prebuilt binaries.
Some Linux distro ship both source code packages and prebuilt binaries. This has not been a problem for the past two decades. Why is it suddenly a problem with Rust's ecosystem?
Some problems have solutions, but people need to seek for solutions instead of collecting problems.
Proc macros have effectively become fundamental to Rust. This is unfortunate, but, given that it is reality, they should get compiler support. Part of what is making them so slow is that they need to process a firehose of information through a drinking straw. Placing proc macros into the compiler would not only speed them up but also remove the 50+ crate dependency chain that gets pulled in.
Serde, while slow, has a more fundamental Rust problem. Serde demonstrates that Rust needs some more abstractions that it currently does not have. The "orphan rule" means that you, dear user, cannot take my crate and Serde and compose them without a large amount of copypasta boilerplate. A better composition story would speed up the compilation of Serde as it would have to do far less work.
Rust has fundamental problems to solve, and lately only seems capable of stopgaps instead of progress. This is unfortunate as it will wind up ceding the field to something like Carbon.
What does the orphan rule have to do with the compilation speed of serde? How does it reduce the work it needs to do?
Serde has to enumerate the universe even for types that wind up not used. Those types then need to be discarded--which is problematic if it takes a non-trivial amount of time to enumerate or construct them (say: by using proc macros). Part of this is the specification that must be done because Rust can't do compile-time reflection/introspection part of which is blocked because of the way Rust resolves types--different versions of crates wind up with different types even if the structures are structurally identical because Rust can't do compile time reflection and relies on things like the Orphan rule.
I am not saying that any of these things are easy. They are in fact hard. However, without solving them, Rust compilation times will always be garbage.
I'm beginning to think that the folks who claimed that compiler speed is the single most important design criterion were right. It seems like you can't go back and retrofit "faster compiler time" afterward.
First of all it's not 50 crates, it's 3 main crates syn, quote and proc-macro2 and a few smaller helper crates which are used less often.
And the reason to not include them in the compiler is exactly the same reason why "crate X" is not included in std. Stuff in std must be maintained forever without breaking changes. Just recently syn was bumped to version 2 with breaking changes which would have not been possible if it was part of the compiler.
In any case, shipping precompiled proc macros WASM binaries will solve this problem.
> Serde has to enumerate the universe even for types that wind up not used. Those types then need to be discarded--which is problematic if it takes a non-trivial amount of time to enumerate or construct them (say: by using proc macros). Part of this is the specification that must be done because Rust can't do compile-time reflection/introspection...
Yes, it's a design tradeoff, in return Rust has no post-monomorphization errors outside of compile time function evaluation. I think it's worth it.
> I'm beginning to think that the folks who claimed that compiler speed is the single most important design criterion were right.
I don't agree. As long as the compile times are reasonable (and I consider Rust compile times to be more than reasonable) other aspects of the language are more important. Compile time is a tradeoff and should be balanced against other features carefully, it's not the end goal of a language.
That may be technically true. However, every time I pull in something that uses proc macros it seems my dependency count goes flying through the roof and my compile times go right out the door.
Maybe people aren't using proc macros properly. That's certainly possible. But, that also speaks to the ecosystem, as well.
I have two dependencies (memchr, urlencoding) and one two dev dependencies (libmimic, criterion). I have 77 transitive dependencies, go figure. People like to outsource stuff to other libraries, when possible.
The orphan rule isn't a problem that needs solving, it is itself a solution to a problem. There's nothing stopping Rust from removing the orphan rule right this minute and introducing bold new problems regarding trait implementation coherence.
It is both.
The Orphan Rule was indeed chosen as a solution because other choices were dramatically more difficult to implement.
However, the Orphan Rule also has tradeoffs that make certain things problematic.
Engineering is tradeoffs. There are only two types of programming languages: those you bitch about and those you don't use. YMMV. etc.
I really don't want to see something like Carbon take off. It's not really that much "better" than C++, and it will take up the oxygen from something that could be a genuine replacement.
Yes. Significantly slower. The last rust crate I pulled [0] took as long to build as the unreal engine project I work on.
Clean and rebuild of Unreal Engine on my 32-core Threadripper takes about 15 minutes. And incremental change to a CPP takes… varies but probably on the order of 30 seconds. Their live coding feature is super slick.
I just cloned, downloaded dependencies, and fully built Symbolicator in 3 minutes 15 seconds. A quick incremental change and build tool 45 seconds.
My impression is the Rust time was all spent linking. Some big company desperately needs to spend the time to port Mold linker to Windows. Supposedly Microsoft is working on a faster linker. But I think it’d be better to just port Mold.
The file formats are indeed totally different. But the operation of linking is the same at a high-level.
My first clone on symbolicator took the same length of time on my windows machine. Even with your numbers, 4 minutes to build what is not a particularly large project is bonkers. I
My experience across a wide range of C++ projects and wide range of Rust projects is that they’re roughly comparable in terms of compilation speed. Rust macros can do bad things very quickly. Same as C++ templates.
Meanwhile I have some CUDA targets that take over 5 minutes to compile a single file.
I feel like if Rust got a super fast incremental linker it’d be in a pretty decent place.
[0] https://www.cpubenchmark.net/cpu.php?cpu=AMD+Ryzen+Threadrip...
(Even most MBPs are unaffordable for a large set of developers.)
I don't think it's as accessible to average C++ or Rust developers as you expect.
Lenovo P620 is a somewhat common machine for large studios doing Unreal development. And it just so happens that, apparently, lots of people in this thread all work somewhere that provides one.
I don’t think the story changes much for more affordable hardware.
I've also got a 14" Macbook Pro (personal machine) that was a _little_ cheaper - it was £2700.
> An expense that's very unaffordable for most companies I think it's unaffordable for some companies, but not most. If your company is paying you $60k, they can afford $3500 once every 5 years on hardware.
> I don't think it's as accessible to average C++ or Rust developers as you expect.
I never said they were accessible, just that they are widespread (as is clear from the people in this thread who have the same hardware as I do).
FWIW, I was involved in choosing the hardware for our team. We initially went with Threadrippers for engineers, but we found that in practice, a 5950x (we now use 7950x's) is _slightly_ slower for full rebuilds but _much_ faster for incremental builds which we do most of.
I don't think that's true at all, unless you're using a very personal definition of "common".
In the real world, teams use compiler cache systems like ccache and distributed compilers like distcc to share the load through cheap clusters of COTS hardware or even vCPUs. But even that isn't "common".
Once you add CICD pipelines, you recognize that your claim doesn't hold water.
Rust also does significantly more for you at compile time than C++ does, so I don't mind the compiler taking some not time to do it
Why do you need to build hundreds of dependencies if you're not touching them?
Also, I've spent a _lot_ of time with Unreal and the build system. Unreal uses an "adaptive unity" that pulls changed files out of what's compiled every time. Our incremental single file builds are sub-10-seconds most of the time.
Ccache works, but if you use the Visual Studio C++ compiler you need to configure your build to be cacheable.
Out of curiosity, why do you use precompiled headers? I mean,the standard usecase is to improve build times, and a compiler cache already does that and leads to greater gains. Are you using precompiled headers for some other usecase?
> and a compiler cache already does that and leads to greater gains
Can you back that claim up? I've not benchmarked it (and I'm not making a claim either way, you are), but a build cache isn't going to be faster than an incremental build with ninja (for example), and I can use precompiled headers for our common headers to further speed up my incrementals.
You did encourage me to go back and look at sccache though, who have fixed the issues I've reported with MSVC and I'm going to give it a try this week
If anyone knows a method for to avoid such launch (besides connection blocking after launch by firewall ), share it please.
There are workarounds like sccache, but again additional tooling to take into account, across all target platforms.
Some points I hadn't considered:
- Current compile times aren't too bad, but compile times are a factor in how macros are developed. Making that less of an issue will likely make the macro ecosystem richer and more ambitious.
- Precompiled binaries can be compiled with optimizations, which anecdotally overcome the potential perf hit from running in WASM
- Nondeterminism in macros (especially randomness) has bit a lot of people and a more controlled execution environment is very beneficial
- Sandboxing doesn't have to be WASM, it's just the most capable, currently available option.
- Comments on the pre-RFC are exploring solving this in rustup/cargo instead of crates.io and that is also promising.
In general, I appreciate how this is creating awareness and driving exploration of solutions. In particular I really like the notion that there is currently implicit pressure on macro authors to make them somewhat quick to compile and releasing that pressure will enable innovation.
WASM with optimisations will mostly run faster than binaries without optimisations, which is the current strategy. Plus you won't need the compile step that's necessary at the moment.
Distributing arbitrary binaries right now is bad because it's not clear what those binaries contain, but the WASM environment is completely sandboxed, so it wouldn't be possible for them to read files they weren't meant to or phone home with telemetry.
The Cargo project already does some amount of automatic building e.g. of docs pages - this could be extended to proc macro crates to ensure that the uploaded binaries must be built from the uploaded source files. Alternatively, making sure that reproducible binaries are standard would help a lot from a trust perspective.
This isn't suitable for every situation, but as a tool to reduce the compile times of some of the bigger and slower macro crates, it sounds like a great option to have.
The nokogiri rubygem has been shipping a prebuilt windows dll for probably more than a decade. The sun has not imploded into a black hole at any point in the meantime.
> If we're going to be trusting some random guy's binaries, I think we are in the right to demand that it is byte-for-byte reproducible on commodity hardware without having to reverse-engineer the build process and figure out which nightly version of the compiler is being used to compile this binary blob that will be run everywhere.
I have some really, really fucking bad news about the O/S distros most of you all use...
And that is a reasonable thing for those projects to try to achieve.
Still we've been downloading binaries for decades and the world hasn't ended, this issue never warranted people treating it like nuclear war was imminent and hurling abuse into Dtolnay's github issues like he was a child pornographer. There's a way to have a discussion about this being the wrong direction without doing that.
And I'd encourage everyone to take an honest audit of how many precompiled binaries are in their lives, even if they've switched to NixOS and how much stuff they trust downloading from the internet. Very few people achieve an RMS-level of purity in their technological lives.
There is a very real problem that he was trying to solve, but the right way to do it will require core language support. And his heart is certainly in the right place, since "slow build times" is probably the #1 complaint about Rust.
If you actually polled most Rust users they probably care about their build times more than they care about this issue, particularly if you designed the survey impartially and just asked them to stack rank priorities. They probably all would prefer "both" being the right answer though.
Where the criticism really needs to be leveled is on the core language design (although I appreciate that their job is also extremely tough).
[And the one good thing to come out of this shitshow will probably be getting people focused on solving this issue the right way]
Maybe they could learn a trick or two from each other?
Java’s standard HashMap isn’t thread safe, it has ConcurrentHashMap[0] for this reason. And I’ve never heard somebody refer to Java as an unsafe language. Lemire wrote a blog article arguing Java is unsafe for similar reasons that it seems you’re using here, and I think the comments section has some great discussion about whether this conflation of language safety makes sense or not[1].
And one final point: a lot of the world lives on software segregated across several different machines, operating asynchronously on the same task, usually trying to retrieve some data in some central data store on yet another machine. Rust is great for locking down a single monolithic system, but even Rust can’t catch data races, deadlocks, and corruption that occurs when you’re dealing with a massively distributed system. See this article that talks about how Rust doesn’t prevent all data races, and why it’s probably infeasible to do that[2]. And I doubt anyone here is about to call Rust unsafe :)
This conflation of language often confuses me and I hope as an industry we can disambiguate things like this. Memory safety is distinct from thread safety for a reason, and by all reasonable standards Go is as safe as a modern language is expected to be. (Obviously better thread safety is a goal to be lauded, but as an industry, it seems like we’re still trying stuff out and seeing what sticks in regards to that).
[0]: https://docs.oracle.com/javase/6/docs/api/java/util/concurre...
[1]: https://lemire.me/blog/2019/03/28/java-is-not-a-safe-languag...
From [2] you linked
> a race condition can't violate memory safety in a Rust program on its own.
From golang FAQ on maps not being atomic:
> uncontrolled map access can crash the program
This may be one very narrow case (or a more common theme), but at least in this case the thread safety becomes a memory safety concern in golang.
Otherwise you get into stupid situations like: https://xkcd.com/1172/ (your optimization is killing children!)
In FOSS there is the principle of four Fs:
- Fix it
- Fork it
- Fund it
- Fuck it/off
But we've gone from 0% to 100% overnight and as usual people have adopted it as their new religion and they want to burn all the heretics and there can be no compromise.
I seriously doubt that this one specific issue was all that important in the larger problem of securing the supply chain, and there was a very good reason why it was done (which has now been entirely thrown away, which will certainly harm adoption of rust). I don't think it was remotely comparable to the way all of npm is a security hole.
I personally don’t hold much value in ‘reproducible builds’ as being some silver bullet - back doors can almost as easily be hidden in source code, particularly if nobody gives it a once-over read, same goes for compiled code only it normally requires a sandbox and some R.E.
Still, I think in the open source world, the primary distribution method, at least the one emanating from the original developer, should be in forms of source code, to ensure that the original developer can't include special closed source features. Also, for end users it's easier when the source code is available, then they can rebuild things using standardized methodology. all the distro packages have a bunch of commands you can run and then you can rebuild the distro package.
So in summary, I concede you are right in most, but not all cases.
Java libraries are distributed as binary bytecode blobs. They can also include native libraries, which usually are extracted to /tmp and dynamically loaded into the process. The security of this approach is questionable, if you ask me, however that's the way things were for the long time and that's the way things will be for the long time.
I think that striving for better security is a noble goal, but if that prevents usability, security should step aside, when the world around is not as secure anyway.
I don't think it's fair to describe JARs this way. While it may be technically true, a JAR file is far less opaque than an actual binary executable.
There are some changes coming where application developers will have to explicitly acknowledge that libraries they're using are using native code: https://openjdk.org/jeps/8307341 (the new, soon-to-be-stable FFM API already has this behavior).
It doesn't change the distribution mechanism of native code, but it's some sort of step towards better security.
> Something about ads
> Title
> Something I skipped
> Something about patreon
I dismissed it as "yet another part of the headers" and not as "part of the article".
I mean, to be fair, I was already aware it was removed from serde, but I'm pretty sure I didn't notice that box in the article the first time I read through it.
The title is kind of hidden as it is right now, and the "notice" on the top doesn't look like a notice. Maybe try making it more different than other elements on the website, and put it below the article title rather than above?
It seems built-in to Swift, as opposed to a dynamically executed crate like with serde. I wonder how it's implemented in Swift and if it leads to any significant slowdowns.
[0] https://developer.apple.com/documentation/foundation/archive...
It was only fully addressed in Swift 5.9.
With ongoing work for Swift macros, it may eventually be possible to rip this code out of the compiler and rewrite it as a macro, though it would need to be a semantic macro[1] rather a syntactic one, which isn't currently possible in Swift[2].
[0] https://github.com/apple/swift/blob/main/lib/Sema/DerivedCon... [1] https://gist.github.com/DougGregor/4f3ba5f4eadac474ae62eae83... [2] https://forums.swift.org/t/why-arent-macros-given-type-infor...
1) Make using proc macros in your own application code faster: Encourage better macro re-export hygiene. Basically, one of the reasons serde + serde_derive is so slow is because for serde re-exports serde_derive to compile before it itself can be built. A solution (I didn't come up with this myself) would be to have a crate that re-exports serde and serde_derive together; see top comment herehttps://www.reddit.com/r/rust/comments/1602eah/associated_pr...
2) Make using libraries that use proc macros faster: Publish pre-expanded source code, such that no proc macros run for dependencies. Some work would have to be done around conditional compilation to make things just work (TM), but I think it could be done. (Who knows, maybe it'd be really hard to make the compiler deal with expanding proc macros on stuff behind conditional values. Also, there'd be issues regarding macro hygiene. Both solvable, but I'm not sure how much effort they'd take).
I don't think anyone has a right to demand anything of the project. The MIT license specifically has the whole "THIS SOFTWARE IS PROVIDED AS IS" spiel for a reason. Thinking that you can demand anything of an open source developer who, afaik, has no responsibilities in relationship towards you, is a rather toxic mindset that should be kept out of open source.
What a tiring argument. Why complain about anything? Why do anything differently? Maybe it's because people have invested time, money, and effort into supporting and using the project and because of your stupid decision they now have to do more work. There were a lot of comments about people showing up that monday to pin+vendor the old SerDe. Others planned to remove it entirely in favor of other libraries. Almost like actions have consequences even if your "license agreement" says otherwise.
> is a rather toxic mindset that should be kept out of open source.
I wish I lived in your echo chamber where everyone who uses your stuff is just totally, like, accepting of your garbage ideas. The precompiled binary idea was a classic garbage idea. Refusing to make it reproducible was just the corn on top of the cowpie. The author made a major screw up, doubled down, tripled down, and then finally gave in. The author has no idea what scale and scope of project this is used in. They also clearly did no evaluation on potential damage OR ask for feedback before moving it into mainline. For a library with 3M+ downloads this was the what third? Fourth? Classic ego-driven folly.
You know what else shouldn't exist in open source? Ego tripping morons.
He's free to tank his own project and you are free to use it, fork it, or not. He's allowed to be an ego maniac if he wants or do things in ways you disagree with. You got what he made for free, so why you want to demand things of him is beyond me.
The difference between this guy and the maintainer is that the maintainer did this guy a huge favor by writing this software for the guy where as this guy did nothing except complain about the free meal he's getting.
I think the burden of open source maintenance often goes too far in the other direction, especially when corporations demand free labor, but I think it's reasonable to expect maintainers of core libraries to not actively cause harm.
I can't speak to his motivations, but nothing that you've said seems to be against any law or any sort of agreement he's made. This is his software, he can do with it what he wants. If you don't like it just don't use it.
You can't police him, you can complain about him, but you seem to making up violations he has committed that are mostly made up, rather than based on any actual agreement he has entered into.
I don't care what he does with his library. You're right. At the same time you can't "do whatever you want" when you have a 3M+ download library even if you do own it. It would be like if you suddenly decided an entire metropolitan is wrong and you're right. You need A LOT of evidence to do that. There are practical limitations to your personal freedom when so many people depend on you. Even if you want to believe that isn't the case how many potential sponsors, contributors, etc do you alienate with such a stupid idea in a pool of 3M? Even if it's 1% thats 30,000 people who now will do absolutely nothing to help you.
Less time and money than rewriting the thing from scratch.
Rust needs specific toolchain improvements to cache partial compilations for reuse. sccache can cache some results, but it's not a comprehensive solution.
What specific proposal would replace serde? I'd like to know.
But generated picture at the beginning of the article, the one with space needle, robs me very wrong way, as real landscape doesn’t look like it. It’s like seeing Golden Gate Bridge in New York.
In the future I'm gonna use more of my own photography like this: https://pony.social/@cadey/110956174300162810, I'm just building up a library of viable photos.
https://docs.rs/serde/latest/serde/trait.Deserializer.html
Or the copy and pasting in Axum:
https://github.com/tokio-rs/axum/blob/24f0f3eae8054c7a495cd3...
If you're complaining about something that TFA is not an example of, it's a good idea to at least provide something that is an example of it, because otherwise you're just yelling at clouds.
David Tolnay is not a random guy.
And anyway it doesn't really matter. As soon as you use anyone's crates you're more or less completely trusting them. It's not difficult to hide malware in a Rust crate even if you don't ship a binary.
And... come on. David Tolnay came up with Watt. He's clearly not intending to ship a binary forever - the long term solution is WASM.
This author comes across as an annoying naysayer - everything is impossible, even things that have already been done like WASI.
Rustdoc is also automatically built by docs.rs and nobody distributes it in their .crate file. I think the same should be done for wasm proc macros, too: they should be built by public infrastructure, and then people can opt into using binaries provided by that infrastructure to do their development, and if they want also opt towards using native binaries instead of wasm. But the binaries, including wasm, should only be a cache.
That's obviously how it would work. Read dtonlay's proposal. Crates.io would compile the WASM.