2. Prefer dynamic dispatch to monomorphization (i.e., use fewer generics).
3. Don't use proc macros (i.e., don't depend on the `syn` crate, even transitively).
Easy to say; hard to put into practice. But that's all there is to it.
I want to make it /very/ clear that async isn't to blame at all for the pathological build times described here. It's a bug about traits and lifetimes, both very core concepts of Rust that you deal with even if you stay away from async code.
async rust will certainly be more ergonomic once some more improvements land (hopefully later this year), but I don't feel like it deserves all the sighs it's been publicly getting these past few months. (And I /love/ to complain. I've written pieces named "Surviving Rust async interface", "Getting in and out of trouble with Rust futures", "Pin and suffering", etc.)
> Prefer dynamic dispatch to monomorphization (i.e., use fewer generics).
Unless you hit a pathological case as shown in the article, it tends to not be _that_ bad, especially if you enable `-Z share-generics=y` (unstable still, yet enabled by default for debug builds if I remember correctly).
Overall still solid advice - although "use fewer generics" sometimes turns out to be "just turn a big generic type into `Box<dyn Trait>`" (it's not _just_ boxing, that would be `Box<T>`). That's what axum[1] does with all services, and it's never had the compile times issues warp[2] had, for example.
> Don't use proc macros (i.e., don't depend on the `syn` crate, even transitively).
Good news there, I hear there's some progress on the proc-macro bridge (which improves macro expansion performance) AND "wasm proc-macros". I hope this piece of advice will be completely irrelevant in a year (but for now, it's spot-on. using pin-project-lite instead of pin-project is worth it, for example).
Also for proc macros, rust-analyzer seems to struggle with them sometimes as well, so I try to avoid them (outside Serde, which is worth any price) for that reason.
In my experience, however, macros are usually not that problematic and `syn` is a one-time cost.
A lot of the async Rust code I work with already looks like `async fn foo() -> ... { do_request().await?.blah().await }`, plus the occasional gathering of futures into a `Vec` to join on. That sort of thing, not much different from Javascript, but with a lot more control of the low-level details.
A good deal of corner cases should get better once async traits are stabilized, which will mean much less need for manually writing out Future types. But honestly, even now it's not that bad. I have a codebase that uses async to read hundreds of thousands of files[1], streaming gunzip them, pass them to another future which streaming parses records from them, and then pushes those parsed records into a `FnMut` closure for further non-async processing. It took a bit of thinking and design to get everything moving together nicely, but that corner of the codebase now is only ~200 lines of pretty straightforward code -- there's like 1 instance of `Unpin`. It's not that bad.
[1]: I know async isn't necessarily faster for reading files, but it started life doing network requests and it can still saturate a 200-core machine so I haven't felt the need to port it over to threads.
The most egregious code comes when implementing one of the `AsyncRead`/`AsyncWrite` traits or similar, and that can come up a bunch in backend services, for example if you want to record metrics on how/when/where data flows, apply some limits etc. I'm curious how the ecosystem will adapt once async trait methods land for real.
The beauty of the rust async stuff is that you can move to a multi-threaded runtime as you desire with minimal effort.
As a heavy user of async Rust in production (at a couple places), resource leaks / lack of visibility into that has been a top issue.
In this area, tokio-console[1] is an exciting development. I have high hopes for it and adjacent tools in the future. (Instrumenting your app with tracing+opentelemetry stuff can help a lot, too).
Until those become featureful/mainstream enough, Go has the upper hand in terms of "figuring out what's going on in an async program at any given time".
This is also a downside, having multi-thread be the default and not single-thread. It introduces some awkward / accidental trait bounds that are annoying to deal with if you want to do thread-per-core type of stuff IIRC.
Pony does fearless concurrency better IMO, and Forty2 shows how we can expand on Pony to be faster and more flexible.
There are other approaches that have emerged recently too. For example, one can apply Loom's memory techniques to most memory management approaches to eliminate the coloring problem, to decouple functions from concurrency concerns.
There are also languages which separate threads' memory from each other which allows them to do non-atomic refcounting, relying on copying for any messages crossing thread boundaries (though that's often optimized away, and could be even less than Rust's clone()ing elsewhere).
One could also apply that technique to a language using generational references, if they want something without RC or tracing GC.
Sometimes I wish Rust waited just a few more years before going all-in on async/await. Alas!
I'm personally sort of skeptical about "color free async" because the models for sync/blocking IO and async IO are so different -- you can paper over the syntax differences, but you're going to be in a world of hurt when the semantic differences arise[2]. I'll admit I haven't tried a color-free async implementation myself though, so it's just speculation / sour grapes :-)
> There are also languages which separate threads' memory from each other which allows them to do non-atomic refcounting, relying on copying for any messages crossing thread boundaries (though that's often optimized away, and could be even less than Rust's clone()ing elsewhere).
Curious what you mean by this -- my understanding is that Rust also does this (i.e., you can `move |x|` a value into a thread and that thread owns it now, and then the thread can hand it back in a `JoinHandle`. That sort of sharing doesn't require an Arc or Mutex, since there's only one owner at a time. Is this something different?
[1]: The other day I turned something reading in files from the filesystem sequentially into a custom threadpool passing blocks of parsed JSON over a MPSC channel that exposed the whole thing as a sequential iterator and it worked first try. I almost didn't believe it until I wrote the tests.
[2]: E.g., "I wrote this and tested it with blocking IO but this syscall isn't supported by io_uring so in async mode it goes to a threadpool and passes some huge object in a message which kills perf with a huge memcpy", or some similar jank. Just spitballing on the type of thing I would fear happening, not a specific example.
Pony is garbage collected. Most of the reasons why Rust async/await are the way it is boil down to the fact that Rust is memory safe without using GC.
> Forty2 shows how we can expand on Pony to be faster and more flexible
I can't tell from a glance, but that also looks garbage collected.
> For example, one can apply Loom's memory techniques to most memory management approaches to eliminate the coloring problem
Assuming you're referring to the JVM Project Loom, that's just M:N threading. This was tried in Rust almost a decade ago. Nobody used it because the performance was not appreciably better than 1:1 threading.
> There are also languages which separate threads' memory from each other which allows them to do non-atomic refcounting
You mean like Rust? Like, that's exactly why Rust can have both Rc and Arc and still be safe.
> relying on copying for any messages crossing thread boundaries (though that's often optimized away, and could be even less than Rust's clone()ing elsewhere).
Ancient Rust did this, but it was removed because with the current immutability and borrow checking rules there is no need for copying anymore. Why would you want copying if you don't need it?
I'm also not going to just accept that clone() could be faster. I mean, I'm sure the clone codegen could be improved by better register allocation or whatever, but I don't think that's what you mean.
> One could also apply that technique to a language using generational references, if they want something without RC or tracing GC.
Why would you want to copy if you don't have to?
> Sometimes I wish Rust waited just a few more years before going all-in on async/await. Alas!
I haven't seen anything here that is better than Rust's async/await, and a lot that's either worse or doesn't fit with the rest of Rust's design.
But I can see Rust async being more of a gumption trap than many features. A gumption trap is a problem which uses up your motivation before you can work on the thing you actually wanted to do and so there's none left for the actual project.
Some folks just use threads, no event loop.
Event loops are also there without async, you could just write against mio or whatever else you choose directly.
https://perf.rust-lang.org/index.html?start=2022-01-01&end=&... (takes a while to load, too many datapoints for 6 months at once)
When it comes to efficiency optimizations I see much lower hanging fruits in the web space, some TypeScript projects I'm working on take longer to fully start the dev setup than a fresh optimized Rust build. I hope Bun will make the difference it promises.
Usually, even after a simple change a simple `cargo check` can take a minute or two on a beefy PC. That said, over time you get numbed to it :D.
For example, to do the same things as `rustc` but using C++, you need to run external runtime modifying program to check your program for memory errors .. so that extends your C++ "compile time" by fair amount to even begin to match what rustc does.
When you do full comparison of what rustc does, and compare it to what you need to do in C++ project to match it, the compile times aren't actually slow since rustc does more with actual guarantees.
Having some additional features and checkers is nice, and can even be a life saver, but having to pay a high cost each time I hit compile is not acceptable.
So the compiler will spend some time inferring the types of different variables and what traits those types implement, which is necessary to know which trait methods get called, etc.
Dynamically-typed languages get to skip all this and say "my type is either a number or an object, because everything in the universe is either a number or an object", and methods are selected through dynamic dispatch (usually by storing a ton of function pointers everywhere).
Rust can't just do dynamic typing and skip the compile-time logic, though you can get close by using lots of dyn traits.
It's not that relevant though, because the longest part of any Rust build is the dependencies, and those are never recompiled. Most projects compile relatively fast once you've built them once.