Why Async Rust?
without.boats
without.boats
Personally, though, I would strongly prefer to use async rather than explicit threading even for cases where performance wasn’t the highest priority. The conceptual model is just better. Futures allow you to cleanly express composition of sub-tasks in a way that explicit threading doesn’t:
https://monkey.org/~marius/futures-arent-ersatz-threads.html
In Java, Future.get() blocks current thread, and it is trivially integrated into explicit threading programming. In Rust, Future.poll() is not blocking, and one would need to rely on some async framework, or build own event loop which can potentially block thread.
You can spawn your tasks, store the JoinHandle "futures", and wait for completion whenever you need the result.
A difference being that Futures do nothing until polled, while threads start on their own, but that's arguably a helpful simplification for this purpose.
but instead I start a future, and then to run it at all I need to wait for the result. I understand the there are tools to effect this, but it really leaves you wondering - what did I just do? start an async task and then .. block on it in order to get it to execute?
In JavaScript terms Futures are more like sugar around callbacks, they don't do anything until you call/poll them. Tasks are independent entities like Promises which are being run by the executor, though they may currently be blocked on other tasks.
> the model I often want is I want to start some work, and then join at some later point - or even chain directly into the next task.
Rust wants you to do this the other way around. First chain together your futures so that when you start the top level one as a task there is a single state machine for it to run.
I disagree with you, my code looks safe and simple with explicit blocking threading, and at the same time is much simpler to reason about what is going on and tune in contrast to async frameworks which hide most of the details under the hood.
You can argue about performance, that async/epoll/etc allows to avoid spawning thousands of threads and remove some overhead, but there is no much benchmarks in internet (per my research) which would say that this performance overhead is large.
Having data affinity to cores is also great for cache hit rates.
Here is part of the C++ runtime this is based on: https://github.com/goto-opensource/asyncly. I was the principal author of it when it was created (before it was open sourced).
it doesn't sound they really sharing data with each other, it looks like your logic is well lineralizable and data localized, and you can't implement access to some global hashmap in that way for example.
> Try that with tens of thousands of actual OS threads and the associated scheduling overhead.
I run this(10k threads blocked by DB access) in prod and it works fine for my needs. There are lots of statements in internet about overhead, but not much benchmarks how large this overhead is.
> Here is part of the C++ runtime this is based on
yeah, I need one runtime on top of another runtime, with unknown quality, support, longevity and number of gotchas.
Yes, because data can have thread affinity. Data doesn't need to be shared by _all _ connections, just by a few hundred/thousand. This enables connections to be scheduled to run on the same thread so that they can share data without synchronization.
> I run this(10k threads blocked by DB access) in prod and it works fine for my needs. There are lots of statements in internet about overhead, but not much benchmarks how large this overhead is.
The underlying problem is old and well researched: https://en.wikipedia.org/wiki/C10k_problem
data doesn't need to be shared in your specific case, not in general.
> The underlying problem is old and well researched: https://en.wikipedia.org/wiki/C10k_problem
wiki page doesn't mean it is well researched, where can I see results of overhead measurements on modern hardware?
Here is how this works: at the bottom of the wiki page, there are referenced papers. They contain measurements in modern hardware. You read those, then perhaps go to Google and see if there is any newer research that cites those papers.
If you don't feel like reading papers, HN has a search bar at the bottom that yields a wealth of results: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
Maybe you should just take a college computer architecture course along the lines of Hennessy/Patterson. This is nothing new, I learned much of this in college 15 years ago. The problem has only gotten worse since then, computers have not become more single threaded.
> The problem has only gotten worse since then, computers have not become more single threaded.
Computers are now can handle 10k blocking connections with ease.
It's a library. It solved our problems at the time, years ago. It's still used in production and piping billions of audio minutes per month through it. You don't have to use it, I merely referred to it as an example. A similar library is proposed to be included in C++23: https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2023/p23...
there are tons of overengineered unmaintainable code in prod, it doesn't mean I need to follow them as example without much justification.
> A similar library is proposed to be included in C++23
hm, I went through the code example, and would prefer my current approach as a much simpler and readable.
I've (ab)used them that way, without any async runtime, just to easily write stateful iterators.
In threads that perform blocking I/O you don't get that, and need to support timeouts and cancellation explicitly in every blocking call.
This seems like a glib dismissal of a real problem. If you want to do threads-and-blocking programming in the large in Rust, you basically can't, because async has become a very strong default in the ecosystem. And i don't think anyone could look at the ecosystem and honestly form the opinion that it's that way because every library developer has made the informed decision that their use-case requires async, when use cases which require async are so incredibly rare - it is absolutely a social default.
> None of us can control what everyone else decides to work on, and the fact of the matter is just that most people who release networking-related libraries on crates.io want to use async Rust, whether for business reasons or just out of interest. I’d like it to be easier to use those libraries in a non-async context (e.g. by bringing a pollster-like API into the standard library), but it’s hard to know what to say to people who’s gripe is that the people putting code online for free don’t have exactly the same use case as them.
The criticism is not about the library authors, but about the stewards of the language and community who laid the path for them to follow.
It's not quite "colorless", but it's practical.
Pulling pollster into the standard library is a very minor change technically, but is much larger socially. The blessing of pollster will make blocking a first class citizen again.
I hope that people, especially the ones that have voiced these criticisms, take the time to read and digest this article. It may not change their mind, but perhaps it will help them understand the current situation better.
Unfortunately, there are already some comments that prove my optimism misguided elsewhere in this thread.
Not because of the technical decisions behind async/await, but because of the async ecosystem, especially the runtimes.
Picture a Rust user when the feature came:
- You can use it with future combinators!
- No, actually, use it with a macro library that almost works but not really.
- Here is Tokio (with a lot of moving parts).
- Wait, async std is much simpler (again, almost seems to work but not really).
- Wait, here is Smol, which is truly simpler (or at least smaller but not used).
=> You get "async fatigue."
I can understand the position: let's not commit to anything before seeing what sticks on the wall.
It might have worked for Serde (even that is debatable by some).
But as a user, it's hard to follow, and you get the impression that this feature is not the stable foundation you can build on.
I don't like the idea of having a single framework shaping up the entire Rust ecosystem, but at the same time once you just jump the tokio ship, everything just works out of the box and you don't have to worry. lots I wonder what are you referring to when talking about “lots of moving parts” when tokio has reached 1.0 three years ago and been pretty much stable since then.
But, particularly for library developers, one reason not to just jump on tokio and instead strive for (some, reasonable) compatibility across executors is the embedded world - async is an amazing match for programming tiny microprocessors because it provides an elegant syntactic sugar for all of those little interacting state machines, you would otherwise write, but you don't want a heavyweight thread-based executor to be mandatory.
(The embassy embedded framework for rust is an example of this. It doesn't yet have as wide support as some non-async frameworks but for the things that it's compatible with, it's an absolute delight.)
While it's not the default, Tokio is usable in full without threads using the current-thread executor: https://docs.rs/tokio/latest/tokio/runtime/index.html#curren...
Tokio internals are plentiful and more complex than in other runtimes. But the common difficulty is choosing what to use for an http server.
- hyper? just an http library but bring your own boilerplate?
- actix web? it was quite opiniated and there was some drama from its main author.
- axum? wait that seems to be the latest consensus actually.
And that's just the remaining popular choices.
I've been using Rocket.
https://github.com/tokio-rs/tokio-uring/
https://www.datadoghq.com/blog/engineering/introducing-glomm...
https://itnext.io/modern-storage-is-plenty-fast-it-is-the-ap... https://news.ycombinator.com/item?id=25220892 https://www.reddit.com/r/rust/comments/k16j6x/modern_storage... https://www.reddit.com/r/programming/comments/k0yyk7/modern_...
The impact of tokio on your app's code base isn't actually particularly big, and it wouldn't be too much trouble to change (the interface of other runtimes is very close to tokio's AFAIK). the main issue is the ecosystem: most of it is using tokio already so opting out of tokio also means cornering yourself in a place where there's little available external libraries to use.
The ecosystem strong ties with tokio isn't a good thing, and I wish there were ways to make things generic over the runtime, but it's not an application developer's decision in any case.
> Other ways of arranging the computation & I/O, like the thread-per-core model of glommio and io_uring, fundamentally change the API. There's even a second implementation of Tokio with an API different from the first one!
It's an incompatible API in the sense that you need to update your code, but it doesn't require deep re-architecture work or anything (going from thread-per-core like glomio to tokio would be harder for instance, because then you'd need your futures to be Send).
It should be noted that AFAIK on all modern operating systems, only one page of the stack is actually allocated on thread creation, and the rest is merely reserved address space that will be allocated on demand.
That doesn't make this point wrong: 4 KB is still a lot more than the couple of bytes a future might need. And setting up the page table is part of what makes spawning threads so expensive.
That was one way to scale up to a large number of clients when memory was more limited.
It makes the memory usage acceptable in many contexts. E.g., 4 KiB is small compared to a socket buffer. YMMV.
> And setting up the page table is part of what makes spawning threads so expensive.
Stacks can be reused.
The async Rust book [0] says this about runtimes:
> Importantly, executors, tasks, reactors, combinators, and low-level I/O futures and traits are not yet provided in the standard library. In the meantime, community-provided async ecosystems fill in these gaps.
Notably, it says "not yet". My question is if someone knows if there are actual plans to incorporate any (existing) async runtime, and if so, whether there is a timeline? Also, is tokio in the talks to be the runtime, or is this still open?
[0]: https://rust-lang.github.io/async-book/08_ecosystem/00_chapt...
There are possible plans to implement a minimal runtime that only allows you to execute futures. No I/O or such. Mainly for tests or examples. No timeline though, at least as far as I'm aware.
There are also plans to standardize the async traits, e.g. spawn a new task, AsyncRead and AsyncWrite etc.. I don't think there is a timeline, but they wait at least for async fn in traits.
That was just merged! [1] It should be stable in 1.75 on December 28th.
https://docs.rs/pollster/latest/pollster/
Technically pollster is a runtime, although in practice it's an anti-runtime.
You've done so much great work on Rust education, documentation, community building, etc. I'd be happy to send money your way.
That said, I appreciate the kind words. But I am struggling to have the energy to accept conference invitations and do smaller open source work, so I wouldn't commit to anything like that any time soon. (that said I literally arrived in Raleigh for All Things Open earlier today, where my talk will actually be about async/await, so... never say never.)
Everything you've done is greatly appreciated!
despite its complexity and slight ergonomic annoyances, async rust is a monumental achievement: safe userspace concurrency without heap allocation!
Twisted built async on top of Python's generator functions, as far as I understand. I see now that the Rust community talks about supporting generators, and I've wondered if Rust did things backwards. If Rust had received generators first, could async have been built on top of generators without async specific special syntax?
https://doc.rust-lang.org/beta/unstable-book/language-featur...
I think that's because there was a huge demand for async specifically, and Rust could ship higher level async without solving all the design details of less desired generators first.
It feels like it would be totally possibly to improve OSes to remove this limitation. Is anyone actually working on that?
For example I don't see why you couldn't have growable stacks for threads. Or have first class hardware support for context switching. (Yes that would take a long time to arrive.)
This is just my guess, but I assume they don't use growable stacks for the same reason Rust doesn't. C doesn't have a garbage collector, and making a fragmented stack would incur unacceptable performance penalties for many workloads. Getting around that would require a ton of work, but maybe with hardware support, like automatically following the stack pointer to bring it into cache like a normal stack could get around that.
The main solution is to just reserve a ton of virtual address space but avoid committing to it until the process actually writes to it, which is exactly what OS threads do. They reserve a large amount of virtual address space to start but it's more or less free until a thread actually uses it. However you may not see it released back to the OS until the process exits.
So you can keep dedicating more and more silicon to redundant components to get closer to physically representing each thread, or you can code more efficiently.
Eg even if you did nothing but loop over a do-nothing system call, it would still need to have two separate executable pages in the cache instead of just one.
Not only that, but often the kernel is acting as a mediator for the hardware - which could mean synchronizing between cores, which brings its own obstacles (using slower shared cache, waiting for other cores, etc)
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
The Switchto patch allows you to just tell Linux which thread to switch to, so it doesn't need to figure it out. Looks like they reduced the costs by ~20x. I wonder why it was never merged.
I guess my intuition was right then. I mean, async is still useful for WASM and microcontrollers, but for "I need to support 10000000 concurrent connections" (which is the usual motivation), it's a hack around poor OS APIs.
OCaml is a much better option in the UNIX world.
The thing about async Rust that absolutely destroys my efforts with it is the legion of problems involving dependencies and their async runtime peculiarities.
As a result I've had to set Rust aside for most things where I'd otherwise love to employ it. I can't risk losing days to some tangle of an async runtime compatibility snafu. The one area that Rust has really astonished me recently is embedded Rust and MCUs: the ecosystem there is still young, but the results that can be produced with Rust in embedded development are really astonishing.
Of course, there is typically no async runtime involved at that level, so problem solved...
At this point in time, if you use Tokio, won't basically everything mainstream work? At least that has been my experience so far.
I get that it is sometimes needed (if you use GTK or browser JS as your async runtime, or write your own executor).
But for majority users there is only one runtime and there's no compatibility problem. Tokio won, and network effects killed everything else. Think of it like Golang's choice of runtimes (there isn't one).
I think this would simultaneously solve two of the major gripes expressed here and elsewhere.
1) An easy answer for those who wish to avoid being sucked into async just because there's a useful crate that has async.
2) Blessing an executor other than tokio avoids tokio lock in.
Coming from the Python world where I've written a ton of async code there (which I think is just a short way of saying cooperative concurrency?), it's generally easier to reason because you don't need to worry about getting preempted anywhere. If your application is mostly comprised of hurry-up-and-wait I/O code, single thread, single process can be great.
It's literally "why async Rust" as in "why async Rust is the thing that it is," why it is that way, how it came to be, etc.
There's some unavoidable level of relationship, but that's not the focus of the article.
i.e pre allocated worker threads and work queues
perhaps all we needed was some nice semantics around that
The biggest pain coming up is how io_uring changes the shape of the read/write APIs everywhere :-/
https://www.datadoghq.com/blog/engineering/introducing-glomm... https://itnext.io/modern-storage-is-plenty-fast-it-is-the-ap... https://news.ycombinator.com/item?id=25220892 https://www.reddit.com/r/rust/comments/k16j6x/modern_storage... https://www.reddit.com/r/programming/comments/k0yyk7/modern_... https://www.youtube.com/watch?v=PbgTyCSDPrs
Also, as I mentioned below, Rust fares even better than C++ on minimizing allocations here.
C++ is used in high performing network services all the time, it shouldn't be a shock that Rust gets used for them as well.
As a counter argument though, clang's support for coroutines is still buggy and not ready for production use.
This is admittedly an unimportant correction to what you said that doesn't in any way change your point; and yet, I still think it is an important one to keep in mind for people who only might end up with an indirect understanding of the C++ feature: C++ additionally chose to add a specialized co_yield... and specifically does not have co_async! This latter tradeoff then relates to the Rust discussions I have seen come up again recently due to the article "Was async fn a mistake?".
https://seanmonstar.com/post/66832922686/was-async-fn-a-mist...
It was clear, though usually left unsaid, that what Rust needed to succeed was industry adoption, so that it could continue to receive support once Mozilla stopped being willing to fund an experimental new language. And it was clear that the most likely path to short-term industry adoption was in network services, especially those with a performance profile that compelled them at the time to be written in C/C++ ...
The other advantage of network services was that this wing of the software industry has the flexibility and appetite to rapidly adopt a new technology like Rust. The other domains were - and are! - viable long term opportunities for Rust, but they were seen as not as quick to adopt new technology (embedded), depended on a new platform that had not yet seen widespread adoption itself (WebAssembly), or were not a particularly lucrative industrial application that could lead to funding for the language (CLIs). I drove at async/await with the diligent fervor of the assumption that Rust’s survival depended on this feature.
... Many of the most prominent sponsors of the Rust Foundation, especially those who pay developers, depend on async/await to write high performance network services in Rust as one of their primary use cases that justify their funding. Using async/await for embedded systems or kernel programming is also a growing area of interest with a bright future. Async/await has been so successful that the most common complaint about it is that the ecosystem is too centered on it, rather than “normal” Rust.
I don’t know what to tell users who would rather just use threads and blocking IO. Certainly, I think there are a lot of systems for which that is a reasonable approach. And nothing in the Rust language prevents them from doing it. Their objection seems to be that the ecosystem on crates.io, especially for writing network services, is centered on using async/await. ...
None of us can control what everyone else decides to work on, and the fact of the matter is just that most people who release networking-related libraries on crates.io want to use async Rust, whether for business reasons or just out of interest. I’d like it to be easier to use those libraries in a non-async context (e.g. by bringing a pollster-like API into the standard library), but it’s hard to know what to say to people who’s gripe is that the people putting code online for free don’t have exactly the same use case as them.
Well, that says it. Rust has pivoted to web stuff. For which Go is probably better suited. Good-bye, Rust as a systems language, or for game development. Younger programmers will probably still be seeing C/C++ buffer overflows in 2050.
The technical problem is that pure async, like JavaScript, is fine, and pure threading, like classic Rust, is fine. But the combination is awful.
Rust is just a language - and it's just as suitable to deep embedded and general system programming as it's ever been. The real difference is that it would be insane to write network services and especially web stuff in C/C++, whereas Rust makes this quite feasible. Why are you surprised that web folks are interested in doing that?
It does not say that though. It says that the people doing the work value doing it with async.
Are game developers gonna have that much trouble calling `tokio::runtime::blocking_spawn`?
[1] https://docs.rs/tokio/latest/tokio/runtime/struct.Runtime.ht...
The end of the post specifically lays out a space for a "less systems Rust" that would make these features nicer, but that cannot exist in Rust due to its strong commitment to being as zero-overhead as possible.
I don't know what it is about this feature that leads to everyone grandstanding all the time. It's incredibly frustrating.
Anyway to get back to your point, I think this is why so many people complain. It's either go or rust and neither is ideal so one way to deal with the problem is attempting to shape rust into this ideal.
Right? Where can I find a language with no garbage collection, nice sum types and no over complicated async syntax? Where? Nowhere. This is what people want.
Even if you have to pull in a dependency that uses async (though if you are not doing async stuff yourself, why would you??) you can trivially wrap it up with `block_on` and move on with your life.
I actually think what you want is green threading and garbage collection and the language I described at the end of my post, but you've sort of ideologically decided garbage collection is bad for whatever reason.
Rust obviously has a zero cost objective. I'm not directly talking about that.
Austral: https://borretti.me/article/introducing-austral It doesn't have complicated async syntax, because it doesn't have it at all :) https://borretti.me/article/introducing-austral#fn:async
Hard, hard, hard disagree. Oh my god, I disagree so much.
We're using Rust in production in our microservices stack (Actix), and we're developing animation systems in Bevy.
The only places we use other languages are Python for pytorch and Typescript for rapid UI development. We do have a bit of Rust / WASM / Typescript interplay though.
Rust is turning into a truly full stack, cross-domain, cross-dicipline language. And that's powerful. It's something Python had going for it (scripting, web, numerical, etc.) when it got picked up for its massive adoption. Rust could totally replicate this.
Rust rocks. Its type system, package manager, and threading / async / memory model kick ass.
I wouldn't choose Go for anything new where it wasn't already entrenched.
Rust is only getting started. It's a phenomenal language and the most exciting "new" language to introduce to problems.
Rust is a blissful experience for those that want and enjoy it.
I think the focus should be entirely on this line:
>The other domains were - and are! - viable long term opportunities for Rust, but they were seen as not as quick to adopt new technology (embedded), depended on a new platform that had not yet seen widespread adoption itself (WebAssembly), or were not a particularly lucrative industrial application that could lead to funding for the language (CLIs)
The problem is "web stuff" was the only domain that could realistically fund the development of Rust. This is further supported by Mozilla dropping the Rust project with all it's funding issues and the slack being largely picked up by companies like Amazon.
The systems/game languages are still possible, but if the Rust team hadn't focused on serving the needs of an industry that could ensure it's longevity, there might not be any Rust today to say "Good-bye" to.
And I think that's fair. It's unreasonable to expect a project to prioritize your needs when you are unable or unwilling to fund the project. The is also true of the 'crates.io' async problem - the companies that are ready to adopt Rust and pay developers to write Rust are companies in which async networking is a big deal. If there are other libraries that need to exist in any other domain, well someone needs to be paid to write them and it doesn't look like there are very many entities that exist to take up that challenge.