A four year plan for async Rust
without.boats
without.boats
I for one am pretty satisfied with async rust, and I'm excited for the stabilization of async-trait. I'd love to see some of the improvements discussed in this post come to fruition. Generators in particular are something I've found myself wanting on multiple occasions, because writing custom Iterators is relatively complex.
The point about return-type notation is really interesting. Once async-trait is stabilized, we'll probably go through and rip out as much usage of the `async-trait` macro as possible, so I'll be curious to see how often we run into the issue described there. I also really like the idea highlighted in the blog post of adding async sugar to function closure types as part of expanding the support for async closures, e.g.:
where F: async FnOnce() -> T
// rather than
where F: FnOnce() -> impl Future<Output=T>
Personally, the lack of good support for async in closures remains one of my only issues with async, just because we often write code in a more "functional" style, and whenever we're dealing with complex async stuff we often wind up having to drop out of that.For what it's worth, Dart has had synchronous and asynchronous generators (including `await for` statements) for as long as its had async/await. They are neat features. I've definitely written code using synchronous generators that would be hard to manually transform into a custom Iterable implementation.
But they add a large amount of complexity to the language implementations and it turns out are very rarely used in practice. Here's a quick scrape I did of the most recently published 2,000 packages on our package manager:
-- Style (64317 total) --
59409 ( 92.369%): normal =================================================
4842 ( 7.528%): async ====
42 ( 0.065%): async* =
24 ( 0.037%): sync* =
Async/await is clearly pretty useful with there being one for roughly every ten normal functions. But generators and async generators are barely used at all.Rust might be a in a different situation because efficient concurrency may be much more important in a systems language, but it's not clear to me that those same features carry their weight in Dart.
Stops when it's got all the keys.
First time in 5 years writing typescript I've actually used a generator.
The other to me is some of the semantics of generators are not well thought out, for instance, you have to call `.next()` _twice_ in order to get the first `yield` value, and how arguments to `.next()` should be used correctly is opaque. This combined with the fact `yield` doesn't follow the same lexcial scope as `this` does (IE, you can't have an arrow function yield if it was enclosed by a generator, unlike `this` which can be used inside an arrow function as reference to its enclosing scope. This would make generators far more useful IMO).
Combine all this with the fact that most frameworks don't support generators natively for things like components (but they are starting to accept Promises / async functions) you end up with a relatively niche feature
But for a feature like generators which is, I think, essentially user-facing syntactic sugar, I do think a count of usages is a pretty good measure of usefulness.
It's probably also worth noting that my data here is from published packages, which tends to skew towards reusable libraries and away from application code. So, if anything, I would expect this to be an over-count of their use if they primarily made libraries/frameworks useful on behalf of application developers.
One thing to remember is that Rust is a bit weird here; often times, features like this are "sugar" in a sense, but that sugar lets you write safe code, whereas the non-sugar version would force you to write unsafe. This means this kind of thing is a larger win in Rust than it would be in other languages.
I also don’t think that async-drop as quite as necessary or desirable as it may seem. I really wanted it at one point and then I realized that it’s just too tricky. It reminds me of how File calls sync_all on drop but ignores the error - ultimately, “drop” is just a really tricky place for anything complex. I’d rather see a linear type, like:
struct LinearFile {
fn close(self) -> Result<File, std::io::Error> { self.sync_all(); File { inner: self.inner } }
}
Where Linear can’t be implicitly drop’d, you have to .close() it, get a File, and File can be dropped. Hand wavy and not necessarily a good implementation but hopefully this is getting the point across. This would be preferable to shoving more into a drop impl when drop is such a constrained interface.Or maybe add a try_drop(&mut self) that will run implicity but also ? implicity. I don’t know.
I guess the point is that I’m not sure an async drop can ever be worth it.
TBH I feel like ~30% of people's complaints about async are solved by:
a) Encouraging a sync + async API in libraries
b) Adding `block_on` to the stdlib, which will help with (a)
People mostly seem to care (and imo this is stupid but whatever) about using async in sync contexts when they don't want to, and they don't seem to know that you can just block.
struct SyncThing {
async_version: AsyncThing
}
impl SyncThing {
pub fn sync_api(&self) {
block_on(self.async_version.sync_api)
}
}
If this were really such a huge problem I think we'd see more PRs. I get that there's some survivorship bias here, but still...I never thought about it but, my only problem with C's manual destructors is that the compiler doesn't check that you remember to call them. If we accept a language like Rust that does sophisticated static analysis, it doesn't need to delete things for you, it only needs to be sure you have a plan to delete everything.
That's interesting.
Pretty much, yes.
> I never thought about it but, my only problem with C's manual destructors is that the compiler doesn't check that you remember to call them.
Right, so in this case the compiler would force you to consume the value somehow. This is "linear" typing.
I began writing a large application, and noticed that many of my libraries only offered async versions, and the promise was quite appealing -- not having to worry about threads or concurrency as long as I followed certain rules. What I ended up with was an incredibly slow application because of the limitations of Rust async, and the runtime. All of my I/O got pushed through one thread (with tokio), and that, plus scheduling overhead became my bottleneck. Debugging this was a nightmare in writing my own tracing tools.
20/20 hindsight, I would not write my code to be async, and would just prefer threads. I'm really not sure how / why async took off the way it did.
But you can just `block_on` your futures if you want and not think about it at all.
Async is almost useless when you’re not incredibly I/O bound. But many people these days are because the web ate the world.
C did and does have a poor almost-anything story but it doesn’t and didn’t really matter because C hasn’t been relevant on the web since 1995 or so.
I'm not sure that's true. The async/await model is about representing a state machine in imperative code. Not all state machines can be written imperatively, but when they can be, it is often clearer than writing the state machine manually. I/O is the most common scenario for such state machines, but I can see a few other OS scenarios where you might want to use async/await instead (e.g., process management is probably better represented with async/await).
It does, once you have async you can multiplex n coroutines on m cpu cores, meaning you can throw libthread out of the window.
> C hasn’t been relevant on the web since 1995
Do you know nginx, a core web technology is written in C? Linux is also C. Don’t forget that at the end ov every web request are syscalls and hardware.
(Your response isn't entirely accurate though: it's using a thread pool whose size is configurable, not one thread per spaw_blocking)
I don't understand why some people think async is difficult to write. It's difficult if language doesn't have good support for it, but rust does. I remember async ruby was pretty weird to write initially.
EDIT: People talk about swapping executors etc but it seems like 99.999% of the programming applications you don't need anything like this, nor does something like Python supports this anyway and we're all fine with it.
It depends on what you're comparing it against. If you are comparing Rust async against other languages, then the major difference is Rusts borrow checker. In every implementation of async, you effectively move whatever state is needed off the stack and into a separately allocated memory area. In Rust that separate area is a closure.
This interacts badly with Rusts borrow checker. The borrow checker needs to know the lifetime of any object you deal with. It has two "base truths", by which I mean life times it already knows about that you can derive other life times from: static and the stack. If they don't suffice you have to handle life time management yourself and at run time using Rc or Arc or something. Being forced to do that complicates your types and slows the code down. The root cause is async in Rust removes one of those two base truths: the stack. So now you are forced to write that ugly manual life time management code far more often.
This is unique to Rust. Every other language I know of that implements async has garbage collection, so while it remains true they also move stuff off the stack it doesn't change anything. You still use the same types, and apart from sprinkling async's and await's here and there and indenting your closures, your code remains the same.
This is also why I think green threads are a much better fit for Rust than async. Under the hood green threads and async are very similar: they are both ways of doing event driven I/O. Their performance characteristics are near identical. The main difference is while async forces you to move your state to a different area, green threads you do it as before and store it on the stack, just like normal code. In fact green thread code looks identical to normal code. The only change you have to make to convert some code to green threads is change the name of the I/O calls to use non-blocking versions (which is something you also have to do with async, of course). But since the stack is still available all those fights with the borrow checker async creates go away, as does all the extra syntax async requires.
It doesn't come for free of course, so the run time performance of green threads and async is not absolutely identical. In async every task shares the one stack, whereas in green threads they each get a new one. This creates some extra memory overhead, chews up considerable address space (which normally isn't backed by memory) in order to protect against stack overflow, and it costs a bit more to set up a stack. But once a task is setup green threads are going to be bit faster you aren't moving stuff and and off the stack and you don't get hit with those additional run time life time checks you were forced to introduce for async. Mitigating green threads overheads somewhat, a process that is handling 1000's of concurrently connections is unlikely to be running one a machine that is memory constrained, so the extra memory probably doesn't matter. (In reality a Raspberry Pi with 4GB of memory can handle 1000's of stacks.) And it's likely to be a 64 bit machine where address space is nearly free, so the "considerable extra address" space also doesn't matter.
Still, I can think of once place it does matter. You can have two styles of generators: ones take the async approach and ones that take the green thread approach (ie, allocate an extra stack to each generator). Rust nightly does have generators and they currently take the green thread approach(!). But, generators tend to be short lived. You tend to use them to iterate over an array, rather than serving a web request (the typical use a task is put to). That means the overhead of creating the stack for a green thread isn't a small proportion of the time a long running task takes, but could well dominate the time it takes to iterate over a small array. Thus async is a much better fit for iterators.
Currently Rust has this arse about: it has async for long running tasks, and green threads for generators in nightly. With this announcement it looks like this will be 1/2 fixed. Great!
> Every other language I know of that implements async has garbage collection,
Small note: C++ also has it these days, and does not have GC. You are (as far as I know) left to your own with object lifetimes, as with non-async/await.
The issue generated a lot of hot air at the time, far more than I have time to read, but I think I have the gist of it.
The performance issues were due to an implementation choice. All green thread implementations I'm aware of hide the code colouring event driven I/O introduces, and Rust did the same. The hiding has to be done at run time. That translates to every I/O call has at some point chose the blocking or non-blocking implementation. This is typically done using vtables (dyn in Rust), but whatever mechanism is used, it introduces runtime overhead. Worse, it slows down everything - including code that doesn't use green threads. It gets radically worse if you try to hide C calls blocking.
Async has the same issue, some solves it by making coloured code the programmers problem. If they had of made the same design decision for green threads the performance issues go away. We don't have to speculate about that. There is a green thread crate out there called "may", then has independent benchmarks covering it, async Rust implementations, and other languages. "may" beat everything at one point, but it's a very noisy benchmark so the only conclusion I would draw from it is that "may" looks to run at the same speed as async.
As for the rest of the issues: they revolve around needing a separate stack. I tried to cover the trade-offs above. Summarising: green threads lead to simpler code that's easier to write and theoretically could run faster than async, but have larger setup overhead and use more memory.
Just reasoning aloud here. That looks to be another similarity with green threads and async. Async also requires a big fat runtime that isn't part of the language. Instead you have to pick an async colour - such as tokio. That's pretty much what happens with green threads now. You have to pick a runtime such as the "may" crate, which forces the programmer to choose yet another colour.
So many colours, yet they are all just event loops underneath. Colours cause fragmentation, fragmentation is the mortal enemy of reuse.
It seems like the very least the language could do is provide a set of trait's for event driven I/O that mirror the existing I/O library in std. Then the library writers wouldn't have to colour their code by using a particular event loop implementation. I suspect it's easy enough for green threads or async, but accommodating both styles of event loop would be hard.
I do completely agree with you that it would be great for std to pull in more traits from `futures` and elsewhere to allow it to be easier to write code against different runtimes! I think part of the intent was to get Futures out, see how they get used, and then to go from there. Hopefully we're getting into the "go from there" stage now, which is some of what this article gets into
Sometimes I feel like the only way to get really good async I/O is to rewrite the whole damn thing to use io_uring. Avoid the blocking syscalls as a whole.
("Rewrite" because this changes e.g. read buffer management. The fundamental APIs will have to change, to enable the performance gains.)
The current async approach requires the compiler to statically calculate a fixed size for each invocation context. Except when it can't, in which case you have to dynamically allocate space for each invocation. That same logic could be used to transparently optimize stack creation, opportunistically avoiding both a pessimistically large stack and the costly guard pages. The compiler could even choose to instantiate the generator stack on the caller's stack, just as today.
async Rust optimizes the common case but completely neglects the hard case. In theory Rust users should be able to have their cake and eat it too, especially given that the difficult static analysis work already exists to support the current async model.
(I've made this point before and feel like I may have forgotten some counterpoints. I apologize in advance if that's so.)
I’m guessing your code is blocking. You absolutely cannot do blocking code in async in any language, unless the blocking is super quick. No blocking io.
async code should yield great performance if you are doing everything the right way. Yes it is single threaded but either run multiple threads in rust or run multiple instances of your program with systemd.
That applies same to rust, Python, javascript, any async language.
I wrote a prototype message queue in Rust with Actix and got 7 million messages per second via http.
Async I/O is all about multiplexing work on few CPU cores efficiently, and multi-threading is still required. Python or JavaScript are seriously limiting environments and should not be given as an example when we're talking about a language that does 1:1 scheduling.
You cannot, (should not), do blocking IO in async in any language.
The language provides libraries to do non blocking IO.
What is unclear about that?
Multi threading has nothing to do with async, except as a way to run certain things in an executor thread to avoid blocking.
Also I don’t really know what you mean by async IO…. even in the languages you reference I would guess IO operations are all run in a single thread.
On platforms with 1:1 scheduling, of course you can. Blocking I/O executed on another thread, with a callback to execute when done, becomes async I/O (from the user's PoV).
Ofc, when we talk about async I/O, we also refer to the kernel APIs being used, such as select/poll/epoll/io_uring. Say, working with Epoll is usually done via a single-threaded "event loop", but that only listens for the operations possible for an open file/socket. The read/write operations are still potentially blocking, so for efficiency you need multiple threads. The dirty secret is that async I/O, as implemented by Linux, isn't actually fully async.
> Blocking I/O executed on another thread, with a callback to execute when done, becomes async I/O (from the user's PoV).
That's not what we're talking about when we discuss languages with async I/O, though. That's just bog-standard synchronous I/O with multithreading.
> The read/write operations are still potentially blocking, so for efficiency you need multiple threads.
That doesn't actually follow. The entire point of language-level async I/O is to be able to continue doing other work while waiting for the kernel to finish an I/O operation, without spawning a new OS thread just for this purpose.
My code wasn't blocking, but it did a very small amount of computation, locked on what was typically a highly contended (async) lock, and then dispatched a bunch of blocking I/O operations. Each of these turned into a context switch under the hood.
There were parts of my code that were computer intensive, but determining that without being able to use my normal tools was a pain, and I had to come up with heuristics of whether or not to dispatch with spawn_blocking, since the call had significant overhead.
An abstraction that was meant to simplify program resulted in me spending a lot more time staring at it.
I think one difference with Python is that the async implementation doesn't try to hide or abstract I/O away from you. If you're doing I/O, you know you're going to block, and thus you're forced to acknowledge it by spawning it on another thread.
It would be interesting to see the code or know which libraries you used (or even just what type of application you were building).
I built a heavy app using async-std and then rewrote it to threads just to better control what was each thread doing. Without my additional scheduling rules (which brought a lot of other benefits and helped performance in other ways), the performance between async and not async was close, with threads being marginally faster.
In all my use of Tokio in the last few years, I never heard of such a thing.
In tokio::fs::File [1] it calls spawn_mandatory_blocking to do file writes, I assume this is similar to spawn_blocking [2] which sends a task to Tokio's blocking thread pool. That thread pool is supposed to max out at 512 blocking threads [3], unrelated to CPU core count.
Tokio's TcpStream appears to be built on mio's TcpStream. I didn't dig deep into the code for this, but it doesn't just call spawn_blocking, and I'm assuming on Unix it ultimately registers the socket with epoll or equivalent, so it never blocks a thread to do socket reads or writes.
Could you share more about your application?
[1] https://docs.rs/tokio/latest/src/tokio/fs/file.rs.html#682
[2] https://docs.rs/tokio/latest/tokio/task/fn.spawn_blocking.ht...
[3] https://docs.rs/tokio/latest/tokio/runtime/struct.Builder.ht...
Swapping the executors out should absolutely be a feature, and the traits should be portable, but a way to start fixing the situation beyond the great suggestions in this proposal is to acknowledge that std and no-std users are different, and std users are often developing applications and prefer sane defaults.
Regardless of the quality of this idea, it isn't going to happen: neither the Rust Project nor Tokio (in my understanding, to be clear I am not involved in either) want this to happen.
> Swapping the executors out should absolutely be a feature, and the traits should be portable
I am not fully aware of all of the details here but there are significant problems when it comes to actually getting this done; I don't think anyone is ideologically opposed, but there's a lot of practical considerations that make this difficult, in my understanding.
I think they'd like the choice to use a different one, but they'd rather just having something available with async traits that trend toward opinionated. The executor issue is a huge problem, because the ideology of "zero opinion" on executor coupled with "ease of adoption" are completely at odds. I don't think the current trajectory will ever resolve nicely.
Whereas my own sense of the Rust philosophy is supposed to be one of zero-cost abstractions (when possible), and for the language to provide nuts and bolts that I can assemble myself. My interest is in systems eng. I don't want to be tied specifically into the systems-eng choices that Tokio happens to have already made, even if they might be good ones. I want the ability to choose. It's not healthy, in my opinion, for one entity to dominate choices like this.
If this can't be resolved, my sense is that people who have the same impulses as me will choose to either not use async at all, or move on from Rust.
I'm not even sure what I want in Rust. All I know is that people have a legitimate gripe being sold on Rust async for high level work, having it ergonomically fall much flatter than it should. It's very hard to satisfy everyone, which they are learning day by day.
I do wish for the async story to mature a bit more so the implementors have a chance to show the community how good it can be. But the project needs to consider its messaging to the users it has courted over.
Not everyone that writes network services is "higher level" or "web".
That's not a good situ. I'd as soon rip out async, and optimize the concurrent multithreaded situation myself than be tied down that way.
FWIW my day job is on embedded Linux, small systems that sit in tractors. We don't do async Rust. It's all actor-style communicating components. And it works pretty much fine.
However, I don’t think the situation can be improved by putting an executor in std. That will even more strongly make everyone stick to the standard one.
The problem isn’t that it’s hard to pick an executor (you can pick tokio without thinking). The problem is that when someone has a legit reason to use a different executor, it’s hard to avoid dependencies using tokio, and it would be even harder to avoid dependencies using a built-in executor.
What the stdlib actually needs is the proper set of traits/facades/whatever to interact with the current runtime. Just like they did with the Future trait. And add the handful of traits that go with them that tokio, futures, smol, etc have: AsyncRead, AsyncWrite, Stream, Sink, et al
Every async-focused dependency that is relevant effectively mandates Tokio. And I've been ridiculed in public and private forums for suggesting that there should be a way to write libraries generically so that they don't have a runtime dependency. (Right now, services as basic as locking and task spawning are coupled to tokio, at least, and people use them all over.). Major new things are being made, all coupled to tokio.
And I think this is a shitty situation for a whole bunch of reasons. And others (looks like you, too) agree, but it feels like there's just no way this is going to get fixed.
Like, this article here, it doesn't even mention this as a concern? In the context of a big picture discussion of async over the next 4 years.
So I don't foresee any progress on this front, and it makes me want to just rip async out of my code entirely.
The first is interfaces to represent async versions of existing core sync traits (that's the AsyncRead/AsyncWrite/Stream/Sink/etc. you refer to). What makes this somewhat awkward is you're providing these traits without an implementation of them in the standard lib.
The other thing that's missing is something like GlobalAlloc. You need a generic executor runtime interface to handle some of the executor things you need like "schedule new task".
In general, I think there's a class of features where you need to have just one global (really, process-wide, not just library-wide) provider of some service, and it may be worth having language features to provide this functionality. Memory allocation is one specific area; async executors is another topical area. But you can also throw in stuff like signal handlers or logging or service providers or error handling or tracing features, etc.
I think it's a great idea. By blessing something that's not tokio you give a good incentive for library authors to test against more than one executor. And by being so far from fully featured pollster is never going to "win" so blessing it doesn't appoint a winner. And it gives an obvious solution for what people who don't want to use async in their code but do want to pull in an async library should do.
edit: it was without.boats in this post: https://without.boats/blog/why-async-rust/
The problem with async Rust is that the async idiom is new, and its interactions with the rest of the software ecosystem not particularly well understood. This makes async support glaringly different from the rest of the language.
I'm glad the designers seem to be taking a step back and reconsidering how everything fits together.
This is simply not what is happening here. The post is clear that this is about filling out and finishing up the plans that were laid down when async was initially designed, not changing how things fit together.
Also it’s simply not true that async wasn’t fully understood or carefully evaluated. It took years to design, and then bikeshed every detail, to the point people involved were burned out. It had multiple prototypes, and an early callback-based implementation used by hundreds of libraries, in production, for over a year. It’s probably the most thoroughly designed and tested feature in Rust’s history.
Editions happening every 3 years unfortunately undo this, and there is a pressure to land changes before the upcoming edition.
That’s how it works in javascript and as a commenter mentions in Python too, right?
> On the other hand, there was some speculation about making “await patterns” that destructure futures and then somehow making that work here; I think this would imprudent and leaving await as an expression, and for await as a special expression for handling AsyncIterator, is the most sensible choice.
A generic block_on in std would be a footgun for new programmers I believe.
That being said, I am also a fan of embassy when you have different design constraints and goals, and consider the fact that is is able to exist and be successful is a massive testament to the design of async Rust.
(We also use async Rust heavily further up the stack, and have some issues with it, but they tend to be disjoint from the way that this is talked about online.)
Example from a current side project, since it's not work-encumbered: A wifi-enabled clock light for my 5yo. It has a task that sits there and every hour pings an SNTP server, updating a (mutex-protected) global with the time state. It has another task that listens for a telnet session for various control signals - which also updates a mutex-protected global config state. And it has a task that spins doing LED effects.
With embassy/async, I can write each of those as a separate task, without paying much attention to what gets invoked by interrupt handlers.
(This particular one is an rp2040, but I use it on stm-based systems as well).
I feel like this is kind of analogous to the discussion of threads-vs-events as a mechanism for structuring code vs. threads as a mechanism for achieving parallelism. :-)
Edited to add: Or, perhaps an alternative view of what I'm doing is that I'm using the embassy runtime as a really lightweight alternative to an RTOS, since I mostly haven't met an RTOS I don't want to throw across the room. An argument against what I'm saying here is, "well, use Hubris as your embedded OS and then you can have tasks and they can be synchronous" - which seems entirely fair.
This is also a good design! The primary designer of Hubris also has a project that works like this: https://github.com/cbiffle/lilos
My complaint is I don't find the async ergonomics intuitive; it feels like a layer of misdirection. And, the viral character.
No.
This is how yo do: you look what executor your dependency is using (most likely tokio) you add it to your cargo.toml (at zero cost, since it's already there in your dep) and then you wrap the async library calls in `block_on` and call it a day. You don't need to change a single other line in your project.
Now if only the other warts were fixed, like the particularly poor compiler errors when there's an issue within an async function...
If it matters to you, you don't have the same priority as your dependency's author anyway and probably shouldn't be using it in the first place, and it has nothing to do with async.
[1]: https://docs.rs/reqwest/latest/reqwest/blocking/index.html
So many people complain about Rust being hard. What on earth are you people on about? This isn't at all a hard language.
I'm so sick of this "async is hard" / "Rust is hard" meme. No, it isn't.
Maybe it's inconvenient if you're used to slapping packages together and calling it a day. But I don't think the majority of us do work like that.
And it's not hard.
I'm sick of people perpetuating this about the Rust ecosystem. It's not a month that goes by that there isn't some article badmouthing the language and its features.
Stop it.
We want our language to gain support in enterprise so we can get paid to write it as our day job. These articles and negative comments are not helping.
The amount of "async rust is impossible to use," "async rust is fundamentally broken," etc. commentary that comes up on HN is absurd. I have trouble squaring my own experience with it, and it makes me think people are either complaining just to complain or that they haven't worked with it in anything other than toy contexts. I have a hard time imagining how difficult it would be to maintain our ~100k line rust codebase (http proxy that processes requests & responses, makes DB queries, and makes http requests to other services) and keep any kind of predictable performance by spawning threads, using channels, etc.
Like, not trying to be elitist, but I feel like people are missing the benefits for the sake of piling on or something.
You didn't make any arguments here at all. You can check my other comment for detailed reasons why async in a language is not a good approach to concurrency. The overview is that is just isn't a holistic solution and the only time it will solve someone's problem is if they have extremely simple concurrency needs in the first place and those never scale or change.
a) They often focus on problems that, at least for me (and many others) are not significant. For example, acting like writing "+ Send + Sync + 'static" is causing your hands to seize in pain. Or they ignore that you can "block_on" a future.
b) They're then often interpreted by people who have very little context on async or rust at all. Look at the initial comment of "Async is a wart", it's idiotic. Look at how stupid the comments section is, how people are not talking about literally anything that boats wrote.
c) They make propositions that aren't very helpful. They mostly say "it was a mistake!" or "we don't need multithreading" or "we want TPC".
Contrast with Boats' post, which is far more contextual, far more historical, and proposes actual solutions and future work to be done.
It's quite frustrating to see the same sorts of complaints over and over again, especially when they're often not super great complaints to begin with.
That's no way to do good work.
The specific issue is the context switch. The compiler, with async is able to be much more performant than the context switch that the OS provides.
Depending on the kind of application you're writing, this is either splitting hairs or very, very important. Applications that handle many (hundreds, 1000s,) of concurrent IO will have a noticeable performance improvement using async versus the OS's context switching.
But, there's a more important thing to consider: "async" in a programming language communicates that a method can block. It allows the caller to start the call and do something while it's waiting for the result, without needing to get into the weeds of threading. Purely relying on the OS for context switching means that it's hard to know what methods block.
> without needing to get into the weeds of threading
A thread is the exact concept needed to describe that and preserve “if else then” sequential code.
> 1000s of threads
Linux handles thousands of threads just fine. If you’re using a scripting language like python, context switching is the least of your performance concerns.
I suggest looking at C# and Javascript that implement async very, very well.
The difference is that with conventional threading semantics, when you join a thread, it doesn't return a result. You still need to write some kind of "thing" to get your result from the subthread to whatever's waiting on it. (C# also provides a less-well-known BeginInvoke mechanism which is somewhat cleaner than join.)
In contrast, the promise (Javascript) or task (C#) has a result. Instead of joining a thread, the await keyword gets the result, just like calling a method.
> Linux handles thousands of threads just fine.
Yes... And no...
It doesn't matter what OS you're on, each thread needs its own allocated stack space and has the overhead of context switching. "async" optimizes that by putting data that would normally go into many different stack spaces into the heap and jumping around among concurrent operations without the context switch.
Again, depending on what you're doing, that's either splitting hairs, or really making a tangible improvement. But you can't argue that more allocated stacks, and more context switches, is faster than doing it in process. At that point you're arguing with fact.
(I really struggled with async rust, so I'll admit that I don't remember the name of the type that represents the promise.)
I am familiar with how these are implemented. My opinion is the same. Yes you can implement slightly lighter weight threads in a language itself.
> It doesn't matter what OS you're on, each thread needs its own allocated stack space and has the overhead of context switching
This is a quantitative question. I know what it does, the question is how much slower it is than whatever async construct you want to use.
See e.g. this recent clip from one of the engineers on Project Loom in which they argue that the context switch is relatively low overhead: https://youtu.be/07V08SB1l8c?si=i0v9w90Kb_M0I4gP&t=966
Is this something that's hard to do on Linux?
Both of these APIs are set by you, the user. You can choose how big to set your stack size, it's true. However, what value do you actually set? Too high, and you're still using too much memory, though admittedly less. Too low, and you either need to accept death by stack overflow, or runtime detect this case and fix things up.
With async/await in Rust, the compiler can statically see how large the stack size needs to be. Each "thread" will have a perfectly sized stack. No user intervention required, no fiddling with settings.
This is a solved problem - it's "just" the longest path in a call graph where each node is weighted by its function's stack frame size.
> With async/await in Rust, the compiler can statically see how large the stack size needs to be. Each "thread" will have a perfectly sized stack. No user intervention required, no fiddling with settings.
There is nothing stopping a compiler from doing the exact same analysis for a sync function at compile time (barring exceptions like varargs or deliberate recursion). They just... haven't, for some reason. It's a shame that it took until Zig for it to be addressed.
EDIT: to add a specific example, postgres. I understand from sfackler's perspective why it makes sense to maintain just an async library and a sync wrapper around it. But from the user's perspective there'a absolutely no reason that all of tokio should be required to talk to their db.
Would you also advocate for the kernel to deprecate blocking i/o?
And all you need to do to not deal with async is wrap it in one of the many different block_on implementations. Yes, that's less ideal than not having to pull in these dependencies; if avoiding those dependencies is mission critical for you, maybe yours could be the enterprise that pays for the blocking IO postgres library.
I've never saw anyone discuss using poll(2) style concurrency here which is interesting but also kind of sad.
With the mio crate for example, it uses the operating systems "select/poll" interface. On Linix this is epoll(7), but other OSes have their own interface.
"poll" style concurrency is a lot less complicated for me. The general idea is to register IO handles such as sockets into a data structure and pass it into the kernel. The kernel will emit events whenever a handle is ready to do something. There is the option to go to sleep and block until one of the handles are ready but it's not required.
"async" is an abstraction on top of this. The async runtime is handling the event loop and other stuff for you which is very useful but also a little bit complicated.
I am not in any way against "async" style concurrency but I think a lot of libraries could avoid pulling in a runtime and still be concurrent.
DB drivers, http servers, runtimes and so many other complex components are preferences of high skilled developers and not some objective market choices. Once its developed most application developers have to use it however user unfriendly it maybe.
If there's not enough community interest for there to already exist a solution for your problem, you really only have three options: write it yourself (including opening a PR to add it to an existing implementation), pay someone to write it, or switch languages. If none of these are options, you're just out of luck, and complaining about it is unlikely to do any good. I'm not saying you can't complain of course, but I am saying that you're more likely to get what you want via another path.
But regardless, I don't think it's an odd argument. Indeed, no one is stopping anyone from writing a new, low-level, memory-safe programming language! That's why Rust exists! Turns out the market was huge! Writing a new one now would be easier because of all the great ideas that Rust helped to prove work at scale. We're seeing this already with the success of languages like Zig, which makes different tradeoffs than Rust did and seems to be finding success in slightly different niches.
And sort of disproving your point, we use postgres, and I can think of three different implementations of postgres drivers offhand (sqlx, tokio-postgres, and diesel). I'll also note that the author of tokio-postgres also publishes https://docs.rs/postgres/latest/postgres/, which is not async! It's impressive how many options we have for such a complex thing in such a young ecosystem.
Async, in the sense from the article, and non-blocking aren't synonyms. Not using Async doesn't imply blocking.
(Also, in case it isn't obvious: I wrote the article in question.)
More generally than the concurrency operations I described, are any sort of event-loop or state machine, of which Async is one example.
I eventually rewrote the whole thing with rusqlite, but apart from being non-async, I found the API much less ergonomic to work with. But at least Rayon worked exactly as I expected.
I just wish everything would at least still support blocking IO as a common denominator, so there's at least a baseline that is known to be possible. 99% of percent of the time I do not need benefits of async, because I write small software for relatively small, but still real use-cases. And using Rust ecosystem was easier just a few years ago than it is now for these use-cases, as blocking Rust was not yet relegated to be a second class citizen, replaced by an immature and fractured async.
I used to write Rust web services (professionally) before async and it was definitely not easier for me. It was way harder. I'd end up with accidental hanging because some socket didn't have a timeout set on it properly, it was extremely leaky (why the hell am I talking to raw socket APIs just so that I can read from S3?), and it sucked compared to async. We've had radically different experiences, somehow.
I don't know. https://doc.rust-lang.org/std/net/struct.TcpStream.html#meth... etc. were there for a while. Though I think that deadline-based timeouts would be way better to have. https://www.reddit.com/r/rust/comments/8b5krv/the_case_for_d... . Probably could be built around existing primitives. These things were never built/popularized because the community just jumped on async like some silver bullet.
I agree that building heavy duty networked services got easier with async, but at an expense of fracturing the ecosystem and dragging everything else into an MVP feature, which made things worse (at least in some respect) for other things. I personally don't do much web servers, and when I do I can just spawn lots of threads, terminate TLS with nginx anyway.
Again, I don't mind async on its own. I think it's great for what it is, and is useful when it's useful. But for decades tons of the web was built in Java, Python, RoR without async IO and it worked just fine. And I didn't have to play IO-type sudoku, and could expect that the most basic, native blocking IO is well supported, and not just an afterthought/wrapper.
This isn't true. I'm writing an application right now that is mostly sync, but has a small amount of async code in it. It's easy to spin up a tokio runtime which can run inline on the current thread. Then use it to evaluate a Future.
tokio::runtime::Builder::new_current_thread().enable_all().build().unwrap().block_on(async {
println!("I'm async!");
}); smol::block_on(async {
println!("I'm async!");
});
(I thought tokio had a helper like this too but could only find `tokio::runtime::Runtime::new().unwrap().block_on(async { println!("I'm async!"); });`.)Rust is not a batteries-included language like Python. There are lots of libraries that are very commonly used in most projects (serde, thiserror, and itertools are in almost all of mine), but this is a conscious choice. They say in Python that the stdlib is where projects go to die. I'd rather have the flexiblity of choosing my dependencies, even for stuff I have to use in every project.
So you are stuck with smaller, less battle tested products if you'd rather not pull in 100+ crates of dependencies that are doing nothing but inflating the build times and file sizes (for your particular usecase).
Example: reqwest vs ureq
Increased build times are not great but holy shit the way people talk you'd never know that that's the actual trade-off here, an extra 3 seconds on a clean build.
smol::block_on(async {
println!("I'm async!");
});
if smol is an option.But on the other hand, just read the code. It's not complex.
You drill down into a tokio namespace. You make a builder object. You unwrap it. This is idiomatic Rust. It's verbose, but explicit is better than vague. There's no conditional logic. There's no weird type-fu. No macros. There's not even any parameters to supply, other than the Future to block on.
It's trivial to write a wrapper to go from chained methods a helper call.
- create a builder
- run on the current thread only
- enable all drivers
- create the instance
- ignore errors
- then call some blocking async code
Would look nicer if split out onto multiple lines I imagine:
use tokio::runtime::Builder;
let runtime = Builder::new_current_thread()
.enable_all()
.build()?;
runtime.block_on(async { println!("I'm async!"); });I will say that the notion that you must make your whole project async is largely true (except that you can block on futures with all runtimes), but this is more symptomatic of the fact that Rust has hitched its horse to two wagons that don't have clean overlap. It is effectively two languages trying to converge in the middle.
It's equally desired at "web server and above" and "operating system and below", which probably makes it one of the hardest languages on the planet to design (lets not forget it deliberately has no runtime, which makes things even more difficult). Whether or not this is a good idea remains to be seen over time, but they are in largely uncharted territory so I'm prepared to give them some slack on it while they figure it out.
That said, holy has it taken such a long time to round out the async story to make it feel better. I don't blame users not wanting to wait it out because it has felt like an eternity.
My theory as to why C++ has become as a problematic language as it has is because it is able to do it all. And "all" doesn't fit nicely into a single package or paradigm.
I don't like async (in any language), so your post immediately piqued my interest. I'm not that interested in 'webserver and above' so of course that would be where people find async useful.
It is possible that there is no way to unify these two domains into a single language (at least without becoming c++). Although, then again, it might be that if you think about it long enough then the unification mechanism becomes apparent.
- attempts to solve too many problems,
- has significant papercuts that have to be considered,
- has significant runtime cost,
- interacts poorly with other features, or at least leaves significant edge cases open,
- has non-uniform compiler support (although in OSS land only GCC and LLVM matter), and
- creates arcane error messages that are anything but fun to analyze. Especially when templates are involved.
Above everything the fundamental problem that the unsafe parts of C are everywhere, and to master C++ you have to become really proficient at diagnosing problems with them because you will run into them. Only then can you start thinking about fun things like software architecture. Always use a linter/static analyzer.
IME the bad compiler messages tend to come more often from the `async-trait` macro than from anything fundamental to asyc itself. Hoping the stabilization of async-trait helps with that. Backtraces with async are definitely a bit nasty though.
The C++ committee basically evaluates features almost entirely in the abstract, writing papers and discussing privately among a small group of people about things instead of working on proof of concept implementations, getting actual real world feedback on it, or having something along the lines of Rust's nightly where experimental features can be tried out.
Sitting on the C++ committee, we actually demand a lot of proof-of-concept implementation of proposals before they'll be accepted.
It's a shame Rust let itself be distracted by it instead of focusing on refining its strengths and developer experience
those people probably may use another language (java, go, c#) if absolute performance is not critical for them.
Integrating 'async' into a language is not a good approach to concurrency. If you want to run a single function on a different thread it can work. If you want something to run after that, that can work out.
Once you go beyond that, you are building a graph with very crude tools. Then you have lots of problems, including how you handle packaging dependencies when something you want to run depends on data coming from multiple other async sources.
Treating this as a language issue and not a library and tools issue is a huge mistake. Another reason is exactly what you outlined - libraries getting infected with a languages half baked concurrency solution instead of doing what they are supposed to while the user can fit them in where they want.
What actually works is graphs that handle dependencies and data structures for synchronization, but ultimately those need to be done well too.
I'd love to see this in Rust, and I know that it initially had something like this, perhaps its not a bad idea for someone to take a second look at this now that more time has passed.
EDIT: to be clear, that is not me saying we need goroutines or we need a Java Loom equivalent in Rust, simply, the DX of these two examples are far superior to the DX of async Rust today, in my opinion
And Java/Goroutines don't have a keyword because they do things implicitly, they have heavy preemptive runtimes.
I didn't say we need goroutines or we need Loom style asynchronous primitives I am however pointing out, that they did a really good job of making async approachable and feel like you're writing the same code as you would in a traditional synchronous model. Thats the real win.
Its not an easy problem to solve, to be absolutely clear, but its a worthy goal, and if that means it adds marginal overhead initially I think its worth the tradeoff (esp. if you can do something similar to how you can use Rust in a `no-std` build, you could do one without a `async runtime` build, spitballing off the top of my head)
Another problem is that generally speaking, async is de facto used a 3rd party lib, and it really should be a first party primitive that everyone feels comfortable using.
To me, this is what matters
While you are right that DX matters, DX is not the only thing that exists. Design is about balancing constraints and tradeoffs. Rust has other design commitments that preclude using this design, before you even get to DX questions.
(Furthermore not everyone agrees that these things are clearly better from a DX perspective, as DX is a subjective topic.)
I will say, that its clear, to me and a good portion of Rust users, that async in Rust needs DX improvements. Its one of the top features of the language that is used alot and people struggle with often[0]
[0]: To be fair, async in any language trips up developers pretty often, though I think Rust can give someone a particularly bad time. Granted, I have not compared it to C / C++ as I don't do any development in either language currently
To be honest, this is why this discussion always gets frustrating: people demand change that is impossible, and then when pushed for how to accomplish the impossible, they throw up their hands. I do not think you or anyone else is doing it maliciously, but for some reason, on this specific topic, it happens endlessly.
Ideally, the runtime could handle anything written using lower level primitives as to not completely kneecap libraries that need to work with `no-async-runtime` (or whatever you want to call it).
This would at least alleviate a common concern I hear around this, which is runtime bloat.
That seems like a step in the right direction to integrate a unified async runtime with an alternative / better syntax[0]
[0]: I want to note, that C# is the only language I ever worked in that supported async in two constructs. There is the traditional async / await, which is by far the most common. Before that though, there was event driven async programming (with support for background workers and other async features) and that is also still supported, and they can interop with each other (with some caveats)
You created a dichotomy - languages with a native construct for async versus languages that provide this as libraries. But that dichotomy does not exist - both language have native concurrency support through their runtime. That was what I was pushing back against.
> async is de facto used a 3rd party lib, and it really should be a first party primitive that everyone feels comfortable using.
I think this really remains to be seen. Rust has always had a "just pull in a crate" mentality and a "be very conservative about what's in std" approach, and I think the community is overwhelmingly in favor of that. Any major stdlib changes should be taken pretty seriously.
I think 4 years worth of a feature being out in the wild and in use is enough time to start having conversations about what went well and what didn't and how to address that. Its very clear to me, at least, that async in Rust is becoming more and more dependent on tokio. Most of the major async supporting crates support tokio and/or only leverage tokio. Just a cursory glance at the crates registry supports this much.
I'm not against a "pull in a crate" mentality mind you (though, careful what you wish for here, see: NPM / Node ecosystem) however, it is worth identifying when something is becoming / has become / is considered to be such a core feature of the ecosystem that it would benefit greatly from stdlib support, and I think this fits that definition based on the evidence I've seen, at least. Though I realize others may not share this sentiment, I think its a viewpoint that has evidentiary backing (see all the talk about async Rust in the communicate, issues etc surrounding it. Its already a pretty big buzz topic relative to other things surrounding the language)
All this is to say, maybe its time to seriously start thinking about what first party support will look if we bring in a first party async / non-blocking I/O platform into the stdlib
Just to be clear, I think the NPM ecosystem is generally great and a massive success. People totally overblow the issues, and none of them are actually because of a small std library or due to the ease of install/publish.
> however, it is worth identifying when something is becoming / has become / is considered to be such a core feature of the ecosystem that it would benefit greatly from stdlib support,
I agree, and I think that there are a few places with regards to async that could work well here. Maybe some kind of Executor trait (hard, but maybe possible?), probably `std::block_on`.
> see all the talk about async Rust in the communicate,
FWIW I think the majority of people are just happily doing async work in Rust and don't get too involved in the discussions. I'm one of those people, except I'm also an internet addict on extended PTO so here I am.
> All this is to say, maybe its time to seriously start thinking about what first party support will look if we bring in a first party async / non-blocking I/O platform into the stdlib
I agree, I think some of this is best done in the language but some should be in std. No question.
That said, I disagree on the usefulness of async: in my experience it does the job and it's my default setup.
There's a complex project where I ended up just using threads because I needed to squeeze performance out but overall async is great for tasks, sequence of async operations.
Still, it had some rough edges: (it's been a while but) using Async Closures wasn't pleasant and I re architected my app to not use them as a result. There's an argument to be made for this change making my application easier to reason about - but overall it's poor language flexibility.
"I will forever argue that [...]" -- Why are you forever arguing?
"I feel like I am alone with [...]" -- No, I read it every day on here.
My hope is that by raising these things on HN, someone will take notice in a different way and at least start considering / revisiting alternatives
This is a serious wart on language design and while I can agree that it's likely too late to fix it for Rust, there is a kind of race among new languages to be a successor to C++ in many of the domains C++ is used in, and while I think Rust does hold a lead in that race, the race is not over.
A language that can provide an ergonomic solution to concurrency would absolutely provide a huge boost to any such language, and so to people in that space, listen to these complaints. Async/await is not a good solution to this problem.
You almost always hear people complaining about async/await in every language it's a part of but you rarely hear people complain about how Go manages concurrency.
I don't see how that makes it a natural transformation.
Not sure why, progress used to be way faster years ago.
For example, this "four year plan" should be implemented in 6 months at most, not 4 years.
And the "long-term features" that "should be considered carefully, could not be addressed in the next few years" (lol, seriously?!?) should have serious work start right now and be released as part as the next edition in 2024 (since they need an edition break to be ergonomic). These are basic misdesigns of the type system that should have been fixed 10 years ago.
Whoever is managing and paying for the Rust developers needs to fix this.
I wonder if it feels this way just because Rust has such a fast development cycle. But like, most languages takes years and years to do something like async. Rust went extremely fast, relative to other languages.
> For example, this "four year plan" should be implemented in 6 months at most, not 4 years.
That seems kind of insane. What language goes from idea to shipping major features in 6 months? Who even wants that?
Yes, and the language was qualified as unstable for evolving so quickly. Move fast and people complain, move slow and people complain. They lose both ways.
I find the development cycle to be quite good: they ship features and improvements and refinement, and I don't need to rewrite my codebase every three months because something subtly broke.
Consider the bit in OP about wanting to add a Move trait and deprecate Pin, and reverse the semantics so types are immovable by default. That's a difficult change to make now, after async has been stabilized. Obviously in this case the longer, deliberate process didn't save them from this (apparent) mistake, but overall new big language features should never be rushed.
> For example, this "four year plan" should be implemented in 6 months at most, not 4 years.
Now that just sounds reckless to me.