Comparison of Rust async and Linux thread context switch time and memory use
github.com
github.com
Think about it this way—if you have a user-space thread which wakes up due to I/O readiness, then this means that the relevant kernel thread woke up from epoll_wait() or something similar. With blocking I/O, you call read(), and the kernel wakes up your thread when the read() completes. With non-blocking I/O, you call read(), get EAGAIN, call epoll_wait(), the kernel wakes up your thread when data is ready, and then you call read() a second time.
In both scenarios, you’re calling a blocking system call and waking up the thread later.
Of course, there are scenarios when epoll_wait() returns multiple events, which reduces the number of context switches. But the general result is that it’s not always easy to beat blocking I/O and kernel threads.
What io_uring does do is provide a way to poll without needing to wait, but if you haven’t received new events when you poll, you’re not on the fast path any more. Whether you are often on the fast path for io_uring will depend on the particulars of your application and its I/O patterns.
Isn't "not on the fast path any more" a bit absolutist? io_uring's "slow" path is roughly one syscall per iteration, right? That's still many fewer syscalls than one syscall per IO operation (or more if any return EAGAIN/EWOULDBLOCK) as you'd be doing without it. I'm not sure I really care about eliminating that last syscall per iteration; it seems minor in comparison.
I am a bit baffled how this could possibly be considered an “absolutist” viewpoint—I am just saying that there exist scenarios where io_uring is not helpful. This should be uncontroversial.
That's not correct, io_uring was "absolutely" designed, at least in the technical sense, for zero syscalls in the slow path (if you want to):
IORING_SETUP_SQPOLL
When this flag is specified, a kernel thread is created to perform submission queue polling.
An io_uring instance configured in this way enables an application to issue I/O without ever
context switching into the kernel. By using the submission queue to fill in new submission
queue entries and watching for completions on the completion queue, the application can submit
and reap I/Os without doing a single system call.
From the man page: https://manpages.debian.org/unstable/liburing-dev/io_uring_s...This mode required privileges in early kernel versions but that's already changed. Things are moving fast.
Right, so if a blocking API makes 1 syscall, io_uring would make N syscals for N iterations.
That's pretty huge amortization. Rough benchmarks we've done are showing double throughput for io_uring for 4096 byte AF sector write/fsync/read combos: https://github.com/coilhq/tigerbeetle/tree/master/demos/io_u...
Same for the blocking case. If I do a syscall to read a whole file, its just 1 syscall creating millions of I/O operations.
io_uring is a bicycle for IO, and you can ride it as fast as you want to. But it's apples and oranges to blocking IO, which is always stuck in first gear.
The CPU can run other threads while the hardware does DMA transfers. The thread just yields when the transfer is started, and a hardware exception wakes it up when the DMA transfer finishes.
At the same time, multiple threads for a single program introduce context switches which are becoming horrendously expensive compared to the sheer number of IOPS that modern NVMe SSDs can do.
Thread-per-core designs built around io_uring are the future of IO on Linux.
So the whole thing is a bit of a false equivalence. Better interfaces that reduce context switches are desirable, but even where they exist often you are just substituting a user thread for a kernel one, and in the general case, there is likely to always be system interfaces that never make it into the brave new world -- take SysV IPC for example (a 1975 era API), it seems doubtful anyone would put the effort into making it async, but there will probably still be times where you might want to consume those interfaces for compatibility or some other obscure reason.
Also consider the case where a user program has a need for some substantial thread pools of its own, it might be the case in some scenarios that reusing resources that must already exist in user space and live in warmed caches makes more sense. Neither async or Linux threads are "better", it will always depend on a particular use case, and even then the right answer might well be some combination of both.
At this point it's proponents are being pretty unapologic that it will, in fact, reimplement every part of the syscall interface that is actively used.
And among Rust frameworks the same pattern holds. The fastest Rust frameworks are async while a synchronous frmework such as Rocket is about 20x slower.
[1] https://www.techempower.com/benchmarks/#section=data-r20&hw=...
[2] https://www.techempower.com/benchmarks/#section=data-r20&hw=...
You could make faster code with it but I wouldn't want to maintain it and you'd have to throw an obscene amount of man hours at it to get that performance.
But yeah, it's be super interesting to actually see that demonstrated - that'd be quite a lot of work, however.
Whether you read one file after the other sequentially, or try to read all of them concurrently, won't make a difference, because your Disk/RAM bandwidth is going to be bottlenecked anyways.
Trying to do this concurrently requires more work that won't pay off, so it might actually be slower.
Probably because they forgot to enable realtime priority for threads in the synchronous frameworks.
Failing to do that means Linux will starve your web request handling threads in favor of various system tasks you don't care about.
Lack of HTTP keep-alive is probably the most obvious thing holding it back.
This benchmark as written is probably underestimating kernel task scheduling cost since only 1 task is runnable at any 1 time, while a realistic multi-threaded system will have more runnable threads to juggle.
0: http://pdxplumbers.osuosl.org/2013/ocw//system/presentations...
(E.g., don't switch to the systemd or sshd thread if a customer's web request is timing out.)
That said, doing this right is out of reach of the average programmer, and it's doubtful that the compiler has enough domain-specific knowledge to do this automatically.
"Async" of the Python and node.js fame is yet another thing, a hack to get around their interpreters' inability to use kernel multitasking features because of global locks.
I disagree with this. I feel like using Async programming is actually much more powerfull and expressive than theaded programming, especially with Rust combinators on streams of futures (for examples: futures_unordered), which allow to trivially express complex concurrency patterns (such as: wait for the first two requests to return something and discard the third request's response, and btw also cancel that request). Async programming also allows for structured programming, where each task is an owned resource of a parent tasks, which means that lifetimes of tasks can be controlled and runaway threads can't exist (if one is avoiding tokio::spawn). I've been developping [Garage](https://git.deuxfleurs.fr/Deuxfleurs/garage) for some time now (a simple distributed object store that implements a subset of S3, not ready for production!), and I've been in awe about how easy it was to write these complex patterns using async Rust.
[0]: Which is fair, I wouldn't be surprised if this was the best way for some.
It's programming with threads, where you have a thread pool and pipes to put tasks onto that and a helper function/macro. Await does this under the hood of course.
Also I agree that multi-threaded Rust is probably the best alternative to async Rust.
[0] I have no problems with the GIL, but it's yet another factor to consider when a python multi-threaded program stops working as intended
You can do exactly the same thing, after all async is just a big auto-generated state-machine that puts jobs onto a thread pool and waits for them.
Nginx is hand crafted (artisanal!) event-driven C, which is exactly how async runtimes also work. A big loop (the event loop, usually an infinite `while` blocked on epoll()). For example NodeJS uses libuv for this.
The big advantage of first class async support is that we don't have to do this by hand. Plus it makes some optimizations easier (eg. putting things on the stack instead of assigning each event handler a slice of some global [heap allocated] structure).
Or maybe I'm simply misunderstanding what you meant. In that case cloud you clarify please?
That is, the 'easy' path is to write code such as the following (in vaguely C# pseudocode):
var p = await GetUserPermission( username );
var c = await GetServerConfig();
var m = await GetMessageOfTheDay();
Assume each await call is potentially an expensive SQL query or REST API call.The problem with that is that this is strictly sequential, synchronous code that is merely "dehydrated" and "rehydrated" to reduce overheads during the waiting periods. It is strictly slower when executed on a server that is not very busy! It must be, because it does the exact same work in the exact same order as the ordinary synchronous version, except now with extra state machinery and complex error handling woven throughout by the compiler.
Scalability is not everyone's concern. Scalability is for the FAANG sized companies. I care about the individual user experience, and async does nothing for that by default.
I mean, sure, you can write much more verbose code along the lines of:
var p_t = GetUserPermission(username);
var c_t = GetServerConfig();
var m_t = GetMessageOfTheDay();
await Task.WhenAll(new Task[] { p_t, c_t, m_t });
var p = p_t.Result;
var c = c_t.Result;
var m = m_t.Result;
But noone does this, for some values of noone. I've never seen code like this in the field.In fact, let's test this. I'm reviewing an asynchronous ASP.NET application developed in 2020 right now. It's a large app, with literally thousands of uses of the "await" keyword, at least 3500 files use it.
The only use of "Task" static methods are seven uses of FromResult(). That's it. Zero uses of WaitAll(), WaitAny(), or ContinueWith()!
This is typical.
It's not that asynchronous programming is hard, it's that it is unergonomic to gain a latency benefit out of it. Most applications need lower latency, not higher throughput. Hence, for most programmers, most of the time, asynchronous programming is next to useless. It's just extra noise and more failure modes.
I write this kind of stuff all the time because parallelizing long-running tasks without dependencies is one of the easiest wins when it comes to wall-time.
But this kind of optimization is somewhat orthogonal to async/await. You don't need fine-grained async to optimize long-running tasks, you could just throw a bunch of closures into a threadpool for that purpose. Async only makes sense when you're interleaving thousands of tasks with readiness/completion based IO.
Honestly, I'd be really surprised if this was a common practice in C#.
From a quick ddg search, it looks it does: https://docs.rs/futures/0.3.8/futures/macro.join.html
I never personally had any issue with working with threads and locks, finding it simple enough to reason about them, though I understand lots of people felt differently. When async/await first came to C# around 10 years ago, I grumbled because I didn't see the point; I found it much harder to reason about the flow of code, and initially at least, stack traces were a shitshow (things are much improved, but there is still a lot of cruft in async stack traces).
But async/await was heavily pushed, and "real" threading is almost relegated to the sidelines for most developers. Although having said that, I find that junior devs in particular really struggle to really grok async/await.
Anyway, several more years on, and I have mixed feelings about async. Because Microsoft has gone all-in on async/await, I think it's really easy to work with when building web apps and APIs with ASP.NET Core/MVC - there is barely any "developer overhead" at all, really. Web apps very often hit things like HTTP APIs and databases, and with how easy it now is, there is little reason not to use async/await. Yes, for small loads there is a tiny performance loss due to the runtime setting up async state machines, but it really is almost always completely insignificant - even moreso with the advent of ValueTask, and again more recently with pooled ValueTasks. Yet the gains can be tremendous.
But for non-web apps/APIs, I feel differently. I spend a lot of time writing server-side processing services, and things like Windows services for desktops (in the infosec space), and I've gone all-in on async/await because Microsoft has gone async-first. Hell, a lot of stuff is async only now, so unless you want `.GetAwaiter().GetResult()` everywhere, you have little choice. Anyway, these systems are more complex than web apps, because with web apps, most of the real complexity is hidden away in the framework. But here you have to deal with work queues, caching, pooling, serialisation etc all by yourself. And with async/await, it can be hard to reason about the flow of code, and it's really easy to break things in ways that are really painful to diagnose. And it means that every.single.stacktrace contains async cruft that you need to sift through. Which is not fun.
Anyway, this is much longer than I meant, but my conclusion is that I'll continue to use async/await for web apps and REST APIs (because, why not), but for services, I'm going back to the threadpool, green threads and synchronization primitives, and only using async/await in a limited way where it provides clear value - not async all the way down from the entrypoint.
Welcome back, my beautiful, green threads! (⌐■_■)
But most of the time you not only use thread, but also several synchronization primitives (locks, channel, etc.) and when doing so, regarding stack trace you are in an even worst situation than what async stack traces gives you (“some thread changed this shared-memory value and now it's not what you expected, but you have no easy way to know which one did and when, good luck”).
Regarding shared, mutable state - if multiple async "threads" can access that state, then you still need to guard it, but usually with an async-capable means.
Sometimes, but not as often, because the scope of your async function is often the only “shared state” you need.
AFAIK, .Net doesn't support "green threads" and they repeatedly confirmed that there are no future plans to do so. Additionally, M:N threading model has serious interop issues as evident in Go, which is a no-go for system languages. Personally, I don't see a need for green threads since kernel threads are fast enough and don't use that much RAM as people tend to believe. And when they are not, sure, go async/await.
I had actually meant "normal", OS-level threads.
The runtime will generally schedule IO bound tasks to run on the threadpool.
Well, that's not correct. Unless you explicitly call Task.Run or Task.Start (or other similar methods) no new thread is created. The compiler generated state machines don't require the threading mechanism to work. In fact the overhead for async/await is mostly the extra code generated for the state machine and error handling. At runtime, there's no thread switching overhead.
Otherwise, from memory, the runtime spec doesn't actually guarantee that await won't run on a threadpool thread - it will under certain circumstances.
And then there are further nuances if there is a synchronisation context and ConfigureAwait(false) is used, as the continuation will be scheduled on a threadpool thread.
let p, c, m = join!(
GetUserPermission(username),
GetServerConfig(),
GetMessageOftheDay()
);
(you don't even have to write await when using the join macro)With such simple syntax available it seems obvious to me that one would want to use it as often as possible, and it's also much simpler (and probably cheaper) than dispatching those three tasks to a thread pool.
f1,f2,f3 are all async fn's, that is all you have to do
Mist opportunities for concurrency are between unrelated tasks (because related tasks often dependencies between them). Unrelated tasks tend to have unrelated return types.
var p_t = GetUserPermission(username);
var c_t = GetServerConfig();
var m_t = GetMessageOfTheDay();
function_to_call1(await p_t, await c_t);
function_to_call2(await m_t);
This does not look any more complicated than a non-async function. Not sure how this example justifies your claims.Besides, even if your example is valid, the usage of Task.WhenAll has nothing to do with your claims either. The use of async/await is majorly for scalability. Being able to make several network calls concurrently is not the major concern. Even if you await at each async call, you still achieve better scalability because threads won't be blocked for async calls and can work on something else.
And FWIW, this explicit form is often unnecessary - if you kick off each task they will run in parallel and then just await each task only when the result is needed, it can look a lot cleaner:
var p_t = GetUserPermission(username);
var c_t = GetServerConfig();
var m_t = GetMessageOfTheDay();
var foo = isAuthorized(await p_t);
// more code here
var msg = ( (await c_t).ServerName + await m_t) ); await Task.WhenAll(p_t, c_t, m_t);
Or you can just await the threads before you need them. They're already started and running at this point.You also probably want to avoid using Result and just await the completed task for the nicer unwrap syntax. Plus, you don't want to get into the habit of using Result as its a blocking call. Same with WaitAll and WaitAny. Ideally you would never use those. ContinueWith is also not very needed if your style is to use the more plain await syntax. Those methods are more to bridge blocking and async code so an async from the start app might use async extensively and never those methods.
Perhaps search for WhenAny and WhenAll?
The killer use case for async tasks is when you need hyper-concurrency, e.g. hundreds of thousands of concurrent tasks. In that case, as the article mentions, you can't use OS threads anymore. Of course there are some use cases requiring this level of concurrency, messaging servers come to my mind, but there are also many, many use cases were you need a lower level of concurrency, like a few hundred concurrent tasks max. In those cases I think using OS threads can work pretty well with less complexity.
An instance I ran into personally was, effectively, task scheduling. Sure, I could have done the 'normal' thing, of a priority queue being populated from the database on some interval, having some thread reading from that queue, sleeping until the first item needs work, pulling it off, throwing it onto a threadpool. Have to take care to ensure the threadpool is large enough for the maximum amount of concurrency I need, have to make sure that I'm careful in what data structure I use for the priority queue (I need to make sure I'm not adding the same task multiple times to it, and that when adding items to it I'm not locking it), make sure the polling thread can't throw (or at least, when it does, it restarts or kills the program and that then restarts), a few other niggles here and there too. And a whole 'nother level of complexity if tasks lead to follow up tasks (i.e., a task represents a state machine through a series of transitions, which themselves take a sizable amount of time, to where just leaving them on the thread is a bad idea, since it uses up the threadpool).
In a 'free concurrency' world, I just spin up a new concurrent process per task for some window (same as how many items I added to the priority queue). And that's basically it. Each process can step through its state machine, sleeping in between tasks for however long, without issue.
Best combination of things I have come across.
[1] I count await as explicit as it forces the awkward top level only suspend model.
If you like the async style better, then fine, use it. Sometimes you win like that, where the thing you like better is also faster. But don't worry so much about the performance.
Web frameworks is another place I see this a lot. Crossing the streams, if you've got an incoming web request, unless your framework somehow consumes and discards the web headers, a real web request is already many kilobytes just to represent the incoming headers by the time it gets to your handling code. Using async because it has ~200 bytes per task vs a thread allocating 10K out of the box at that point doesn't make much difference because the HTTP request itself is blowing out the difference.
The spread in orders of magnitude in what is expensive and what is not has gotten so significant on modern systems that you can easily get developers sitting there optimizing nanoseconds while throwing away seconds. The old school assembly-style premature optimization where we're trying to save every bit and cycle has mostly passed away, but its replacement seems to be this; frantically benchmarking how many millions of requests per second some framework or feature can handle as if it matters when your code is going to take 500ms.
Meaning you don’t need to read all the bytes from the tcp socket before deciding which route to take. And also the handler for that route is given a stream object and will just read as many byte as it need.
That's unfortunately far less reliable in practice than it seems on the first glance: You might never know whether any async function you call spawns something else, or makes use of `spawn_blocking`, `block_in_place` or any other function which isn't a pure state machine.
If you try to cancel any of those, you will get either excessive blocking or end up with runaway tasks.
A better solution for this is real support for structured concurrency, as available in Kotlin, Python Trio and now coming to Swift async functions. This doesn't really require immediate cancellation - as favored by Rust futures. It works better with cooperative cancellation, where cancellation is requested asynchronously and ongoing tasks are supposed (but not forced) to listen and follow the cancellation recommendation.
You're confusing mechanism with semantics. Here's an article on how Java's Project Loom achieves structured concurrency using its new virtual threads capability: https://vorpus.org/blog/notes-on-structured-concurrency-or-g...
And the article cited above that proposes the idea of nurseries for structured concurrency: https://vorpus.org/blog/notes-on-structured-concurrency-or-g...
Arguably, structured concurrency as described above is easier to obtain when using threads as your underlying mechanism, because the vast majority of code is serial[1]. That there are a handful of critical regions where you want to express concurrency relationships doesn't mean we have to discard threads. That's throwing the baby out with the bath water.
Self-promotion: I had stumbled on the idea of "nurseries", independently and many years before the above were published. See https://github.com/wahern/cqueues It's nominally a non-blocking "threading" API for Lua. (In Lua coroutines are also called threads.) But note the plural, continuation queues. It's trivial to instantiate a queue, which is similar to a nursery. This was by design. Many cqueues projects naturally end up with a tree of thread controllers/schedulers. It doesn't work on Windows (yet) because it relies on the fact that kqueue, epoll, and Solaris Ports descriptors can be recursively polled.
[1] Serial != synchronous/blocking.
Low context switch latency only matters when the number of tasks is very small (their data all fits in the cache), and the workload is entirely computational. Otherwise, even the fastest implementation is ~60 ns, which is the cost of a cache-miss, and the compiler can't optimise things into a simple goto because the dispatch goes through a scheduler that has a megamorphic call-site.
So memory is much more important for I/O use-case throughput, and while it is true that the kernel doesn't commit the full stack memory on thread creation, it's misleading to think that you get good memory usage. For one, once the memory is committed, it's never uncommitted (although it can be paged out). For another, the granularity is that of a page, i.e. at least 4K, which can often be much higher than what a task requires.
> It is hard to pin down exactly how the alleged advantages would arise.
For I/O use-cases the answer is here:
> the async version uses about 1/20th as much memory as the threaded version.
This could translate to 20x throughput -- due to Little's law -- although usually less because there are other limits, like network saturation.
And I only actual use them after one iteration of whatever I do. So the core could fetch the memory content without actually having to stall because I do not use it until later.
I'm not sure how realistic that is inside a kernel thread scheduler, but it sure is useful in user space for task based libraries.
Maybe I'll try it in a blog post one day and see what percentage of the comments consist of hurled fruit.
So, it's not clear why you'd abandon the async syntax just because you're compute bound.
Sadly with so many things having gone async-first (or only) it’s become difficult not to end up with an async runtime anyway, or not be forced to use an async system. I wanted to build a small web-based tool for local, didn’t really find anything which was not async.
Here's how you would do a write()/fsync()/read() with this (https://github.com/coilhq/tigerbeetle/blob/beta/src/io.zig#L...):
const offset: u64 = 0;
const bytes_written = try io.write(fd, buffer_write[0..], offset);
try io.fsync(fd);
const bytes_read = try io.read(fd, buffer_read[0..], offset);
Other sync functions can use this asynchronous IO completion code in a synchronous style (as this snippet shows) and still get all the zero-syscall and asynchronous performance of io_uring. What this is actually doing under the hood is filling SQEs into io_uring's submission queue ring buffer and then later reading completion events off io_uring's completion queue ring buffer, so it's fully asynchronous in the I/O sense but this hasn't spilled out and leaked over into the control flow. The control flow is as it should be, nice and simple and synchronous.Beyond this, Zig still allows you to explicitly indicate concurrency with the `async` keyword, for example if you wanted to run multiple async code paths concurrently.
But the crucial part is that Zig's async/await does not force function coloring on you to do all of this: https://youtu.be/zeLToGnjIUM
Pretty incredible on Zig's part to be able to pull this off. Huge kudos to Andrew Kelley. Also, thanks to Jens Axboe and io_uring, what you saw above was first-class single-threaded or thread-per-core, there's no threadpool doing that for you, it's pure ring buffer communication to the kernel and back, no context switches, no expensive coordination. Pure performance. There's never been a better time for Zig's colorless async/await. The combination with io_uring in the kernel is going to be explosive. It's a perfect storm.
async fn read_to_string(path: impl AsRef<Path>) -> io::Result<String> {
std::fs::read_to_string(path)
}
It doesn't have await inside! My mind was blown as I saw that.That is a super interesting strategy, though obviously only works when you can « afford » a multithreaded scheduler.
Anyway I wonder how they manage this, signals?
Yeah, for example in comparison actix-web only uses single threaded workers - one per core. Future in actix-web doesn’t have to be Send or Sync, and I think it’s incompatible with what async-std is doing here. That design is almost certainly one of the reasons actix-web tops phoronix
Each worker thread runs in a loop executing a queue of jobs. On every iteration it sets an atomic progress flag to true.
The runtime in which it's contained polls its workers every 1-10ms, atomically swapping in false and checking to see if the previous value was also false - if so, it steals its task queue and spins up another worker to execute it.
https://github.com/async-rs/async-std/blob/ceba324bef9641d61...
> This blog post describes a proposed scheduler for async-std that did not end up being merged for several reasons.
I don't think it's a particularly good idea in the first place - it's basically an automatic watchdog-driven block_in_place(). It doesn't remove the problem of blocking in futures, it just limits the damage to the local task rather than blocking the entire executor.
That's fine in the simple case of future-per-task, but it's pretty common to be polling multiple futures concurrently within one, so it's not a general solution.
It’s really not though, at least as long as the parameters and results are Send. For instance Tokio has a spawn_blocking which runs the function on one of the blocking threads it spawns on-demand specifically for that use.
Meanwhile « blocking on a future » requires adding and managing an entire async runtime and its interactions with the rest of the program, and locking up the runtime is a very real possibility.
What other reason were you thinking of?
This is a great example in Node on useful combinators that with async await make it easy to express parallel programming concepts with familiar tools. No manual IPC, no fork/join child PID/thread ID handling, etc.
https://github.com/sindresorhus/promise-fun
The same abstractions (or many of them) exist in Rust, but I think the above is illustrative of the ways we can combine async object returning functions and then use await to hide the complexity of the state machines needed to drive them.
That this abstraction that makes code easy to read and write also performs better is the icing on the cake. The former prevents bugs and keeps code quality high, and that is worth much more.
Rust is not Javascript. Using threads is actually a lot simpler in Rust than async/await.
Also threads are definitely not easy to use in all languages. E.g. C++ gives you very little help (no channels for example), and JavaScript makes starting threads difficult and moving/sharing memory is limited to primitive arrays.
So, code potentially laden with use after frees, double frees, shared and mutable data, and so on.
No offense to you, but I would be leery of trusting that code in any languages except a handful. Certainly not C/C++, and if it were written in Rust, I would hope it would use a thread combinator library and channels.
- A. Async/await - compiler saves and resumes functions.
- B. Message based - Golang, Erlang, threads with messaging.
With category A, I can use my IDE to jump to every function that is called and easily follow the computation.
With category B, all of these connections happen at runtime with messages.
When you have a tree of tasks all which may save/resume many times, async/await it easier to understand than launching a thread per IO event.
A. Implicit messaging using the languages function syntax (async/await).
B. Direct messaging using a message passing feature of the runtime (Erlang, Golang)
Note: I mean "messaging" in the context of a single OS process, that possibly has many threads (so within a single language runtime).
Async/await is still implicit messaging, but it appears like a regular function call - which in my opinion is easier to understand. Using function args/return for input/output is something every developer already knows.
In contrast, Erlang and Golang require you to use some type of messaging feature in addition to functions.
> A useful way to think of Go and Erlang is that they automatically and transparently insert async/await each time you call a function that performs I/O
The part they are missing from async/await is the ability to easily get return values without messaging, and do this recursively for a large tree of functions.
E.g. getting a return value from `go x()` requires messaging, but with async/await you could do `const p = x(); const ret = (await p); // return value received at a later time with no messaging.`
Both of them will require you to create some type of messaging topology to return the values (which makes your program a mixture of (regular functions + messaging features) vs async/awaits "everything looks like a function").
No, they do not. In Elixir for example if I call:
bytes = File.read!("filename.txt")
`bytes` will have the data returned from the function call immediately, with no need for message passing or awaiting the result. Under the hood, it is still asynchronous evented I/O. If I want to explicitly await for flow control reasons (await all of or one of multiple events) that is available in the stdlib in the `Task` module. E.g. t1 = Task.async(fn -> do_this_thing() end)
t2 = Task.async(fn -> do_this_other_thing() end)
Task.await_many([t1, t2])
You can accomplish most things without ever calling send/receive or writing your own gen_server etc.Last time I used Erlang (pre-Elixir), the `bytes` example would require you to set up a request/response with a blocking `receive`.
I think the key issue is that the inputs and outputs are disconnected in the static program text (and only connected dynamically at runtime).
Two contexts that matter for understanding how a system transitions between states are:
1. Program editing/reading.
2. Runtime.
I think AA is superior for understanding the system as a whole in both these contexts, because at edit time the IDE jump to def/show all usages allows you to understand every function that will be called, and at runtime you can get a stack trace to understand where the current function came from, and where it is going.
With message passing runtimes, both 1 and 2 require extra mental models on the part of the programmer, because they also need to understand the network topology (which either is not possible statically, or requires extra tooling on top of functions).
Message passing breaks down your system into CSP's, which makes it easy to understand each sync process, but hard to understand the whole system, as the same program-writing-process that allowed you to break down your components is working against you when you need to put them together again to understand the whole system.
I could be wrong as I have not used modern IDE's or debugging tools with message passing runtimes lately.
This is a big surprise.
If you look at the Techempower web benchmark [1], the performance of actix-web is about 20x higher than that of Rocket.
The common explanation is that actix-web is async and hence much faster than Rocket which relies on kernel context switching.
But if Rust async and kernel thread has the same switch time as shown by this benchmark, then why is actix-web so much faster than Rocket?
[1] https://www.techempower.com/benchmarks/#section=data-r20&hw=...
So the Rust async context switch is on top of the regular Linux context switch, not instead.
"Linux thread context switch time" is a meaningless metric, since Linux will switch thread context regardless of what you choose to run on your computer.
Any "async" switches are additional overhead; you don't get to not have kernel preemption just because your Rust thread is now switching contexts "asyncly".
There are benefits to having an additional user-mode scheduling mechanism inside your kernel thread, but saving CPU cycles isn't one of them.
my point was that thread context switches caused by preemption happen at an entirely different time scale than the rate of context switches caused by syscalls (if the system is doing any meaningful level of IO)
Looking at the techempower benchmarks, the projects using tokio generally outperform Go, Java, so I'm guessing it's on par or better.
Hypothetically, you could port goroutines exact behavior to rust and use that as your wanted to too.
Yes, pages will only be allocated for a thread's stack when the thread actually uses them. However, the thread does not release said memory afterwards. The memory can only be reused by the same thread. If a thread ever once does something that temporarily allocates a bunch of stack space, then it forever consumes that space going forwards even when no longer needing it. If you have 10,000 threads and each one of them happens to, at some point in its lifetime, use 1MB of stack space and frees it, then you are now using 10GB of RAM on mostly-unused pages.
Now you might say "what on Earth would ever use 1MB of stack???", but the problem is, in normal programs with few threads, there's no problem with a function temporarily using a ton of stack, and so random things feel free to do so. Maybe some library call you make likes to allocate a temporary buffer on the stack and you don't even know it. There's also normally no problem with doing some deep recursion every now and then, so it happens. Often, stack allocation is data-dependent (e.g. recursive descend parsing). So if you try to strictly limit your stack space then you risk running into random stack overflows or maybe even security issues. And if you do find a limit that works, it's still probably much larger than the average usage, so you're still wasting a bunch of memory.
IIRC, Go uses segmented stacks to avoid this problem, but C/C++/Rust do not. (I think Rust tried to at one point, but later gave up on that because of the complexity?)
In contrast, async tasks only hold onto the memory they are actually using to store live data at any particular moment. If an async task invokes some deeply-nested function and uses a bunch of stack space, it doesn't really matter, because all the tasks are running on the same thread, so the next task to call that function uses the same pages rather than allocate new ones.
(There's actually a similar issue regarding heap space. Memory allocators that perform reasonably with multiple threads typically maintain per-thread freelists, so if you have lots and lots of threads, you end up with a bunch of free'd memory stuck in freelists. Though, some allocators, like the new tcmalloc, are starting to use per-core freelists instead, which may avoid this problem.)
M1-MBP async-brigade % time cargo run --release 500 tasks, 10000 iterations:
mean 761.403µs per iteration, stddev 8.929µs (1.522µs per task per iter)
cargo run --release 3.21s user 4.60s system 99% cpu 7.818 total
M1-MBP thread-brigade % time cargo run --release 500 tasks, 10000 iterations:
mean 787.149µs per iteration, stddev 67.289µs (1.574µs per task per iter)
cargo run --release 0.94s user 7.19s system 100% cpu 8.081 total
I ran it a few times and the numbers came up rather similar each time: async-brigade finished in 760.273µs-764.928µs while thread-brigade took 784.510µs-796.323µs.
As macOS doesn't have taskset, I can't easily set affinity. I tried to use the workaround documented elsewhere to use Xcode's Instruments to reduce the number of CPU cores but it would always re-enable itself at 8 cores, so that didn't work.
What I noticed is that any syntactical benefits of async/await has a lesser impact when most of your application logic lives in pure functions, since you greatly reduce the amount of code in async functions.
When I started using async/await in JS 4-5 years ago I thought: "How could we have lived without this for so long?". These days I don't care much about it.
If you add a huge thread pool, and then those downstreams don't have a large latency, then you end up accepting a huge amount of work and then are CPU starved.
So in order to correctly size your thread pool, you need to understand all your downstream latency, and adapt to it.
Compared to an async runtime, which just handles this scenario, it's very painful.
Even if you get this roughly right, the scheduler is very unhappy when you have lots of threads - it tends to make incorrect scheduling decisions.
You can also have lower RAM overhead per thread if you choose a smaller stack space. Many programs will run fine with a smaller stack space, BTW.
----
Years ago I had to build a load simulator in C#. The CTO looked at me and told me that it had to simulate 100,000 clients; thus it had to be async.
He arranged for me to have a very powerful computer to run the load simulator.
I originally wrote non-blocking code. The non-blocking code had very low load at 100,000 clients, but I hit a problem with a difficult-to-understand edge case.
Because we only had a weekend to do load testing, I refactored the load simulator to be threaded. It only took me 20 minutes or so. The problem with the difficult-to-understand edge case went away, but CPU usage went up dramatically.
We had to tune the .Net framework to use a much smaller stack space.
In the end, I was able to have 100,000 threads to run the load simulator. CPU usage and RAM usage were very high, but the load simulator ran fine.
If I had more time, I would have taken the time to understand the edge case and continue to use non-blocking code. Then the program would have used much less system resources, but ran just as fast.
No, this is about as low as it gets. As the author explained, "the kernel only allocates physical memory to a stack as the thread touches its pages, so the initial memory consumption of a thread in user space is actually only around 8kiB."
The smallest possible page size (on x86-64) is 4 KiB, and you can't share pages between thread stacks, [1] so you can't go below 4 KiB of physical memory usage per thread. I'm not exactly sure how the author got to 8 KiB; maybe they meant "for each userspace thread" rather than "memory used in userspace" and are counting kernel memory too. I'm pretty sure the kernel uses at least 4 KiB per userspace thread (for a stack of its own, among other overhead).
Green threads won't take you below 4 KiB either, for the same reason.
[1] Without some custom ABI that guards against stack overflow in a different way. Golang has a custom ABI (I'm not sure exactly if this is why), and interoperability with C suffers, so this isn't an approach I'd love for Rust.
https://www.slideshare.net/brendangregg/rxnetty-vs-tomcat-pe...
This is not at all a fair comparison unless you're using io_uring.
Also (very rough) benchmarks (take with a pinch of salt) comparing various styles of fs and network IO (blocking, epoll, io_uring) for C and Zig: https://github.com/coilhq/tigerbeetle/tree/master/demos/io_u...
But user-facing apps (the sort people run on laptops, say) have async I/O as table stakes, really. It's not even about throughput or CPU cycles: it's about the fact that if you have I/O latency on any thread the user interacts with the user experience will be terrible.
Now in practice maybe that means "just make the I/O async, but the performance details of that don't really matter too much".
Anyway, the overall comment was about performance profiling in general, not just async I/O.
If you speed things up by 10% on your server, they'll get 10% faster on your laptop as well.
> If you speed things up by 10% on your server, they'll get 10% faster on your laptop as well.
Depends on the speedup and techniques to achieve it. For example, speeding things up via more parallelism can lead to wall-clock improvements on servers but not laptops, precisely because the latter just end up doing more thermal throttling....
Ideally, you want to measure both ideal hardware and actual-user-hardware; often speedups on one will not be visible on the other and vice versa.
* Rust's async ecosystem [1] adds a lot of complexity over simple threaded code.
* Rust's async ecosystem doesn't interoperate as easily with C libraries written in a simple threaded way. (And it's debatable which interoperates more easily with C libraries written with a different event loop.)
* async tasks can't be preempted, so concurrency will fall off a cliff if they run on O(cpus) threads and involve long-running computations or accidental blocking.
I think it's reasonable to ask if these numbers are enough better to justify all that, particularly given the disappointing "this advantage goes away if the context switch is due to I/O readiness".
And to go back and argue pro-async for a moment, io_uring might eliminate that disappointing caveat.
Then again, on the pro-thread side, there's Google's interesting fibers model that might solve some of these performance issues. [2] Also, "~17µs for a new kernel thread" is the wrong number, since you can avoid that cost with a simple thread pool.
Personally I think some things are better written as async, but it's a mistake to impose it on everything. For example, if you're writing a web app in Rust, I think you're usually better off writing threaded request handlers and having a mechanism for them to interact with the async hyper code. The hyper code is better off as async because an Internet-facing server might have an enormous number of connections in keepalive state.
[1] or maybe I should say ecosystems, plural, given the current tokio vs async-std divide.
You get a marginal to good benefit if you have a specific work load that I think most people don't really have.
A modern microcontroller/microprocessor is inherently event driven (for example, on ARM, at the very bottom of the call stack there is a wait-for-event (WFE) or wait-for-interrupt (WFI) instruction).
If async needs to be polled to run ("Futures are inert in Rust and make progress only when polled"[1]) this means my processor should be busy running these async tasks instead of waiting (WFE or WFI) as the result of a native call to one of the operating system functions (i.e. recv() on a socket). What is the impact on embedded battery-based systems?
[1] https://rust-lang.github.io/async-book/01_getting_started/02...
for network io, behind the scenes this is most likely utilizing epoll system calls. epoll mitigates context switch problem in a few ways, mostly because there is only one stack context to notify about new io events, instead of many.
Main idea is that a 'scheduler/executor' at the runtime/language level that knows about the state of the program can (a) 'save' and 'restore' fewer things compared to an OS context switch. (b) co-operative stuff does not need to pay the cost of too many unnecessary pre-emptions
> polling is only explained as a logical thing
But there is the poll() function that returns either the result of the operation, or "pending". So it's more than logical. Correct? I mean, if I (or the executor) don't call poll() nothing happens...
> OS thread now switching to execute whatever it is that it is waking up.
This is what confuses me. As I see it (and I what I understand from reading), async/await splits a routine into a (very smart) state machine.
I assume that there is no magic underneath. I mean, I can do the same state machine by hand if I want to, under the constraints of what the OS makes available for me in what context switching regards (APIs for waiting and synchronizing).
For a (OS/native) thread that has to wait for data on a socket, you have (basically) two options: wait on recv() or poll recv() without timeout.
Waiting on recv() would block (so no other code of my thread can be executed while waiting), so I guess the state machine needs to poll on recv() (I believe this is what this[1] example does).
In order to no block my thread, the executor either spins its own thread, or has to wait for my thread to poll() it.
[1] https://rust-lang.github.io/async-book/02_execution/02_futur...
The benchmark compares fibers to threads and has little to do with Rust. You will see the same numbers for a fibers implementation in any natively compiled language like C or Java.
The title is completely misleading, especially for most people who are not aware of this important distinction.