Learning async Rust with entirely too many web servers
ibraheem.ca
ibraheem.ca
I've updated it over the years (moving from chaining futures to using async, using the latest tokio, etc) to keep it relevant.
In particular, correctly splitting a duplex connection into send/receive streams and then aborting one side of the connection when the other closes without missing events is trickier than it should be.
(I'm actually surprised it's that small, and it's clearly written for clarity not size. nicely done!)
For anyone not used to more than simple concurrency stuff though, yeah, loads of subtle footguns. And they're often the sort of thing you might not notice in normal manual tests. Thanks for the link and maintaining it!
The copy_with_abort() routine is still taking the easy way out in this not-optimized-for-heavy-production-use sample because it uses a broadcast channel per connection to reactively signal that the other half of the connection should be closed (rather than timing out every x ms to see if an abort flag has been set). In the real world, I'd probably replace the join! macro with a manual event loop to be able to do the same but without creating a broadcast channel per-connection.
(I maintain an extremely lightweight "awaitable bools" library for rust [1] that is perfect for this kind of thing (roughly equivalent to a "bounded broadcast_channel<()> of queue length 1, but each "channel" is only a single (optionally stack-allocated) byte) — but it's for event loops in synchronous code and not async executor compatible.)
[0]: https://github.com/mqudsi/tcpproxy/commit/0164ef836a49f2f738...
But there are quite a lot of people stuck somewhere between "can reliably use a single blocking queue correctly / coarsely protect things with one mutex" and "can handle multiple competing interactions with error handling". More material to help cross that threshold is definitely a good thing.
These runtimes seems heavily scewed towards being used by IO bound application.
While rayon can help with CPU bound tasks. I wonder if these mix at all, it may not be desired at all. But I think fairly interesting nonetheless.
I know that tokio doesn't like being blocked by anything. You will get a hard error, if you do a blocking operation when not being spawn_blocking via. tokio. Rayon, probably doesn't trigger this warning, as it is technically not doing any WouldBlock operations.
you would probably do a spawn_blocking for rayon as you would for IO bound operations.
Anyways, mixing CPU and IO bound operations on the same CPU/machine is probably the same as mixing oil and water, and would probably be a pain in the neck to maintain and get good utilization.
It is actually smack bang in the middle of the tokio docs.rs page: https://docs.rs/tokio/latest/tokio/index.html#cpu-bound-task...
It is also a bit interesting how these paradigms compare with Golang green threads, which also seems to yield without a lot of syntactic sugar. The same for elixir. These languages also seem to solve the IO bound problem quite well using this paradigm (you obviously doesn't get as much control as you do for rust), but powerful nonetheless.
I think you can mix them (or you can, as mentioned in the link, have a separate runtime for CPU-bound tasks) but it is probably easier to use Rayon.
However there seems one thing worthwhile to point out:
> All to implement graceful shutdown.
It isn't true that async IO is required for that. Blocking IO calls can also be interrupted via signals. For the use-case of Ctrl+C that is mentioned in the blog post even the simple setup which unblock the syscall and let it return EINTR are successful. But even custom in-application cancellation from different threads works via signals. E.g. building a Thread::interrupt() that unblocks IO on a different thread is possible.
Most classical applications (e.g. proxy-servers) that went for the nonblocking route did it indeed for bigger scalability and lower memory footprint than using a thread per connection model. In the last couple of years and especically in Rust I however feel that most applications choose async either because the library ecosystem forces them or because there's just an assumption of missing out on something if not using it. There's rarely measurements performed with both setups that actually quanitfy any win or loss of using async IO.
In my benchmarks with a slightly tweaked version it was 2x faster than Nginx and and 30x faster than Python's SimpleHttpServer.
[1] https://github.com/DataDog/glommio/blob/master/examples/hype...
I like this post because it explains the debugging process, and makes me want to give another try to async Rust.
The other way is that sometimes a little more load perturbs the system so things take a lot longer. I've had systems that work fine until you hit X% cpu, then it jumps to 90%+, because of say lock contention; sometimes N is 80% so whatevs, but sometimes it's 25%, and then you need to find and fix the bottleneck.
E.g. pool of 10 threads does computational heavy task requests and responds when done to the async i/o bound task.
The article mentions how non-blocking IO is the way to gain back control otherwise conceded to the kernel in the case of blocking IO. However, non-blocking IO has to be laboriously implemented in user code. Is that true still? What if we used the threading model but relied on io_uring?
Note that the io_uring approach isn't too different (just more universalized) than how Windows has always implemented async operations (with overlapped and io completion ports). Code written for that environment originally just used the event loop model but it's easier/better ergonomics to wrap it in async and use that instead (e.g. how C#'s async/await works) because you can avoid callback hell, state passing, etc. and just write code "linearly" with cooperative yield points invisibly stashing and retrieving the "stack" for you.
Whilst you can do this, I still think the best approach is “actual-async”. The idea of “spin up this worker and shunt the data across to it just so it can block” strikes me as somewhat wasteful. If you know you’re going to have to wait, you’re better off not expending resources doing so, and instead use the time to do useful work instead.
It doesn't make any sense to say 'use the time to do useful work instead'. That's what you do when you block: in a threaded model, when you block you do 'io_uring_get_sqe', 'io_uring_prep_read' (or whatever), 'io_uring_set_userdata' 'io_uring_submit' and then you switch contexts (with <ucontext.h> or some equivalent) to another thread that is ready. When you have no ready threads you just call 'io_uring_wait_cqe', and for each completion you do 'io_uring_get_userdata', get a pointer to the thread, record the result of the operation and append the thread to a runqueue to be resumed.
From the perspective of the thread, you write
int e = read(fd, buf, sizeof buf);
if (e < 0) err(1, "read: %s", strerror(-e));
and it all feels completely sequential. Underneath it is stackful coroutines.