Async and Await in Rust: a full proposal
boats.gitlab.io
boats.gitlab.io
As for this article, I think the broad implementation details that it talks about are largely still accurate, just don't take it as gospel. :)
Now that Rust is getting some maturity i worry that it gets too complex in too much direction just because it's not powerful enough to express the few common abstraction, of which several instances are being added as distinct concepts (async&try, const datakind, higher-order poly--especially for lifetimes, named impls, before-after memory state for references...).
(Previous version of this comment said "HKT" but that's not purely right; we have a path forward for something at least HKT-like, but do notation is the bigger question.)
Like what is the real road-blocker, since we're talking here about abstractions that are gonna be monomorphised (at most into lifetime-polymorphic function, which anyway get erased) and known statically, right? Is it only a problem of getting the typing rules consistent and robust?
First off, do notation desugars to closures. But Rust has three different kinds of closures. How does that work?
Furthermore, Rust has imperative control structures. How does all of this interact with each other? Are you now no longer allowed to use those constructs inside of do notation? That feels inconsistent.
I don't know of any languages that have something like do notation but don't have everything boxed. It makes everything easier, but that's not generally acceptable in Rust. For example, we can use async/await with no dynamic allocation. Could we with do notation?
"Open question" doesn't mean "impossible", mind you. But nobody has ever come up with a design. In the meantime, we have users to support...
- I wrote the generic associated types RFC (how Rust will implement higher kinded polymorphism).
- I wrote the const generics RFC (the closest Rust will get to dependent types).
- I wrote the async/await RFC, as well as the linked blog post.
That is to say that I am intimately familiar with how Rust's type system can be extended to support more "powerful" abstractions.
Monads as implemented in pure functional programming languages like Haskell cannot usefully abstract over asynchronous and synchronous IO in Rust for a variety of reasons having to do with the way the type system exposes low level details by virtue of Rust being a systems programming language. I do not believe that `do` notation could be a useful mechanism for achieving either the ergonomics or the performance that async/await syntax will have in Rust.
I'm responding to you because you're the top comment, but I could write a similar response to a lot of comments here. Monads, stackful coroutines, green threads, CSP, etc - we've heard of them! :) We have well-motivated reasons to choose async/await: its the only solution that meets our requirements.
Sorry, but I don't buy it. I had a half-baked, unfinished proposal for an effects system that would have allowed Rust to implement async/await just as efficiently (no stackful coroutines) along with any number of other effects [0]. Maybe it wouldn't have been a good idea due to stretching Rust's complexity budget too far, but that's very different from saying it's impossible. Having watched the development of Rust closely I really think that the design team just didn't understand the theory side well enough to be able explore the design space here. (I'm not being as critical as I might sound, PL theory is hard and the Rust devs have wielded it much more competently than the designers of any other non-research language).
[0] https://internals.rust-lang.org/t/start-of-an-effects-system...
Sticking to the first part: an effect system is not what the user I was responded to was talking about. They were talking about building do notation on top of type classes with higher kinded polymorphism, which cannot effectively abstract over the monadic operations in Rust.
I interpreted OP's comment as complaining about the lack of more general abstractions in Rust that would allow you to implement async/await. Your comment specifically mentioned Haskell-style monads (eg. a `Monad` trait), but that's not the only way to implement something like this.
> the last part is offensive & wrong
Quoting steveklabnik:
> it’s an open research problem if do notation can work in Rust. Until that’s solved at all, we’re just not sure it’s possible. ... "Open question" doesn't mean "impossible", mind you. But nobody has ever come up with a design. In the meantime, we have users to support...
Isn't this what I was saying? "We don't know how to do it, so we're going with the easier option."
Edit: To be clear, I don't think async/await we've ended up with is necessarily in the wrong direction. But I also don't think that "we thoroughly explored the design space of do/monads/effects and concluded that they were impossible to implement ergonomically/efficiently" is really true.
When folks say something is "impossible" in such a context, they mean "given the constraints", which include goals the lang team has for the language. An effects system is pretty heavyweight and may violate these goals.
This is separate from effect systems, which I never said was not possible. rpjohnst's parallel response sums up the key differences between monads and an effect system.
When people say that monads and do-notation don't work in Rust, they're talking about a user-level Monad trait and a simple CPS transform built on top of it:
- Such a monad trait is impossible to write even with HKT because it would have to abstract over both Option/Result (type constructors) and Iterator/Future (traits)
- Such a CPS transform would be extremely limited (not composable with built-in loop structures) and/or extremely tricky around TCE, lifetimes, and allocation (the typical type for `>>=` would involve `Fn` trait objects...).
Effects bypass this by leaving the CPS transform in the compiler, instead only exposing delimited continuations to userspace. Which is basically what async/await does, just non-generically.
You might want to do a Google Scholar search on Niko Matsakis and Aaron Turon before making silly claims like this.
This is probably only because you didn't develop the answer fully but i still struggle to see how monads (or other structures in that family) couldn't apply here: they aren't implemented in any way and just provide interface (eg not caring about rust at all, the semantic of that feature here is very monadic, it should mostly behave like the CPS monad). Anyway, i'm gonna stop arguing, look at how it is/will be implemented and hopefully wait for that essay of yours, to have a broader picture.
Higher kinded polymorphism results in trivially undecidable type inferences without something like currying; the restrictions needed to support it would be arbitrary and weird given that Rust does not have currying (essentially some rules to reconstruct the restrictions of currying in type operator context).
Instances of both Future and Iterator do not implement the Monad type class as defined in Haskell, because the "fmap" operator does not return the "self" type parameterized by a new type, it returns it own new type. This is because the state machine they represent is leaked through their function signature.
Do notation doesn't work with the imperative control flow that Rust has, which other people have already discussed.
...I saw this thread too late, hopefully you will still see this comment. I'm genuinely curious.
For Future types, the first thing could be called MakeReadyFuture(type), the second one Future.then(result => functionWhichReturnsANewFuture).
Future, Option, List, State/IO, etc.
And the do-notation that some comments mention become a long list of flatMap/filter/map invocations.
do {
futRecord <- getDBrecord(username, passw)
uid = futRecord.map(record -> record.id)
username = futRecord.map(record -> record.name)
futResult <- setUserLastLoggedIn(uid, now())
futResult <- checkUserACL(uid, operation.roles)
} yield username
This is just an abstraction, and it's not exactly a pretty one, but better than the flatMap hell in functional languages.The do-notation uses the same monad throughout, in this example the Future one. The final yield is a map, the "<-" lines all represent a flatMap.
Finally, Monads are very much like Vector Spaces in Maths. You have Axioms for both, and if the object satisfies them, it's a Monad or Vector Space!
However, I agree that the important step in understanding monads as an abstraction is realizing how map and flatmap make sense for many things other than lists.
I feel like, now that I have a certain level of intuitive understanding of what bind does, to me, the `>>=` symbol best represents what the operation does. It just looks like some kind of physical gadget that extracts a thing from a container, does something to it, and injects it into the next container. This is pretty ironic given Haskell's (somewhat deserved) reputation as impenetrable symbol soup. I don't know whether I'd endorse trying to use an intuitive mnemonic like this to explain monads to FP beginners.
For math it's kind of okay. Because there the point is to talk about that theory. You introduce definitions, and theorems, and use them in proofs. Or in calculations.
And even in math proofs usually come with a lot of explanation. (At least the better ones.)
In software engineering maintainability is important.
And, sure, we can just accept that Haskell is something that you can't learn by looking at real world Haskell code. After all, you can't really learn real world "research math" by looking at it.
But I think real world code, especially one that is looking for maintainers should not err on the side of indecipherability and inapproachability.
I continue to chastise Scalaz for its bad naming and documentation convention. (Haskell is ... well, it's irredeemable.)
And TypeScript is getting into this mess too. The documentation for new releases with new type system goodies (and I really mean it, I like powerful type systems, I just don't want to spend my life on understanding them, I'm happy to use them to get particular jobs done), but with barely enough documentation to let serious TS users with many years of experience understand it.
Thank-you for crystallizing and articulating a frustration I have often, as a beginner to intermediate Haskeller.
Or could you show a counterexample please?
http://blog.sigfpe.com/2006/08/you-could-have-invented-monad...
This one is my favourite, and does it well by example and walking through three different cases, then generalizing.
Maybe I’m misunderstanding something, but since people here are so fluent in concurrent execution, can someone point me to an in-depth explanation why are light-threads, coroutines, current-continuations, etc. so opposed by async keyword (and futures in general) today?
I have some experience with low-level runlooping via coroutines in luajit and understand it to the point to be able to create asynchronous system, consisting of mix of os threads and coroutines (and it worked smoothly until our project was closed due to company’s external issues). I can say, I never felt the need of something different, neither met the ‘complexity’ of everything-can-yield rule. And the possibilities that open, i.e. scalability of simple code around io and cpu cores is just outstanding. I am genuinely curiuos what’s so great (or different) in futures, which seem to me, for now, just poor man’s light threads implemented via lexical closure overhead along with syntactic snow, running on a single-core cpu. This topic seems to be so narrow that modern google is too shallow to answer that. I believe there should be a LtU or similar thread that discusses it in classic depth. Thanks in advance!
Light threads and stackfull coroutines require stack allocation. That is their cost and it is unbearable for system-level language priding itself in zero-cost abstractions. Eg. AFAIK, it's the main reason that is making Go calling C code slow.
Also, (again AFAIK) Rust stackless coroutines and futures compile down to state machines, so I don't understand the "lexical closure overhead".
There's nothing forcing futures to be "running on a single-core cpu" in Rust, since Rust can reason about thread-safety.
Personally, after couple of months of using JS at dayjob, I very dislike JS (as I thought I would), but I love coroutines/yield. I think they will be glorious in Rust: they will allow writting reasonably nice code with an amazing performance.
https://github.com/nox/rust-rfcs/blob/master/text/0230-remov...
There is an ongoing discussion in the C++ world between traditional C# style async, a more extreme non-type erased version (similar to the rust implemation I think but unsafe) and stackfull coroutines. Some (like me) hope that an hybrid solution might be possible.
Maybe what you have in mind, though, is the traditional argument, that there's no such thing as undelimited continuations. All continuations are naturally "delimited" by the boundary of the language runtime, operating system, etc, even with undelimited, multi-shot continuations. One-shot continuations make this argument more natural, because they have an obvious stack-based implementation for both delimited and undelimited continuations, and stacks are clearly delimited.
If you look at the performance numbers of these approaches, you'll see why stackless coroutines are desired.
I also want the ability to convert internal iterators to internal iterators with no overhead and even (especially) if the internal iteration function has not been specifically marked (i.e. no red/blue functions).
Hey, a man can dream.
You are literally describing stackless coroutines. And the generator state transform is that optimization.
If you want to get this without using generators explicitly, it's still stackless coroutines just not how Rust supports stackless coroutines. There was some discussion about making it more implicit but no progress was made in the implicit direction.
But as you might be able to tell from their APIs, you still have to allocate stacks somehow.
The trick that "stackless coroutines" can achieve, and do achieve in Rust, is not just getting a single allocation, but getting a single perfectly sized alloc every time. It is exactly big enough to hold all the data it could ever need to hold, no more no less. You mention we have a lot of annotations, but the annotations that I know of that can achieve that are annotating every yield point so we can compute the stack space we need to store at that point, which is exactly what async/await notation is for.
Interestingly this is similar to the analysis required to completely optimize a way the coroutine stack frame allocation. And because rust type system is good at tracking lifetimes and nested scopes, I think it might be possible to extend it ti guarantee this sort of optimization. Unfortunately this way beyond my pay grade.
For comparison, here is async/await in Zig: https://ziglang.org/documentation/master/#Coroutines
Zig decided to go the other way - when you async call a function, it does eagerly evaluate until the first suspend point. This is less overhead than immediately suspending, plus it removes the dependency of the language feature on a userland event loop. Users who want the immediate suspend feature can call a userland utility method of the event loop which suspends and then tail calls the async function in question.
For a further comparison, Nim works in the same way as Zig when it comes to eager evaluation. Something that I'm particularly proud of when it comes to Nim's async/await implementation is that everything, right down to the macro which defines what `await` means, is implemented in the standard library. The compiler only implements the coroutines. This means that the language isn't bloated by this extra feature and makes it much easier for developers to implement their own async/await.
I believe there was talk to do the same in Rust, but for some reason the developers decided to implement it in the compiler instead.
Clearly you need coroutines in the compiler, and the way Rust treats ownership means pinning is needed in some way (I've not been following closely so don't know if this is in libstd or a language-level feature), but the rest (according to this post) seems to be going in libstd so you can write your own if you want.
I'm referring specifically to the async/await syntax. Coroutines probably need to be implemented in the compiler, sure, but the actual async/await can be implemented with metaprogramming.
https://groups.google.com/forum/#!topic/flutter-dev/3R9qhjNG...
As for the rationale for why Rust is aiming for this behavior, I think this comment buried in the RFC discussion (https://github.com/rust-lang/rfcs/pull/2394#issuecomment-382...) sums it up (and makes note of Dart 2.0 as a contrasting example):
"A fundamental difference between Rust's futures and those from other languages is that Rust's futures do not do anything unless polled. The whole system is built around this: for example, cancellation is dropping the future for precisely this reason. In contrast, in other languages, calling an async fn spins up a future that starts executing immediately."
"A point about this is that async & await in Rust are not inherently concurrent constructions. If you have a program that only uses async & await and no concurrency primitives, the code in your program will execute in a defined, statically known, linear order. Obviously, most programs will use some kind of concurrency to schedule multiple, concurrent tasks on the event loop, but they don't have to. What this means is that you can - trivially - locally guarantee the ordering of certain events, even if there is nonblocking IO performed in between them that you want to be asynchronous with some larger set of nonlocal events (e.g. you can strictly control ordering of events inside of a request handler, while being concurrent with many other request handlers, even on two sides of an await point)."
"This property gives Rust's async/await syntax the kind of local reasoning & low-level control that makes Rust what it is. Running up to the first await point would not inherently violate that - you'd still know when the code executed, it would just execute in two different places depending on whether it came before or after an await. However, I think the decision made by other languages to start executing immediately largely stems from their systems which immediately schedule a task concurrently when you call an async fn (for example, that's the impression of the underlying problem I got from the Dart 2.0 document)."
* Function call
* The coroutine creates its frame (I believe this maps to "future creation".)
* Function return (the injected immediate suspend)
* (sometime later) Put the work on the event loop
* Function call (suspend resume)
* Function executes until a suspend point
Eager execution has no injected immediate suspend, so it looks like this:
* Function call
* The coroutine creates its frame
* The function executes until a suspend point (function return)
That's it. The coroutine is responsible for making sure that it gets resumed appropriately if it suspends. Await doesn't do anything with an event loop; it just suspends and then puts the suspended coroutine handle in the target coroutine's frame with an AtomicRMW. If it turns out the target coroutine already completed, then it grabs the result, destroys the target coroutine, and cancels suspending (which is a jmp instruction). Otherwise the suspend completes and the target coroutine will resume the suspended one when it completes.
> If you have a program that only uses async & await and no concurrency primitives, the code in your program will execute in a defined, statically known, linear order.
This is true with eager execution as well. You can see some examples here [2].
One more side note, how eager execution was valuable to me in the self-hosted compiler:
pub async fn renderToLlvm(comp: *Compilation, fn_val: *Value.Fn, code: *ir.Code) !void {
fn_val.base.ref();
defer fn_val.base.deref(comp);
// ...
}
At the callsite we make an async call to renderToLlvm, passing in a ref-counted fn_val. If the body didn't execute immediately, the callsite might deref fn_val and destroy it before renderToLlvm gets a chance to add a reference, but because of eager execution, the ref is guaranteed.[1]: http://llvm.org/docs/Coroutines.html
[2]: https://github.com/ziglang/zig/blob/363f4facea7fac2d6cfeab9d...
IIRC it doesn't. Coroutines would induce a significant cost (one stack allocation per future) which would make futures unsuitable e.g. for embedded systems.
My understanding is that Rust will compile futures into state machines. So a future would be a closure with a data object that's a sum type with one variant per suspend point, that contains all the variables and references that are in scope at that suspend point. (That's also why Pin needs to be added for futures/generators: This set of variables may be self-referential.)
A Rust future is also "responsible for making sure that it gets resumed appropriately if it suspends." Rust's await is even cheaper than C++/LLVM's- it doesn't do anything with an event loop either; it just returns, without atomically messing with any coroutine handle (which doesn't exist).
Further, Rust futures are poll-based, not callback-based. Your bullet list looks like this:
* Function call to the main entry point (which is really more of a constructor)
* The future creates its frame as a value type (not as a heap allocation) and initializes it with the call arguments
* Function return (not a suspend, because it returns an initialized future rather than Poll::Pending- so again more of a constructor return)
* (sometime later, often immediately) Function call to `poll` (initial resume)
* `poll` executes until a suspend point (function return)
The initial call to `poll` can happen because the constructed future was put on the event loop and scheduled, as you describe, but it can also happen merely because its caller is also a future that is already running somehow.
Also, note that "the event loop" is a rather loose concept here- it's just "the thing that calls top-level `poll` functions." The language doesn't ever submit anything to it, or signal it, or anything. Each call to `poll` receives a handle to that top-level caller, and each leaf future stashes that handle somewhere that will signal it when it's ready to resume. This even works in embedded microcontroller scenarios.
And finally, your eager execution example is addressed in Rust in two ways:
* First, ownership and the borrow checker prevent dangling references like this statically. In your particular case, the caller would increment the refcount as part of cloning an `Rc`, and then move the clone into the future. (Unless, of course, the future was short-lived enough, and you decided to take advantage of that to pass in a `&'a Value.Fn` instead.)
* Second, because future construction (the first three bullet points) and future execution (the last two) are decoupled, you can write your own function that does any extra construction work in the cases that it's actually necessary. For example, using Rust's syntax for async blocks:
pub fn render_to_llvm(.., fn_val: *Value.Fn, ..) -> impl Future<Output = ()> {
fn_val.base.ref();
async {
defer fn_val.base.deref(comp);
// ...
}
}This seems wrong. There's no more overhead for Rust's approach- which to be clear immediately suspends only the callee, not the entire stack of async functions. In fact, if you immediately await the future, the control flow is literally no different from eagerly evaluating until the first suspend point.
There is also already no dependency on a userland event loop. This is, again, because only the callee is immediately suspended. It's still entirely up to specific leaf future implementations to interact with (or not) the event loop.
You can write an executor with a fixed-size queue as well, and not require dynamic allocation at all.
Being able to do this is a hard constraint on the design.
I have had the old macro based async code in Rust running on a Cortex M device, completely runtime free. Once the TLS stuff is sorted I plan to port this forward to work with the builtin syntax.
Yeah the Pinning stuff is supposed to help with this as I understand.
If anywhere down the callstack a function needs to await something it has to become an async function. And this needs to be done to the whole callstack recursively. So over time more and more functions of every codebase turn into async functions.
I’m curious if discerning which is the best type is possible to tool automatically with static or tracing analysis.
Rust has gone back and forth on this. Pre-1.0 versions of Rust had a "green threads" mode. It was removed in favor of OS threads because it wasn't worth the complexity.
Go is a language in which "all functions are async" and it's pretty cool. But it needs a runtime, and interop with C / C++ has some real overhead related to switching to a C-style stack. Rust did not want to make those compromises, it wanted to be appropriate for bare-metal or kernel programming, and for libraries used by big high-performance C++ applications. In these cases, Go's model does not work well.
I personally use Go a lot, it's great for network services and command-line tools, and that's mostly what I do.
I don't think there's anything about Go's concurrency model that precludes bare metal; I think it's just that the maintainers didn't want to support a bare metal runtime. In fact, I recall at least one other project that modified Go's runtime to support exactly this. I'm also not sure how much of the C interop overhead is due to the different stack models and how much is due to GC concerns or other things. But generally I agree with you--in practice Rust is well suited for these types of tasks and Go is not.
Async has come to signify a very specific implementation strategy (i.e. stackless coroutines). Go definitely went in another direction.
In practice, most of your threads are in one form of wait loop or another and you've just got polling both inside the threads and with the scheduler.
Have a look at Erlang if you want a better model :) the erlang "processes" (different to os processes) can intelligently only wake up when there is work for them to do.
For a language to efficiently use cores, it really needs to include it's own scheduling.
There are also costs in setting and checking those locks etc. Sure, you can build solutions to optimise a broken model, which we've done over decades (with quite a bit of success!), but it doesn't make the model less broken.
Sure you do, there is noticable memory and ctx switching overhead(for large no. of them ofc) in stackful coroutines.
Erlang works the same way. The VM scheduler will only context switch on a function call. for or while loops don't exist in Erlang which means there is no risk of blocking the scheduler.
The reason that m:n threading went out of style is that, for CPU bound tasks (where you want to run exactly as many threads as there are CPUs), it is just useless overhead, while for the hundreds of thiusands of IO bound threads scenarios, the cost of stack switching is dominated by cache misses anyway, and the cost of calling into the kernel is amortized by the fact that IO requires a call inti elevated privileges anyway. At the sametime kernel threads scheduling has become very fast and userspace threads, which requirea whole stack of their own are not significantly more lightweight than kernel threads.
The modern async model is a compromise. On one side, the 'threads' consist of a single stack frame are very light weight, on the other side there is no generic userspace scheduler, but scheduling is fully controlled by the application.
You can build threading on top of it but also all kind of other abstractions.
Futhermore, in any complex asynchronous I/O app most functions will end up being tagged async. The only ones that won't are simple leaf functions that wouldn't implicitly yield, anyhow. If someone has the bad idea to, e.g., put a yield point in non-obvious leaf functions as part of some kind of hack (e.g. logging, tracing), they're gonna do it in the async/await case, too, because they're already convinced it has value. Having to a drop a few more annotations here or there won't stop them from breaking the app.
Lastly, the bugs that do occur in these sorts of cases usually have to do with unpredictable latencies violating implicit or accidental ordering assumptions. async/await doesn't mitigate that at all because latencies are just as unpredictable. The solution, as always, is to avoid these ordering dependencies by not sharing mutable state.
This defense of async/await is a red herring.
And it’ll likely pull in some nice features from academia like, building a dependency graph of async statements so it can automatically reorder async statements to get optimal concurrency.
And the reordering of statements exists too. It's called ApplicativeDo[0] and is heavily used at Facebook[1]
[0] https://www.microsoft.com/en-us/research/wp-content/uploads/...
Making every function call asynchronous is likely to make most programs a lot slower.
If we did the whole problem async (async by default as I interpret it) it would require so much bookkeeping to keep causality that would make it unreasonably difficult. For things that require causality like that, we are lucky that the default mode of computation is synchronous because it fits my problem domain perfectly.
That's why I say "not everything is web" because while async maps well to problems like a server-client type application, it doesn't map well to all problems like my problem for example. Also, as someone else said, sync is easier to reason about for problems that don't require a lot of parallelism.
At best, some of your CPUs would be able to run calculations for two especially fast-to-compute pieces of volume space, while some other CPUs would be busy computing a particularly gnarly block of the volume space.
For those CPU bound jobs that do not fit in the classical openmp style scheduling, more dynamic async style scheduling might be appropriate (cilk style work stealing for example) but the async granularity is hardly ever the function boundary.
Which doesn't mean it's bad. I'm excited about tokio and rust's async story. But I love that I get to choose.
Just dealt with that in a .NET app that uses a 3rd party SDK that heavily uses async/wait. I had to build a whole layer of state management to just handle shutdowns. Simple threading would have been much easier.
It's different with server side code. There async/await is very nice.
[1] https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
In a browser on or server side it's different.
Continuation passing style means passing along your continuation to your callee, and is the implementation of the async model. But normal calls do pass their continuation: the return address, pushed on the stack, is the continuation function, and the stack itself is the scope, and together they are a closure. Make the stack a first class value that can be switched, and you get to continuations - and that's what the runtime gets with green threads.
It means much harder interop because you can't have native frames on your stack. But you have the same problem with async.
With the normall call/return model, the implicit continuation in a normal call stack can only be accessed via return. Green threads do allow accessing the continuation at specific yield points (although most green thread implementations do not expose it to the user) but because the yield point are much fewer than every single call, a whole stack can be allocated and used in one go. There aren't in fact many issues with interoperability as foregin functions or even os calls can be accommodated on this stack.
Async is a tradeoff, basically the programmer is responsible of marking functions that are to be CPS transformed and that need their stack frame reified. This way no full stack is needed nor there are the performance and interoperability issues of full CPS. You end up with the blue/red functions issue though.
Note that you don't get the continuation in hand for the async-everywhere advocated by GP, only blocking operations get the continuation, and most of those are implemented by the runtime.
According to a famous quote, the source of which escapes me, "await does not wait for anything, and async is not asynchronous." (This could be specific to C# though.)
https://docs.rs/tokio-threadpool/0.1/tokio_threadpool/fn.blo...
Or you go with cactus stacks which have their own set of issues.
More wasteful than writing each function twice? See C# for endless examples of libraries that have manual duplicate X() and XAsync() for every single method. If you compare the two they're almost always have the exact same body, except the async version has "async" and "await" peppered in and calls duplicate XAsync methods (which are implemented the same, except they call XAsync methods... you might be seeing a pattern here).
If I'm not mistaken most rust doesn't use separate compilation and is built from source. In this case the compiler can treat it just like normal generics and just not compile it until it's used. And even with separate compilation for, e.g. a library, a very basic LTO would trivially trim out the unused methods when compiled into a final binary. And if you don't want 2x library itself, even then you can just use an escape hatch to turn it of on the crate/impl/function level and it's still better for you than the status-quo where you have to manually add a bunch of tokens and junk that the compiler already knows how to do precisely.
If anything is wasted it's human effort because we're wasting our time doing something the computer can do better and faster.
> "annotations in the object file to describe the frame size"
Wouldn't you need this anyway for anything async, even if it's manual? This seems like an argument against async in general, not automatic async.
I think that a good compromise would be implicitly inferring the 'async-ness' of template functions instantiation (or whatever they are called in your language of choice) based on a magic continuation parameter, which would also go around the separate compilation issue (assuming the language does generic monomorphization).
The point is that only at the very top of the call stack or the very bottom should you ever need to care or think about whether code is async, because 99% of the time (every bit of code between main and syscalls... which is almost all of it) that's the only place you should need to care. This can be completely automatic because the compiler already knows where to insert await/async, it only needs to know when to use async (decided by main) and what to call at the end (decided by the syscall funcs).
If I have a sync function and I call something async, I should be able to just .wait() it, I think, and get the sync behavior.
The Future object seems to do the all bookkeeping needed by the borrow checker (which I believe is what rust uses to track the lifetime of objects).
- The continuation which is passed to .then, and which is typically a closure, must be type-erased, which requires an allocation. Storing the continuation in a std::function would allocate too. A short glance at https://github.com/scylladb/seastar/blob/master/core/future.... also confirms that there is a make_unique there.
- Since continuations are most likely not called inline on completion but deferred into the next eventloop iteration there needs to be a dynamically sized queue to hold the ready continuations. I am not 100% sure if that's the case for seastar too, but I would guess so.
The queue might be dynamically sized which might eventually require allocation, but that can be ammortized across many futures. A large enough queue might never require reallocation.
I'm also not 100% sure as I have never used seastar.
Linux is especially problematic, asynchronous IO has arrived late, years later than IOCP and kqueue. It took multiple kernel versions to make it usable, and still the APIs are questionable, e.g. files and sockets use different ones.
Even MS failed to do it right, see a bug I found: https://github.com/dotnet/corefx/issues/25066
This is a classic "being traffic" comparison. If you're in a traffic jam, you are as much the cause of it as anyone else is. Likewise, if the ecosystem is immature, you are as much the reason for that as anyone else who doesn't participate in it.
Perhaps this is just obvious. But I do frequently see people pretending to be objectively objecting -- complaining about traffic that they are themselves the cause of as if they didn't already know the answer to their question.
Going to Go with its LWT goroutines was awesome. Would never want to go back to async, which is really just a manual way of implementing LWT.
The downside of Go’s approach is the overhead of calling into C; Rust can’t afford this. Go can. IMHO both languages are making the correct choices with regards to their constraints.
In python you can chose what event loop to hook to async/await, and hence we have gevent, qt, uvloop, twisted, asyncio, tornado, trio and curio as competing implementations.
They are very difficult to mix, and their ecosystems are mostly isolated, dividing the man power to add features, fix bugs, provide support, create libs or frameworks and write docs or tutorials.
Another problem is that you have (except for gevent which causes other problems by monkey patching the stdlib) to setup the event loop explicitly.
Those mechanisms are complicated, easy to get wrong, confuse beginners, make docs introduction long and annoying or misleading before getting to anything interesting.
E.G: to use asyncio, you are exposed to an event loop, an event loop policy, awaitable, coroutines, coroutine functions, futures, tasks and task factories.
But there is worse... You can setup any event loop any way and time you want, and because the api is public, another lib can come and swap it. No lib to my knowledge provide any form of locking.
This leads to some weird situation where libs are considered so low level you end up writing wrappers on top of it (e.g: https://github.com/Tygs/ayo) just to be able to start using it sanely.
Now compare with JS.
I'm really not a fan of the language. However, even when you can't use async/await, using a promise is straightforward. You don't have to bother about creating the loop, starting it, stopping it, cleaning after it has stopped. You don't have to wonder if somebody is going to swap the loop. You don't have to get a reference to a loop to schedule anything. Actually you can mostly ignore the loop and just code the solution to your problem.
Now this makes JS dependent on one loop implementation for each runtime. Also you can't code any new async feature in JS, only use the existing ones. It's probably not what you want for a language like rust.
However, you should learn from the python ecosystem fragmentation and overly exposed low level API to avoid the same mistakes.
Somebody that just wants to use async/await should not have to learn how the implementation works in details, nor take so much precaution to avoid implementation lock in or break somebody else work.
And you really want a federated ecosystem. Having 7 incompatible websocket lib sucks.
Although perhaps the situation in Python could have been mitigated somewhat if a protocol was defined early that implementations could have adopted? Context managers, iterators and decorators all work nicely together, even across language boundaries via the extensions API.
Yes. That's one thing the rust community has to get right.
async/await was supposed to be that, but it's only a protocol to define what blocks/doesn't and when you allow context switching. An event loop also has the notion of scheduling, getting a reference to what is scheduled, request the result or error on said scheduled thing, or cancel it. And even loop must bridge different implementations of concurrency (e.g: asyncio.run_in_executor). An event loop also has a life cycle, which includes at the very least a setup and a tear down. An event loop must integrates in an environment, like what do you do when you have several loops, or if you run one loop per threads ?
So you need to define a general behavior for all that. Then let anyone write the implementation the way they want.