Await in Rust
docs.rs
docs.rs
Like in real life I can do something synchronously (meaning paying full attention), or I can do the same task asynchronously (watching youtube videos in the background while cooking, for example), but it's still up to me. Being "async" is not a property of the "watching youtube" action, it's the property of the caller running this action.
That's the reason why CSP-based concurrency models really works well for me – it's just so easy to map mental models of system's behaviour from the head directly to the code. You have function/process and it's up to you how do you run it and/or how synchronize with it later.
Async/await concept so popular in modern languages is totally nuts in this aspect. Maybe it's just me, but I find myself adding more accidental complexity to the code just to make async/await functions work nicely, especially for simple cases where I don't need concurrency at all, but one "async" function creeps in.
Namely in platforms like Node the convention is to not do synchronous I/O and as a communication tool - everything that doesn't use async/await is promised to never block or perform I/O.
You might claim it's a lower level of abstraction but it makes concurrency a _lot_ simpler and it's why often when I write go I have to write 35 lines for something that would take me 5 in C#, Python, JavaScript or now Rust.
How do you deal with a situations when you have non-async/await code, and then you need to call just one function, which happened to be "async", and you can't do until you mark all your functions async as well? I find myself in this situation too often with Dart/Flutter, and it's especially embarrassing experience as it happens usually with a simple parts of code, where I don't need concurrency at all (like, get local directory name and store file to it). This "forced refactoring" also hits language design corners like "you can't call async functions" in constructor or `iniState()` and I end up writing async wrappers and other helpers functions just to make the whole thing work.
> it makes concurrency a _lot_ simpler
I find people put different meaning in the word "simpler". Do you mean that code requires less lines/chars to write or that it's easier to read and understand (those are totally different properties, often conflicting)?
It'd be way more convenient if Node just assumed it couldn't carry on with the next operation until the current one was done, even if the current operation is async I/O—feel free to go off and process something else, just no more of my code—unless instructed otherwise. IOW, caller decides, yes. I 100% do not understand the appeal of a system where you have to specify that you do want the next instruction to wait on the result of the previous one. "But you'll block node's single thread during your I/O!" no, it can go do something else, just no more of this particular line of logic in my code unless I specifically tell it otherwise, thanks.
Caller decides is both easier to follow and more useful.
But to be clear, "it'll block Node's single thread!" is not the reason for explicit await. The reason for explicit await is to guarantee that after calling a function foo(), only the state that foo() changes could have changed. Whereas after awaiting on an async function bar(), or any other promise, literally any state-mutating code in the entire application might have run.
Obviously plenty of languages have this problem and deal with it, usually with locks like in Go and Java, or by banning shared mutable state entirely and only allowing message-passing for communication between concurrent "lines of logic" like in Erlang. (As far as I understand Go basically tries to get the best of both worlds by strongly encouraging message-passing, with locks intended only for advanced situations like custom concurrent data structures.)
But there's no way for Node to change to either of those models without breaking every single nontrivial piece of code ever written in it. And apparently Rust intentionally chose not to go with those models because they require a runtime and make compatibility with C very difficult, well-known advantages over Go.
As for calling async from sync, I imagine you just create a runtime and block on it.
Each method now returns a Task, and each method name ought to have a suffix of 'Async'. It sucks that changing this will require all consumers to also modify their signatures as well. But I understand why it is necessary, the compiler needs to mark each 'step' in the asynchronous function to create a kind of state machine.
If you don't want/need to expose your async-ness, you just synchronously block on your async callees instead of awaiting them.
Just call the poll function on your Future.
Just because Javascript sucks doesn't mean the similar feature in Rust does too. :)
This depends on what do you want to happen. Do you want to block on the one async function, or do you want to turn your current function into async and propagate the asyncness to your caller?
Both options are available to you in Rust.
But yes, in Rust async is not a plague that irredeemably taints everything. You can choose to use anything that is async synchronously if you choose to, just by calling .poll() on the future at the callsite.
I don't think that's true. See fs.readFileSync for example. You could say that's an exceptional case (which I think it is), but it means that there's no "promise" that functions without async/await don't perform I/O.
That's a good thing :]
If `fs` was promise-aware and top-level await existed, it would be a different conversation. As-is, for writing a one-off script, avoiding callback hell is pretty nice.
https://nodejs.org/api/fs.html#fs_fspromises_readfile_path_o...
Steve Klabnick provided some useful context on that thread for the Rust decision specifically. More generally though, you've nailed my primary misgiving:
> Synchronous or asynchronous is something up to the caller to the decide – not to the function itself.
Erlang/Elixir and Go place the decision with the caller, not the callee. That just seems much more sensible.
An alternative is perhaps to see async/await as dataflow going through the awkward teenager stage. Maybe one day we'll wake up and languages in general will have fully embraced dataflow variables as first order constructs.
I dislike how this stuff "infect" all your code calls:
https://journal.stuffwithstuff.com/2015/02/01/what-color-is-...
And the problem is rust is that we already have other stuff that infest everything: Borrows, mutability, ownership...
Unlike in JS, it really doesn't. In Rust, at any point you have the option to run a Future returned from an async function to completion in a sync function (which blocks the thread) and stop the spread of async.
I don't think anyone disagrees that the annotation of these concepts play a very heavy role in the type system and cognitive burden of writing code. But there are really three alternatives you can do here:
1. Do nothing, and rely on programmers to manually remember how to do all the necessary bookkeeping themselves. This is better known as "C", and the sheer amount of CVEs and other problems in code written in C is ample evidence that this approach just doesn't work.
2. Make the runtime do the automatic programming. JS is the exemplar of the pattern here. But this approach requires the VM to do garbage collection itself (which adds uncertain hidden overhead), generate code to handle arbitrary shapes (which can result in hidden performance cliffs). And heavily dynamic languages also turn compiler errors (such as fat-fingering a variable name) into runtime errors that can be accidentally ignored.
3. Shift the annotation into a mandatory part of the language, as Rust does. It makes the cognitive burden of writing the code higher, it requires more coding, but the resulting code tends to have the highest performance and lowest bugs of any of these approaches.
The tradeoff you have to make is between "no safety", "compile-time safety", and "runtime safety"; the costs you have to weigh are the costs of testing to uncover bugs, the costs of programmer cognitive and annotation burden, and the (often hidden!) costs of the runtime to dynamically enforce the conditions.
Rust has chosen to live in the camp of low implicit runtime cost and high compile-time safety, and it is really to pay the price for that with extra annotation burden. You may disagree with that choice, but it is a conscious choice that has been made.
This is what I want. I just not like the specific implementation of awaits. I'm not much against it: i could live with it (I have used C#/F#/JS that is similar).
The thing is that still split the world even if you don't want it.
I think CSP/Actors have lesser cognitive load but of course I don't know how make them work without a runtime..
While also making your software more resilient to failure because error handling is easier as you can contain the code into small processes, communicate through messaging, let processes fail, and handle/restart them from the supervisor. Using the OTP supervisor model encourages you to think through those things, like error handling and the lifecycle of your code, so it doesn't impact other parts of your running app. A lot of exception handling in other languages is often bolted-on after-the-fact and isn't part of the typical design, at least for more inexperienced developers.
It's best to commit fully to either a full-featured VM or Rust/c low level OS style IMO.
The latter group can still support a more complex actor model with libraries and frameworks (ie, https://actix.rs/) which is more flexible but nothing beats having it be the 'standard' philosophy of the language and fully baked into the runtime.
I consider Rust to be an antithesis to your remarks because it has realoy good tools to prevent concurrency issues and yet still wants it.
I think on a deeper level it's because there are times where the thread isn't really doing anything and is instead waiting, and that time could be better apent doing something else. That requires sync, either with callbacks, promises or sync/await.
But "wait until there's eol, than call this callback" or its cousin the future which is "when complete, read here to find the callback(s) to call" which let's you fill in the callback after the call, that's inherently an async call. Same way sending a letter forces you to wait. Sure you could stand by the mailbox until the reply arrives, but that's just converting the async call to synchronous by waiting.
Though I don't really see how futures are that different from a cps, especially if you view them as chaining .then calls, which are exactly continuations. Await just let's you drop some syntax, and unify scopes.
LoadWasher() (sync)
washerRunning = RunWasher() (async)
var cleaning = CleanKitchen() (async)
wetClothes = await washerRunning;
LoadDryer(wetClothes)
clothesToIron = await RunDryer() // while we're awaiting we walk back up the stack and can continue cleaning the kitchen until the dryer is finished
IronClothes(clothesToIron) // this is just happening synchronously
await cleaning //last job to do so we need to make sure we finish!
Don't know if that helps, but I think it's an interesting analogy LoadWasher() | RunWasher() | LoadDryer() | RunDryer() | IronClothes()
CleanKitchen()
where `|` is piping the output of one operation into the next one (as per *nix shell). The two lines are independent and so can be scheduled according to resources available.Doing so makes the inherent dependencies trivially clear. It doesn't matter which are "long" running and which aren't: it's not possible to run the washer until it's been loaded. Using async/await conflates two things:
1. What's "long" running, for some definition of "long" (--> make it async)
2. What's dependent on what (await)
--
EDIT: fix formatting & added missing "RunDryer()" step
Then wrap it in a try/catch or pattern match the result to handle the various errors different errors...
No, the point is you don't _need_ async. At least not as a language primitive. Instead, one expresses directly where the data dependencies lie. If an operation in a pipeline blocks, the runtime is at liberty to switch to another task. That already happens today, whether in language runtimes (e.g. Node) or the underlying OS. I don't need to tell the runtime "this operation might block". It already knows.
Instead, we've explicitly defined intent: when `RunWasher()` finishes, `LoadDryer()`. It doesn't matter if the washer takes one micro second or one hour; I can't load the dryer until it's done.
>Then wrap it in a try/catch or pattern match the result to handle the various errors different errors...
True, error handling is missing. But then it's not in the parent example either. The requirements aren't fundamentally different in either case.
If you replace the machines with people, it makes much more sense for the function definition to define what is and isn't synchronous. If you have to do all the chores yourself, just call the functions synchronously. If you have family members you can delegate some work to, you can call the functions asynchronously and wait for them to complete while you're doing your own work (or twiddling your thumbs).
Elixir allows you to do this very cleanly. https://hexdocs.pm/elixir/Task.html#content
I'd rather dirty the signatures of blocking/synchronous methods then of the rest.
If I do what I've always done in rust, which in use threads, I get lovely easy parallelization. Async only seems useful when your program doesn't do anything non trivial, but just dispatches to other things, as any one function that takes a long time blocks all other async functions until completion?
- What is the best approach to encapsulate blocking I/O in future-rs? — https://stackoverflow.com/q/41932137/155423
Also, blocking on a future is trivial (just call poll).
So I would argue that async/await does a better job of leaving it up to the caller to decide: use await if you want async, use poll if you want sync.
You can't (in safe Rust) easily replicate what async does for you because it understands how to handle borrows across yield points, which turn into self borrows when you turn the stack into a concrete object in Rust. That's not really an issue in garbage collected languages so it's not quite as much of a necessity to have a specific keyword, but it's still inconvenient in most languages to write everything in CPS or use combinators to acquire thunks.
As for why you would want a function that's not asynchronous... well, various reasons, but one is performance. Creating a thunk and then sending it to a runtime which decides what to do with it (or polling) usually has some overhead compared to linearly calling a function on the stack.
Another reason is that, in Rust, top-level asynchronous functions normally need to be quite conservative with what they own if you want to use them with an actual asynchronous runtime--many of them like to send the thunks between threads in a static thread pool, which limits you to thread safe constructs and owned (or static-borrowed) data. As a result, even if there was no overhead for using an asynchronous function synchronously, you would still in practice have to either sacrifice performance and generality by keeping your data owned and thread-safe (to make things work with thread pools), or embrace all the usual stuff you would do in synchronous code (like borrow things from a previous stack frame) and lose the ability to use your async function at the top level in most existing asynchronous runtimes. So from that standpoint, it's not really up to the caller. This is again not really an issue in garbage collected languages that only allow references to managed, heap-allocated objects.
There are more reasons than that (being able to efficiently call out to blocking C APIs, wanting to use builtin thread locals across function calls without worrying about how the function call is going to be handled, current issues with recursion, etc.). They may be considered artifacts of the implementation or otherwise resolvable--I'm not necessarily saying they aren't--but they are reasons in practice why you want to have synchronous functions available.
So tl;dr I think most people's immediate intuitions about how async/await should just be sugar, or async should be up to the caller (i.e. all functions should be async), don't straightforwardly apply to Rust even if they are valid for most other languages.
This is a exactly how I feel as well and I write Rust code every day.
But no matter, I will suck it up. The language is evolving and I think async/await has, seemingly, been one of the biggest divisive feature the community has faced. I will still use as there are so many other benefits this language brings.
Perhaps because you hadn't had the displeasure to work with callbacks, or promises? Which async/await greatly simplifies.
That's probably the real reason why so many people are so happy and vocal about "simplicity" of async/await – after experiencing concurrency only with callbacks in JS, async/await looks like a blessing.
You can't call it magic just because you haven't read what it desugars into. Otherwise anything you don't know could be called "magic".
The same way code using functions is easier to read than looking at the exact same code expanded in 20 places in your program, even if a function call also has "magic" (passing the arguments on the stack, replicating the behavior inside the function as it was written in the place of invocation, sometimes inlining, returning, moving the stack pointer again at the end, recursion, and so on).
You learn once how function works and that's it. More compiler behind the scenes work ("magic") than writing the same lines again and again, but easier to read.
(In computing there's not magic. There are abstractions, and it's not that always less is better).
To validly complain about an abstraction the abstraction should make reading the code (or other aspects, e.g. performance) worse.
Merely complaining that "it hides things" and that you "have to learn it to know what it does" is not a valid complaint.
That's literally what abstractions are supposed to do: hide things and introduce new things to learn (the abstraction.
So you can't use the fact than an abstraction like async/await hides things ("magic"), or introduces something new to learn, as an argument against it (unless you're against all abstractions, but this train has long sailed, and the question whether abstractions can be good is settled: yes, they can).
So, the whole point to judge an abstraction is whether the new thing it introduces makes things easier _after_ having learn it or not.
So far your arguments were just that it hides things (magic) and that callbacks made what happened clear (so, again, the hiding aspect).
So, tell me, which thread did that callback run in? And what happens if I need to grab a lock in the callback?
That's "magic".
First he starts with a rant about the color of a function. In that sense, all typed languages have a color, aka the type of their return parameters. Sure, there a advantages to dynamic languages, but they're going out of fashion now for some good reasons.
Then he lists 5 more specific drawbacks of async code in 2015 Javascript 2015. All 5 aren't issues in Rust.
#1. Every function has a color: that's how typed languages work
#2. The way you call a function depends on its color: the rust compiler doesn't let you get this wrong and provides cut and paste corrections when you do
#3. You can only call a red function from within another red function: Use Future#poll
#4. Red functions are more painful to call: This is in reference to callback syntax in old node.js code and is fixed by the await sugar even in modern Javascript
#5. Some core library functions are red: The clause "that we are unable to write ourselves" is untrue in Rust; you should never have a need to drop down into C to write anything. And this is really just a variant of #3; if it's trivial to convert async to sync this wouldn't be an issue.
async/await aren't strictly necessary, but to avoid them you need one of:
1. a sufficiently smart compiler with whole program compilation, or
2. to compile two versions of every function, an async variant and a sync variant, or
3. every async wait captures the whole stack, thus wasting a lot of memory.
Async/await is basically a new sort of calling convention, where the program doesn't run in direct style but in continuation-passing style. This permits massive concurrency scaling with little memory overhead, but there other tradeoffs as per above.
Async/await makes perfect sense for Rust which wants to provide zero-overhead abstractions.
[1] true non-tail recursion is a problem of course, but it is ok for the compiler to give up if it can't statically prove that the recursion depth is less than a documented maximum.
edit: rewording
int bar(int i) { return i+1; }
/* bar takes a implicit return address in [esp+0] */
noreturn foo((*ra)(int),int x)
{
/* where (*)(RetType) is a return address */
int y = bar(x); /* x86 call pushes eip */
ra return bar(y); /* pseudo-tail-call (push ra) */
}
int quux(int x)
{
(*ret)(int); ret = (return); /* get implicit RA */
foo(ret,x);
} // unreachable
You'd also need to adjust the stack and/or frame pointer though, which is probably going to be the hard part.For example: convert an internal iterator to an external one (in pseudo C++, yes, yes, I know):
// This is a bog standard template function. Cont can be anything
// here as long as it is callable. It literally does the same thing
// as for_each
template<class Cont, class Iter>
Cont inside_out(Iter begin, Iter end, Cont yield){
// as long as yield lifetime is confined to the scope of
// inside_out; when Cont
// ends up being some compiler internal (possibly unique) magic
// type it should be possible to CPS convert both inside_out and
// for_each. Both rust and C++ monomorphize templates anyway so
// instantiating dedicated CPS variant is not a problem.
std::for_each(begin, end, [&](auto&& x) { yield(x); });
// for various reasons we need to return the continuation here. A
// proper language would tail call to it.
return yield;
}
...
// this is basically call/cc. It uses magic to get the current
// continuation, then invokes its first argument with it and the
// remaining arguments. The current function need to be
// Note: template functions are not first class in
// C++, but we take some liberty here.
auto cont = continuationify(inside_out, my_vec.begin(), y_vec.end());
while(not cont.done())
std::cout << cont();
if inside_out was not templated, but instead Cont was type erased (a trait object in rust parlance), then continuatinify would need to allocate a dedicated stack and fall back to good old userspace context switching (the compiler is welcome to pierce through the type erasure layer and optimize anyway, but it would not be a guaranteed optimization).Rust is close enough to C++ that it would work in the same way. I've been planning to write a paper for the C++ committee to add fire to the current coroutine controversy and proposing something like the above (which would unify stackless and stackful coroutines).
This is a bit mistaken; ancient versions of Rust (but not so ancient that they weren't using LLVM) were, like Go, everything-is-implicitly-async, and Rust used the same strategy as Go to manage stacks (to the extent that the old language reference listed Go in its (rather long) list of precursors, specifically calling out its split stack approach).
The ultimate reason that Rust switched away from Go's approach is for interoperability with C. If your stacks are the size of C stacks, and if your language doesn't feature a pervasive runtime, then interoperability with C is nearly trivial, and this was something that Rust dearly desired (and has benefited greatly from, IMO).
Ultimately one should not look at Rust's chosen approach and assume that it represents some grand rebuke of alternative approaches to asynchronicity. If you have different constraints, then a Go-style approach is lovely. For a language at the level of Rust, different tradeoffs may dominate.
in .NET stacks are different, the runtime is relatively large, yet interoperability with C is nearly trivial. From my PoV at least, I never worked on compilers, only used them a lot.
Declare a delegate in C#, apply [UnmanagedFunctionPointer] attribute specifying calling convention (on Linux you’ll likely want CallingConvention.Cdecl), pass to native code specifying [MarshalAs( UnmanagedType.FunctionPtr )] in the prototype of the DLL/SO function which accepts that pointer.
This is even more convenient than doing that in C. In C, function pointers are stateless, typical design pattern is pass accompanying `void* context` for the state. Not needed in C#, that delegate can capture whatever it wants, the runtime will do it’s marshalling magic, generating a native function and somehow associating it with the captured state.
The only caveat is lifetime. Function pointers can’t be retained by native code, they’re too simple and don’t have a ref.counter. You must ensure the .NET delegate is not garbage collected while it’s referenced by the native code, or the app will crash.
That amount is really small for .NET. See this doc https://docs.microsoft.com/en-us/cpp/dotnet/calling-native-f... it says “PInvoke has an overhead of between 10 and 30 x86 instructions per call.” That’s barely measurable. Close to the cost of a mispredicted branch. 4x cheaper than a single cache miss.
That’s if you do it correctly, i.e. only marshal value types or arrays/pointers of them, don’t use custom marshallers, etc. Also it’s important to specify correct attributes. To call `int func(const sStruct&)` C++ function, you should specify `[In] ref sStruct` in C#.
1. They may not have access to the representation the subroutines store their state, as it may be platform dependent (e.g. if you're targeting LLVM you may not have access to that representation).
2. They may not track their pointers, and so find it hard to move chunks of stack around if they have internal pointers into the stack.
3. They may want to enable some specific optimizations in the special (and rare, especially when IO is involved) case of a monomorphic calls to shallow coroutines. This requires that the representation of the subroutine's state be available as a high-level node for the compiler to change.
For a statically compiled language that produces machine code and prefers stack allocation and has strict constraints on copying references, it is not so easy.
If you just want to suspend some code, subroutine calls will do – but so will any other instruction, as long as you save the values of all registers. That's what the kernel does, after all. Using the subroutine call boundary lets you save a bit of work by saving only the callee-save registers, but it doesn't make that much difference. The problem is that you still need a stack, which prevents coroutines from being lightweight enough.
If you want to move the stack, like Go does... then just like with a GC, you need to track which registers and which stack offsets contain pointers at a given point in the program. That does require some sort of safepoint, which could be a subroutine call. But that's all irrelevant from Rust's perspective, because Rust's low-level pointer model is fundamentally incompatible with moving the stack. In Go, pointers into the stack are themselves only stored on the stack and in registers – never on the heap – so the runtime can easily find them all and adjust them. If you could store pointers into the stack onto the heap, it might still be possible to adjust them, albeit at higher cost, if you had a GC-managed heap where the runtime knows where all the pointers are. But Rust doesn't even have that. Tracking all the pointers in the heap is impossible for multiple reasons:
- Rust programs don't have to use any particular runtime-provided heap; they can ask the OS for memory directly and store pointers there.
- Pointers aren't necessarily stored in an easy-to-maintain way: e.g. Rust supports tagged unions, where the interpretation of one word in memory as either a pointer or integer depends on the value of another word.
Oh, and Rust lets you convert pointers to integers (useful for hashing purposes), so moving pointers behind a program's back would have visible side effects in any case.
So moving the stack is out. That leaves two possibilities:
- Segmented stacks;
- Calculating the required stack space statically.
LLVM can generate code with segmented stack support [1], and Rust used to enable that option. It's not that bad, but it does require every function to begin by checking whether there's enough space left for its stack, and call a runtime function (__morestack) if not. That adds a measurable overhead to function calls. It also requires using a custom thread-local storage slot to track the current stack base, which means you can only call the function on threads where the runtime set up that slot, not any random C thread. That's a problem if you're writing a library meant to be used by C code, which evolved into a major use case for Rust. For these reasons, Rust nixed segmented stack support sometime before 1.0. Re-enabling it is one hypothetical reason why you might want two versions of each function.
As for calculating the required stack space statically, that wouldn't require an alternate code generation mode. But as you were discussing in another subthread, the design of most compiler pipelines makes it very difficult to expose that information to the frontend. Even if that weren't a problem, merely making it theoretically possible to statically calculate stack usage requires imposing severe limits on the program. You can't make dynamic calls where you don't know the callee, and you can't use recursion at all (because then you might need infinite space). And if only a subset of code is async-compatible... then you end up with the equivalent of colored functions anyway. So even if LLVM's pipeline could be wrangled into doing what's necessary, it would only be a small improvement over the higher-level async/await implementation that Rust currently has. (As an alternative, you might be able to create a sort of hybrid mode where only dynamic and recursive calls use segmented stacks, but that would still require an alternate code generation mode.)
The nail in the coffin is that Rust also targets WebAssembly, where you can't do any of the nice low-level things that you can on real machines. This has also been a blocker for guaranteed tail call elimination. I hate it.
Async/await also needs the same logical stack, it just captures it frame by frame. Continuations can do the same. But as you say, Rust has some particular constraints that make this hard.
> then just like with a GC, you need to track which registers and which stack offsets contain pointers at a given point in the program
Well, you don't need to track all the pointers, just pointers into the stack.
> ... The nail in the coffin is that Rust also targets WebAssembly ...
But yeah, I mentioned some reasons for why it could be hard for some languages -- internal stack pointers and lack of control over the backend are two of them.
But I find it problematic that backend problems force feature design today. If a language wants to be popular 20 years from now it needs to have the patience not to bend to its backends' limitations, particularly as it has few competitors in its niche. So if a language's designers are absolutely certain that async/await is the right design for them for the next few decades, then that's great, but if they decide they want to design a language for the next few decades based on the limitations of WebAssembly in 2019, then that's something I like less. Sometimes there's no choice and a language/feature has to ship, but this is a feature that the world can and has lived without.
Internal pointers into the stack aren't really the reason Rust can't take this approach either, though, unfortunately. Async frames are "pinned" once they start running anyway so the ecosystem is designed around not needing to move them. Async callee frames are, in a sense, contained in their caller's frames- this makes async recursion tricky, but in exchange each static call graph uses a single allocation.
The real issue is basically naasking's #1- not that Rust loses its IR or that it's platform dependent, but that LLVM simply doesn't have the infrastructure to work with stack frames directly. A sufficiently motivated+funded team could add it but Rust doesn't have that luxury.
In fact, this is also the reason that C++20 coroutines individually heap-allocate their frames. No compiler had the infrastructure or was able to put in the work in time. There are also some thorny problems around phase ordering here- stack frame layout typically isn't determined until very very late (to enable optimizations) but program-accessible types in both C++ and Rust need to know their sizes very very early. Java can sidestep this by being slightly higher level.
The problem is that LLVM doesn't even let you see the stack layout until very late, and then only as debug info. You could probably hack something on top but it would probably wind up looking a lot like C++20, with a lot of extra allocation, or like Rust, with all the stack frame layout optimization duplicated into the frontend.
Unless you have some approach I haven't heard of, in which case, please provide a link!
That said, the extra object headers from heap-allocated activations will make up some of the difference.
It'll be interesting to see how this plays out on the JVM though.
This is a lesson that we all learned during the NPTL/NGPT era [1] and have subsequently forgotten and will have to relearn.
Anyway, I don't think anyone has forgotten those lessons, but they're not very relevant to modern managed runtimes. (also mioco seems to be single-threaded)
Async I/O isn't anything new, and it's stood the test of time. Nginx isn't some new experimental technology.
> Anyway, I don't think anyone has forgotten those lessons, but they're not very relevant to modern managed runtimes. (also mioco seems to be single-threaded)
They're still very relevant. Linux can spawn hundreds of thousands of kernel threads.
Anyway, we used to have an multithreaded M:N runtime in Rust, but it was removed because the performance was worse than 1:1 in practice. M:N was moved out of the standard library into the mioco library. Now it's unmaintained, because nobody actually wants M:N threads in Rust—an interesting outcome to be sure!
What do you mean? User-mode threads also use async I/O. They do just what async/await does, only without the syntax. I mean that async/await hasn't stood the test of time just yet.
> Linux can spawn hundreds of thousands of kernel threads.
Yeah, not very active ones, though. And their stacks are never really uncommitted. And the scheduler is badly optimized for many of the very uses you want in a server. On the other hand, you can spawn many millions of active fibers, schedule them how you like, and play with their state and its representation in many interesting ways (like moving a running fiber from one machine to another).
> Now it's unmaintained, because nobody actually wants M:N threads in Rust—an interesting outcome to be sure!
Just to make sure, when people say "M:N threads" they can mean a lot of very different things. In Java, they'll behave like async/await, only without the syntactic constructs. Stacks and stack frames are moved, can shrink and grow, can have various representations etc., and the scheduling is entirely pluggable. You can schedule the continuations manually, or in a scheduler of your choice, shared, or not, with ordinary tasks.
This is, from what I can tell, their main advantage in performance over other implementations. (In Rust, at least- the performance story may easily be different in garbage collected languages like Go or Java.)
Another caveat: Rust's claims to have implemented M:N in the past must be tempered a bit by two factors. First, Rust had to use segmented stacks, which Go moved away from and which Loom seems to be avoiding, precisely because they're so expensive. Second, Rust had to dispatch all IO APIs dynamically based on the currently selected runtime, which Go avoids by not having two runtimes, and which Loom can probably avoid via its scheduler implementation and/or JIT devirtualization.
They didn't decide to move off of M:N when it was exclusively M:N.
When you write "Java", you mean a specific implementation of the JVM, like the JVM reference implementation in OpenJDK, or you imply that all JVM implementations must support this to be called Java?
Edit: Is your comment related to Project Loom, aiming to add fibers and continuations to Java?
However, the paradox you have to manage here is this: if you make this sequential/parallel determination pretty high up in the code so that there is a lot of sequential code right after an async part (async can call sync all it wants), then very little of your code is altered, but now you have the green threads problem - one long CPU-bound task that doesn’t yield.
At least with Rust, if you’ve partitioned your data properly then there can be several threads of execution running tasks on unrelated data, but in Javascript you will be punished for trying to consolidate the async bits in this fashion.
Now `async`, you could see it as an implementation detail that you need in order to get much better performance...
val result = runBlocking {
val one = async { doSomethingUsefulOne() }
val two = async { doSomethingUsefulTwo() }
one.await() + two.await()
}
https://kotlinlang.org/docs/reference/coroutines/composing-s...EDIT: actually this was wishful thinking, since both functions would block the main thread. The functions need to be modified to "suspend" for this to be truly async.
My ideal language would be something like Kotlin suspend semantics, except ALL functions were implicitly "suspend" functions. If that's even possible...
let result = executor::block_on({
let one = async { do_something_useful_one() };
let two = async { do_something_useful_two() };
one.await + two.await
});
See examples in:- What is the purpose of async/await in Rust? — https://stackoverflow.com/a/52835926/155423
And why this might be a bad idea in:
- What is the best approach to encapsulate blocking I/O in future-rs? — https://stackoverflow.com/q/41932137/155423
And the syntax is surprisingly similar, especially considering we are supposedly comparing a low level language to a high level one.
See the sibling comment [1] where I actually tested and ran the code, finally. ;-). I was unaware that Kotlin would race two `async` blocks, so we had to use the `join` combinator here.
What you need is:
use futures::join;
let result = executor::block_on({
let one = async { do_something_useful_one() };
let two = async { do_something_useful_two() };
let (one, two) = join!(one, two);
one + two
});
I feel like this might be a common stumbling block and hope there will be a Clippy check for this or something.In the spirit of getting it right, your code is missing an `await` and `join` is back to a function in alpha18. Here's the currently working (this time tested!) code:
#![feature(async_await)]
use futures::{executor, future}; // 0.3.0-alpha.18
fn main() {
let result = executor::block_on(async {
let one = async { do_something_useful_one() };
let two = async { do_something_useful_two() };
let (one, two) = future::join(one, two).await;
one + two
});
println!("{}", result);
}
fn do_something_useful_one() -> i32 { 1 }
fn do_something_useful_two() -> i32 { 2 }The sequential but still async version in Kotlin is simply
val result = runBlocking {
doSomethingUsefulOne() + doSomethingUsefulTwo()
}
No awaits necessaryIn other words, synchronous procedures are essentially locally CPU-bound, while asynchronous procedures are bound by I/O, network or IO/network/CPU of remote servers.
Async functions are the most efficient way to support asynchronous procedures, while non-blocking non-async functions are the most efficient way to support synchronous procedures.
There are also adapters to turn synchronous function into asynchronous ones, and to turn asynchronous functions into blocking non-async functions, although they should only be used when unavoidable, especiallly the latter.
One of the problems with async is that it's viral. Make a function async and your entire call graph has to be made async, unless they handle the future manually without await. Async introduces a new "colour" to the language.
I absolutely agree that asynchronicity should be the provenance of the caller. If you look at typical code in languages with async/await such as JavaScript and C#, awaiting is the common case. So if async causes your code to be littered with awaits anyway, it makes more sense to make awaiting the default (just like it's the default for the rest of the language) and "deferring" the exception. Call it "defer" or something.
(As an aside, I was always disappointed that Go's "go" doesn't return a future, or that indeed there's no "future" support in the standard library. Instead you have to muck about with WaitGroup and ErrGroup and channels, which introduces sequencing even in places that don't need it. Sometimes you just want to spawn N goroutines and then collect all their results in whatever order, and short-circuit then if one of them fails. The inability to forcibly terminate goroutines is another wart here, requiring Context cancellation and careful context management to get right.)
In most languages, including Rust, async functions can call synchronous functions, so this isn't entirely true.
> unless they handle the future manually without await.
At some point, the futures are always handled manually, e.g., via the run-time, or in Rust case, by scheduling on an Executor, which is nothing but a fancy name for a library that manually handles a lot of futures.
> If you look at typical code in languages with async/await such as JavaScript and C#, awaiting is the common case.
Do you have any example of code that uses await more than synchronous functions ?
> So if async causes your code to be littered with awaits anyway, it makes more sense to make awaiting the default (just like it's the default for the rest of the language) and "deferring" the exception. Call it "defer" or something.
If the purpose is just to change the defaults, then you'd need `defer vec.len()` to access the length of a vector, or `defer array[i]` to index into an array, or `defer obj.foo` to access the field `foo` of an object, ... unless you want to make all of those operations `async`.
If you make await the default and add a `defer` operation that doesn't awaits, then it's hard to imagine how the resulting code could contain less `defer` annotations than `await` ones.
At that point you might choose one of the many non-zero-cost solutions to removing the async distinction, and call it a day like Go does.
First, async infects the call graph, that is, going up the stack. Not down.
Secondly, by "manually" I mean using the future API -- and_then() and so forth.
Thirdly, "defer vec.len()" doesn't make sense unless you want to execute it asynchronously and get a future. And that doesn't make sense unless you know it to be an expensive operation you want to run concurrently. I'm proposing that the hypothetical "defer" would act like Go's "go", except it would return a future.
Lastly, my point was that most code is serial and relies on awaits. For example, if you're doing a name lookup, then opening a socket, then writing to it -- three async calls that must be serialized with await because the actions depend on each other. Not awaiting is needed when you intend to use the future for some purpose, such as adding to a queue or putting in a map, or perhaps you're doing N connections that you put in a vector to ask the runtime to wait for all of them (does Rust's await syntax support awaiting multiple futures?). But most high-level "user" code doesn't juggle futures, it just awaits whenever it needs a result. From a language-ergonomics perspective, it makes more sense to optimize for the common case.
I've never written any async/await code in Rust, but I've written plenty in JavaScript, and it's the same story.
No, but you use a join combinator, which itself returns a future, that you can then await.
But that was not what I was arguing. Rather, I was arguing that await is more common than capturing the raw future. Case in point: The article, where all of the "real code" snippets rely purely on await.
What I proposed was that async functions would always be automatically awaited unless invoked with "defer", because most code is expressed as a series of invocations that use the return value of the function call.
There are times you want to store the raw future as some kind of state (for example, in a GraphQL server you might "resolve" the GraphQL tree as futures), or wait for multiple futures, or similar, but those are, in my experience, the exceptions, not the norm. Most code is:
let v1 = do1().await;
let v2 = do2(v1).await;
// etc.
In other words, serial.You can "handle the future manually without await" very easily, it's just one function call. The same sort of thing is available in most other languages with async/await- you can prevent async from going up the call stack by synchronously blocking on a future instead of awaiting it.
The caller is absolutely in charge of whether something runs asynchronously. The only difference is a syntactic default- and in fact some languages even make async functions calls default to synchronous execution. For example take a look at Kotlin's `suspend fun`s.
Not in Rust. Or C#. Or any language except JavaScript.
In Rust, a synchronous function can call and block on an async function with executor::block_on(). C# has Task.Wait() and Task.Result(). Python 3 has asyncio.run(). Scala has Await.ready() and Await.result(). Kotlin has runBlocking {}.
The choice of `async`-vs-not is not the one you describe. What we label with `async` is a particular compilation style that makes the function interruptible in userspace, in exchange for making synchronous and recursive execution a bit more complicated. This style is important because it gives you event loop-style concurrency without the performance costs of threads (kernel or userspace/green/etc).
However, this compilation style still leaves the choice you describe up to the caller. Most languages seem to reverse the default choice syntactically, but you can still synchronously block on a call to an async function or run it concurrently with something else. It just so happens to be useful primarily for doing the latter, so it gets lumped in with it. (And we get unfortunate misunderstandings like "What Color is Your Function?" that miss this point.)
As an example of a language that doesn't reverse the defaults, look at Kotlin's `suspend fun`s. While they still have `async`-like callee annotations to control the compilation style, a simple function call behaves the same regardless of the callee, and you instead use various `spawn` or CSP-like APIs to get concurrency.
If changing the keyword and making the defaults match threads/CSP isn't enough, perhaps viewing the annotation as part of an effect system would help? The thing being tracked here is "this function can suspend itself mid-execution," and a good effect system even functions be polymorphic over things like this. For example, an effect-polymorphic function passed a closure or interface can automatically take on the effects of its callees, simplifying the program somewhat if you're faced with viral async-ness.
https://gist.github.com/lattner/31ed37682ef1576b16bca1432ea9...
It’s not, however you do need the entire runtime to be in on it. That’s what Erlang, Go or gevent (for python) do: normal concurrency is in user land and normal IO uses « async ». The code itself looks / is synchronous, it just multiplexes in userland before doing so on OS threads.
I strongly recommend to check it out. I was absolutely mind-blown when saw this talk (without prior knowledge about Go), and that changed my life quite literally.
https://www.youtube.com/watch?v=QDDwwePbDtw
I did a few concurrency related projects (see "Visualizing concurrency in Go" https://divan.dev/posts/go_concurrency_visualize/, for example), and it's way more easier to reason about and work with than any other concurrency approaches I've seen so far.
Languages with fibers figure ou execution suspend points and physical core assignment for you and abstract it all away. So physically they're doing the same thing as async, just without all the cruft.
Rust decided not to go with fibers to avoid having a runtime. I still disagree with this because they already do reference counting, not every worthwhile abstraction is zero cost.
(I'd also be very skeptical of the 2% figure; where did you get that from?)
Remember, "zero cost" means "zero additional cost", everything has a cost.
Now that rust has async, couldn't they make fibers an opt-in replacement for threads? Your libraries use async, but you can handle those calls with fibers. Fiber support is compiled into your code but not the libraries
> Your libraries use async, but you can handle those calls with fibers. Fiber support is compiled into your code but not the libraries
We tried this! Making the attempt is actually one of the biggest reasons that we decided to remove green threads from Rust; it brought in tons of disadvantages and no real advantages.
Fibers are great because you can pretend they're threads. Everyone knows threads. Async is a legit PITA even if you're familiar with it. And the local state stored for a fake thread is often useful. When doing async I find myself frequently building hacky hashmaps to hold local variables values. websocket code is a good example of worse-case. 20k ongoing connections held by async (1 thread per core). It's a nightmare in everything except Vert.X Sync (fibers in Java) or maybe Golang and Erlang.
With fibers you get to keep all your local variables and "pretend" threads actually exist. It's a huge boon for productivity in the few languages where it's possible
Refcounting is a library feature, does not require a runtime, and has no impact on code not using it.
Fibers is a langage feature, does require a runtime, and impacts everything.
Rust actually stripped out its support for fibers as its community moved its usage downwards the stack. In much the same way it stripped out « core » support for GC/RC (the @ sigil) or internal (smalltalk/ruby-style) iteration.
If I text you, I expect an asynchronous response. It's ok if you get back to me 3 hours later (well, context-dependent).
If I call you, ask you a question, and then you wait 3 hours before giving me a response, I've been sitting there blocked for 3 hours, unable to get any more work done, because I asked for a synchronous response and got an asynchronous one. Even worse, if you try and give me the response over text, I won't see it because I have the phone pressed to my ear and now I'm blocked forever.
---
The caller and callee have to agree on whether the communication is synchronous or asynchronous. Asking for synchronous and getting asynchronous doesn't work. Asking for asynchronous and getting synchronous kind of works, but that's not truly synchronous, that's just asynchronous with zero delay before getting the response. Or even worse, asking for asynchronous (and giving you a completion handler to fire) and you fire it synchronously before returning control to me, that way lies madness. Don't do that.
CSP as implemented in languages like Go is just threads, but with an idiosyncratic implementation.
What about the stack size? I guess the default is about 4 MB, compared to a few kB in Erlang or Rust. This implies a virtual memory usage 100x or 1000x higher. I understand the unused virtual pages in each stack will never be mapped to physical memory, but isn't this putting a lot of pressure on the virtual memory system? I'd be curious to read a real world benchmark on this topic.
I think of await to mean "When this code is combined with all the other code that needs to be executed asynchronously on the event loop, this code will not be able to progress until this thing we are waiting on finishes. The event loop can therefore use this hint to schedule things and keep the processors busy."
I think I built up a good understanding of how this kind of a system works with futures 0.1 and tokio. It took some time but it all clicked together for me. But as for async/await as language features I'm satisfied to just let it be magic. I don't care to look under the covers at this point. I know it's similar to how futures worked, and I trust the rust team weaved it together well.
This doesn't work with Rusts implementation, because the functions have slightly different semantics and limitations.
E.g. an async function can't be recursive, since that would lead to a Future type of infinite size.
Async functions also don't necessarily run to completion like normal functions. They can instead also return at any .await point, which are implicit cancellation points. That means if someone just adds an async modifier to a function and some .await calls to methods inside it the result might not be correct anymore, since not all code in the function is guaranteed to be run anymore.
The cancellation mechanism was an explicit choice of the design - it would probably have been possible too to make async functions always run to completion and to thereby avoid the semantic difference.
I came in not wanting to see rs go down the same path. The idea of adding something that looks like it should be a simple value (`.await`) seemed odd to me, and I was already finding I liked raw Futures in rust better than async/await in JS already, especially because you can into a `Result` into a `Future`, which made async code already surprisingly similar to sync code.
I will say, the big thing in rust that has me liking async/await now is how it works with `?`. Rust has done a great job avoiding adding so much sugar it becomes sickly-sweet, but I've felt `?` is one of those things that has such clean semantics and simplifies an extremely common case that it's practically a necessity in my code today. `.await?` is downright beautiful.
Generators got us close to that, and it is a more generic feature. But over all the code I've written, I've only had one solid use case needing generators for something other than what's covered by async/await. Having a specific syntax that covers 95% of usages is very worth it in my opinion.
They could've shipped the `Promise` object with a `coroutine` fn, like bluebird did, and had the same semantics without needing to reserve words and add syntax to the language. All it opens up is combining async and generators, which can be handy for iterating things like prefetched results, but that's a rare enough case not to need to be baked into the language.
Separating the coroutine would let us work beyond just promises too. Wrap around old code using error-first callbacks, or new code using observables (yield to get the next, yield* to complete). I just wrote a coroutine to give quasi-concurrency to Google apps scripts, so it can queue up the routines running `fetch` methods and instead do them in one parallel-request-making `fetchAll`. It's a much better approach for the long-term of a language. Imho 90% of the junk we deal with in old languages is stuff for a convenient implementation at the time that we don't use anymore.
That's like saying "Ordering from Amazon does require knowledge of driving a van. That's how the stuff comes to your home!".
It skips the generator part and extra wrapper. That's a simpler syntax (and thus more clarity) for what people do 99% of the time.
Clarity is not necessarily "I can see the underlying mechanism".
Generators hide their underlying implementation (in C++) too, after all.
that has been contentious to say the least.
FWIW Python originally used generators to implement async / coroutines, and then went back to add a dedicated syntax (PEP 492), because:
* the confusion of generators and async functions was unhelpful from both technical and documentation standpoints, it's great from a theoretical standpoint (that is async being sugar for generators can be OK, though care must be taken to not cross-pollute the protocols) but it sucks from a practical one
* async yield points make sense in places where generator yield points don't (or are inconvenient) e.g. asynchronous iterators (async for) or asynchronous context managers (async with). Hand-rolling it may not be realistic e.g.
for a in b:
a = await a
is not actually correct, because the async operation is likely the one which defines the end of iteration, so you'd actually have to replace a straightforward iteration with a complete desugaring of the for loop. Plus it means the initialisation of the iterator becomes very weird as you have an iterator yielding an iterator (the outer iterator representing the async operation and the inner representing the actual iterator): # this yield from drives the coroutine not the iteration
it = yield from iter(b)
while True:
try:
# see above
a = yield from next(it)
except StopIteration: # should handle both synchronous and asynchronous EOI signals, maybe?
break
else:
# process element
versus async for a in b:
# process element
TBF "blue functions" async remain way more convenient, but they're not always an easy option[0], or one without side-effects (depending on the language core, see gevent for python which does exactly that but does so by monkey patching built-in IO).[0] assuming async ~ coroutines / userland threads / "green" threads
Except in the latter you have to bring your own `co(...)` function. Baking it into the syntax means you don't need that as a library.
Nope, the code is executed in order. That's what `await` accomplishes. Wherever you use `await` you have a guarantee that anything that comes after it will only execute when the promise that was awaited, has resolved.
I'd rather pick a standard solution over a library. I never used CO, because I didn't see the benefits of pulling this library in over plain ol' promises. The cognitive overload of generators, yielding, thunks and changing api-s is just too much IMHO, I like simpler things that work just as well.
With the await keyword baked into the language I can simply think of "if anything returns a promise, I can wait for the response in a simple assignment".
This. Rust can be complex enough for engineers new to the language. This sort of feature really requires a standard approach.
This is why we have await in Python and C# as well, eventhough we know how to write async/await with coroutines for 10 years now :] Tooling and debugging is super important.
I'd recommend just opening up an IDE in any of these and seeing the types and parameters that come in and out of these calls. The sync vs async worlds really to clash with each other but at least we have strongly typed languages to help with writing the code. In javascript you could be mixing callbacks and promises and encounter bits where you'd never know that the next line would never be executed because of some nonexistent callback, etc etc.
Even with Futures there can be times when you need the fine-grained control over execution that you don't get with async-await. Eg if you want to implement a `struct TimeOut<F: Future> { inner: F, timeout: std::time::Instant }` from scratch, you can't do that with async-await.
The trickiest thing for me to grasp first was the `Pin` type. Like how does this prevent the value to be moved... Until I realized it's `Pin<&mut T>` and then I had my a-ha moment.
- In general you have to be careful that if you're returning Poll::Pending, it's after something has registered a wakeup, otherwise your future is going to just stop.
- As a specific case, your match has to be inside a loop so that a state transition results in the new state being polled immediately. You can't just return and expect the caller to poll() you since there will be no wakeup registered that triggers the caller to do so.
- Because futures receive `&mut Self` rather than `Self`, it's difficult to do a state transition where the new state needs to own some properties of the current state. Eg `State::Foo { id: String } -> State::Bar { id: String }` Instead you have to do Option::take / std::mem::replace shenanigans on the state fields, or have an ugly State::Invalid that you can std::mem::replace the whole state with.
Thus, when you have a CPU-bound function, they are of no assistance in UI responsiveness, unless you deliberately insert yield points; in JavaScript, for example, you might do a unit of work, or process work until, say, ten milliseconds have elapsed, and then call setTimeout(keepGoing, 0). (Critically in JavaScript, you mustn’t use `Promise.resolve().then(keepGoing)`, because that’ll end up in the microtask queue which must be flushed empty before the browser will render, so you won’t actually yield to the UI with this approach.)
So be careful about your DisplayResultsOfSlowFunction(). If its slowness is that it takes a long time to calculate, and it doesn’t actually have any asynchronous stuff in it, async will not help you.
Does this mean all 3 of those crates are required to achieve async I/O?
tokio is a large project. Which part of tokio specifically is supposed to be used in conjunction with futures + mio to achieve asynchronous I/O?
Edit: it seems futures + mio are dependencies of tokio.
(Tokio is built on top of mio so if you’re using Tokio, you’re using mio)
Some time later I had to dip my toes into the world of Javascript, learning first about Promises, and finally reading about async/await. I just realized that it is basically the same trick I had been using all along. And now it's coming to Rust, neat!
[0]: https://web.archive.org/web/20190217162607/http://www.drdobb...
Or do you mean state machines?
(note that the case statements are hidden in all of the GET_* macro calls).
You create a main infinite loop in your application, in which you call all functions that are meant to perform activities that should be happening "at the same time", like for example taking ADC measurements, while updating a PID control algorithm, while waiting for the GPRS chip to connect to the internet; those would be their own function, called "protothread" if you use the Adam Dunkel's library, and they are going to run a number of blocking operations, waiting for some external hardware to do their thing.
Each one of those functions is converted into a Duff device by means of the macros provided by the Protothreads library.
Now each protothread function can express its own operations in a linear and very clear way, and yield execution to the other functions by just returning from the Duff device. When the main loop starts over, each of those functions will "resume" their operation at the point where they yielded during the last iteration.
The way I see it, this trick is much more than just loop unrolling; it provides a C developer with the same basic features that an async function would have in JS; and the yield points in the functions would be equivalent to the 'await' calls. This allows for very easy to read code, being able to follow the logic of each individual "thread" without much of the noise needed for taking care of the "coroutines" part. The next step (in the embedded world, at least) would be going the route of an RTOS, such as FreeRTOS, but I've managed well so far without it.
EDIT: and yes, this is basically also a neat was of streamlining a state machine
Fantastic work! Look forward to the async/await simplicity
The current thinking from the team seems to be that, if this does become an issue, a) inlining before state machine transformation can flatten this out quite a bit and b) convincing LLVM to do what would be required is a lot of work.
It goes back frame-by-frame. Yielding (ie returning `Poll::Pending`) is essentially returning "I'm not ready yet" and it's possible the parent future wants to do something else in that case.
The simplest case is a future that wraps a collection of futures and resolves to the value of the first one that resolves. In this case it would cycle through each child future and only return `Pending` if all of them return `Pending`.
Another case is that of a timeout future that wraps another future - it would resolve to `Err(TimedOut)` if the wrapped future returns `Pending` for too long.
>And what happens when it resumes execution?
The top-most future of the task is all the executor knows about, so the executor has no choice but to start from there.
But also the performance is not likely something to worry about. I've never measured it myself but I do know aturon posted that he could not find any measurable effect of having deep future stacks. It is noticeable in the sense that a big enough stack may exceed your OS's limit (especially on Windows where the main thread defaults to a 1MiB stack).
This is of course ignoring the related need to also write code parametrized over lifetime variable structure ..