Why async fn in traits are hard
smallcultfollowing.com
smallcultfollowing.com
"Confessions Of A Used Programming Language Salesman, Getting The Masses Hooked On Haskell"
https://www.researchgate.net/publication/237445028_Confessio...
If it took wonders to make Haskell concepts mainstream in languages with a GC, I can only describe what Rust achieves as true miracles.
In any case, Rust has already managed other language designers to take a look into adopting such type systems ideas, in itself that is a big victory from Rust community.
https://en.wikipedia.org/wiki/Erik_Meijer_(computer_scientis...
In a situation like this:
trait Foo { async fn foo(&self) -> Bar; }
trait Baz: Foo { async fn baz(&self) -> Meow { ...await self.foo() into a Meow ... }
you'll end up with boxes of boxes of boxes. It kinds of makes the feature something for the outermost abstraction layer or you end up with boxes of boxes of boxes.> And of course it’s worth highlighting that most languages box all their futures, all of the time. =)
It's also worth highlighting that most of those languages (1) have a GC that makes boxing a much cheaper operation than in Rust, and (2) those languages aren't advertised as "low-level" languages and that justifies implicit boxing as a trade-off.
---
I feel that async shouldn't really be special but rather just build on other features that stand on its own. GATs and impl Trait in Traits seem reasonable. But being able to do dynamic dispatch on trait methods that return a type that isn't `Sized` seems like a tough problem to solve, and the blog post didn't managed to convince me that implicit boxing is the right approach here. AFAICT, we need either some kind of boxing, or better support for unsized rvalues or similar. I think I would be more comfortable with "sugar" for boxing if the caller of the trait method were in control of where exactly the result is allocated (not necessarily a Box), but that calls for some kind of placement syntax.
Why is that? I would intuitively think it is the other way. (Is a malloc/free pair not cheaper than an allocation on the GC heap + collecting its garbages?)
It's always about details though. If a GC is faster than malloc/free, but your language doesn't tend to allocate much to begin with, the whole system can be faster even if malloc is slower. It always depends.
A GC that's integrated with a programming language can do much much better (different heaps for short and long lived allocations, for example).
RustLatam 2019 - Without Boats: Zero-Cost Async IO
There's also https://www.reddit.com/r/rust/comments/6zy8hl/kotlins_corout...
One tough part about digging through history here is that designs change over time...
Rust went the way it did syntactically because annotating the await points lets you see them when reading a program, and possibly out of familiarity.
Rust went the way it did implementation-wise (stackless rather than stackful) because it's more efficient.
Care to explain? Why would using semantic versioning, something near universally accepted as a best practice, be flawed in your eyes?
> Isn't it well known to be a flawed practice yet?
If you agree with that statement, why can't you explain why it is flawed?
The question was why do I think it's both flawed and universally accepted.
Because it is, and it isn't?
These are basic facts.
Have you read the semver spec?
Is your libraries new version compatible with my library?
How can you know?
Yes, it was. While the grammar I used should be pretty clear on what I meant already, here’s an alternative version:
Why would using semantic versioning (something near universally accepted as a best practice) be flawed in your eyes?
Or more simply:
Why would using semantic versioning be flawed in your eyes?
> These are basic facts.
Actually, they aren’t. It’s your opinion (not a fact) and you still haven’t provided a single drawback (flaw) to following semantic versioning despite repeated requests.
> Have you read the semver spec?
Yes, I have. More than once in fact. Have you?
> Is your libraries new version compatible with my library?
Assuming both authors follow semantic versioning faithfully and the previous minor/patch release worked, then yes, it very likely will be compatible.
> How can you know?
If both sides have a well defined API and the authors follow semantic versioning, you should have a good indication that way. That said, as the saying goes, “trust but verify”.
Those are not facts, those are claims. Facts are supported with data, and you haven't given any.
The only fact you have actually provided is that you can't tell facts from claims apart, and that makes it very unlikely that you will be able to provide any actual facts that support your claims.
Rust has tooling to automatically enforce semver. For example, this tool: https://github.com/rust-dev-tools/rust-semverver
Once you modify a Rust library, it downloads the last released version, compiles it and extracts its AST, and compares it with the AST of the current version.
The diff of the two ASTs tells you what the changes are, and there is a Rust book that documents which changes are semver breaking and which aren't. So if you only add a new function to your crate, the tool says that you can bump the semver patch version or minor version, but if you change the name of a public API, the tool requires you to make a new major semver release.
Setting this tool in CI is dead easy, and will make your CI fail if a PR changes the API of your crate without properly updating its semver version.
The question in my mind is where does semver come into this?
Why do you need three arbitrary and meaningless numbers when a commit hash, branch, or tag would suffice.
Nobody said it fixed all problems. Virtually all dependency management systems struggle with this problem, whether packages use semantic versioning or not.
I absolutely hated, not because OOP is bad, but because they never really manage to blend it just right with static typing. But I was the odd one. Everybody else seemed to love it and use it for everything.
I was lucky enough to outlive this OOP craziness.
Now I am seeing the same craziness with “a sync” and continuations and wonder if I will be lucky enough to outlive it to.
fn get_user(&self) -> Pin<Box<dyn Future<Output = User> + Send + '_>>;
This is plain craziness.async language features are just inline co-routines. Inline co-routines are a useful concept. It's a single tool in a big toolbox of other tools and not every problem requires the same tool, just like OOP features. It's just more stuff you can use if you want. If you don't want it, you don't have to use it, you can just write Haskell or OCaml without using objects and not worry about it and you won't have to worry about how long you live.
And C++ has always been an “OOP” language.
But they are both a dumped down version of OOP as they both don’t have metaclass support or extreme late binding.
For something more aligned to OOP look at smalltalk or Ruby
def getUser() : Future[User]Usually it would look like this in Rust:
fn get_user(&self) -> impl Future<Output = User>;
This isn't exactly the same functionally, but for most use cases it works just fine. The difference is that this uses static dispatch instead of dynamic dispatch and therefore can't support dynamically returning two different structs that both implement Future, it has to be one type.
If you go lower in the abstraction level you get more complicated mechanics, and that says nothing over the validity of the high level semantics.
I don't think it makes sense to compare the output of Rust macros to assembly code generated by a C compiler or to vtables.
Macros are part of the Rust language, as is their output. Understanding Rust means understanding both input and output. So this abstraction boundary is intentionally leaky to some degree.
Generated assembly and vtables on the other hand are compiler implementation details subject to change without notice. Any abstraction leakage is unintended and undesirable (even if developers sometimes benefit from understanding that output)
Your line of reasoning makes all debates about language complexity completely pointless.
For example, no need to understand the magic behind the 'quote!' or the #[async_trait] macros to use them.
As a consequence, criticising the complexity of whatever a C compiler generates cannot possibly be valid criticism of C's complexity on a _semantic_ level whereas criticising the output of a Rust macro can be valid criticism of Rust on a semantic level.
Most C programmers won't have to write or understand inline assembly often, if ever. Of course you can encounter it in production problem or something, so you could make an argument that all C programmers need to understand "C with inline assembly", which you are making for Rust macros.
As long as you just use Rust macro's and not write your own you are solidly in "C without inline assembly" territory.
I couldn't disagree more. Macros are not similar to inline assembly at all precisely because they do _not_ step outside the bounds of the language.
Whatever similarities you may find, it's simply not helpful to deny the fundamental distinction between language A generating code in language A and language A invoking/generating code in language B.
It's futile to debate the properties of a particular language if you can't make a distinction between that language and anything it can generate or embed in some opaque way.
Take Objective-C automatic reference counting [1], implemented as a transformation of the original code to valid code of the same language (similar to a Lisp/Rust/Scala style macro) by automatically adding the appropriate statements.
If I understand your argument correctly, according to you this increases the complexity of "Objective-C with ARC", but would not have done so if the compiler would have implemented it as a direct transformation to its compilation target instead.
To me, that is an implementation detail which does not matter. "Objective-C with ARC" is exactly as complex in both cases. I'd argue it's even a little bit less complex with the "macro" implementation since you don't need to know assembly to know what ARC is doing.
Similar to ARC, Rust implements some things with macro's, which first "compiles" something to valid Rust. To me this is not more difficult for users than it would be if the compiler would directly generate LLVM IR without this intermediate step.
The inclusion of macro's in a language do make the language more complex of course! And creating macros is notoriously difficult since you're basically implementing a small compiler step! But for the user using something is not suddenly more difficult because it's implemented using a macro.
[1] https://en.wikipedia.org/wiki/Automatic_Reference_Counting
It's important because any and all code in language A is fair game when it comes to criticising semantic properties of language A. Code in other languages isn't.
noncoml criticised Rust based on a piece of Rust code. My point is simply that this criticism is potentially legitimate in ways that criticising C based on a piece of assembly code could never be.
I think our disagreement arises because you are asking a completely different question. What you're saying is that for devs who invoke some code it may not matter one bit whether that code was implemented in language A or language B or language A generated by language A or B. Those distinctions do not necessarily affect the semantic complexity for users of that code.
I completely agree with that. I also agree that the code snippet noncoml posted does not mean using async code in Rust has to be overly complicated.
But when I see a piece of Rust source code, I can criticise Rust based on it regardless of where that code came from or what purpose it serves.
Someone had to think in terms of Rust in order to write that code, and it's always worth asking whether it shouldn't be possible to express the same thing in a simpler way or whether that would have been possible in another language.
The fact that this code does not have to be understood by its users is completely irrelevant for this particular question.
Dealing with this "ridiculous" type step by step:
Future<Output = User>
This is a future. It returns an user. dyn
However, because the exact `Future` implementation can vary, dynamic dispatch is necessary. Use of `dyn` keyword explicitly says that a vtable is necessary. + Send
Rust is supposed to make programs free from data races. It's very much possible to have objects which aren't allowed to change the thread they are in. For instance, `Rc<T>` (single-thread reference-counting pointer) cannot be sent to other threads due to it using non-atomic reference count.In this case, using `+ Send` says that whatever future this method returns must be possible to send to other threads. This prevents use of objects like `Rc` within the futures, and allows the future to be used in multi-threaded executors.
+ '_
The method borrows `self` (as seen by the use of `&self`). Saying `+ '_` means that the future should borrow `self`. This is necessary due to lifetimes. Box<...>
Because dynamic dispatch types (note `dyn` before) can have varying sizes (one future can be smaller, another one can be bigger), it's necessary to wrap them somehow so that they could be represented in memory. In this case, the simplest heap-allocation type (`unique_ptr` in C++) is being used. There are other options, like `Arc<...>`, but in this case, `Box` is sufficient. Pin<...>
Futures can store pointers to its local fields. This is going to cause issues if the allocation could be moved somewhere else in the RAM - the future would be moved to the new location, but the pointers would still refer to RAM in the old position. Using `Pin` prevents the user from moving the contents of `Box` to another place in the RAM.C++ is most famous for this philosophy, but Rust doubles down on it because of its commitment to memory-safety soundness through the type system. Look at how many details having to do not with any algorithm but with the code the compiler emits are stated here: that the call is made in an asynchronous context, that an object requires a vtable, that an object is able to change threads, the object's lifetime, boxing, and that the object cannot be moved in RAM. The only part that's related to the algorithm here is `User`.
Of course, in constrained environments like embedded systems or other circumstances where control over resources is crucial -- the domains Rust is designed to target -- accidental complexity becomes essential as control over resources is an important component of the problem.
Some more here, and an alternative (with it's own tradeoffs) that could be called "zero-cost use" (although that's equally as bad and confusing as "zero-cost abstractions"): https://news.ycombinator.com/item?id=19932753
However, worth noting a lot of this accidental complexity is visible due to dynamic dispatch. If `get_user` wouldn't have to be dynamically dispatched, it would be simply written as:
async fn get_user(&self) -> User { ... }
Dynamic dispatch complicates a lot of things in Rust, as then the compiler cannot simply check the returned type to determine the necessary guarantees - as the type can be anything.But there's definitely a very big tradeoff -- in both approaches -- between fine-grained control over resources and "non-algorithmic complexity" if you'd like to call it that.
To get an intuitive feel for why that is so, try to imagine that the OS were able to give you threads with virtually no overhead, and you'll see how anything that's expressible with async would be expressible without it. Over the years we've so internalized the fact that threads are expensive that we forgot it's an implementation detail that could actually be fixed.
Computations in Rust can suspend themselves without declaring themselves to be async: that's exactly what they do when they perform blocking IO. They only need to declare themselves async when they want a particular treatment from the Rust compiler and not use the continuation service offered by the runtime, which in Rust's case is always the OS.
In fact, Rust already gives you two ways to suspend execution -- one requires a declaration, and the other does not, even though the two differ only in the choice of implementation. That you wish to use the language's suspension mechanism rather than the OS's is not a different "model." It's the same model with different implementations.
Whether you need to declare that a subroutine blocks or not has no bearing on the fairness of scheduling.
That "supposed to" is an aesthetic/ideological/pragmatic/whatever preference, and one with significant tradeoffs. Again, Scheme's shift/reset gives you the same control over scheduling as Rust's async/await, but you don't need to declare that in the type signature. Rust puts it in the type signature because in the domains Rust targets it is important that you know precisely what kind of code the compiler generates.
> Either way algorithmically the choice has to be explicit, even it it's not a type declaration, it still has to be an explicit declaration somewhere or there is no choice and algorithms have to be expressed in terms of a single concurrency model.
First, we're not talking about two models, but one model with two implementations (imagine that the OS could give you control over scheduling, like here [1]; you'll have the exact same control over the "concurrency model" but without the type declaration -- even in Rust -- although you'll have less precise control over memory footprint). Second, that the language's design dictates that you must tell the compiler in the type signature what it is that your computation does is precisely what creates accidental complexity, as it impacts your code's consumers as well.
In any event, the reason Rust does it is not because of the aesthetic preference you express (that would be true for, say, Haskell) but because the domains it targets require very fine-grained control over resource utilization, which, in turs, require very careful guarantees about the exact code the compiler emits. It's a feature of Rust following C++'s "zero-cost abstractions" philosophy, which is suitable for the domains Rust and C++ target -- not something essential to concurrency. Again, Scheme gives you the same model, and the same control over concurrency, without the type signatures, at the cost of less control over memory footprint, allocation and deallocation.
I think you might be saying that you happen to personally prefer the choices Rust makes (which it makes because of its unique requirements), and that's a perfectly valid preference but it's not universal, let alone essential.
As long as you have shared memory adding type signatures to avoid thinking about and handling concurrent memory access only reduces accidental complexity and limits possibilities to make mistakes.
That is certainly one way to reduce algorithmic mistakes, but it doesn't reduce accidental complexity, and it's not why Rust requires async in the signature. Rust needs to prevent data races even in non-async subroutines that run concurrently. Rust requires async because it implements that feature by compiling an async subroutine rather differently from a non-async one, and the commitment to so-called "zero-cost abstractions" requires the consumers to know how the subroutine is compiled.
This is something I wish more people would realize. We assume that 1:1 threads inherently can't scale, mostly because of folklore. (As far as I can tell, the origin of this claim is back-of-the-envelope calculations involving stack sizes, leading to address space exhaustion on 32-bit—obviously something that hasn't been relevant for servers for a decade.) That's why we build async/await and elaborate userspace scheduling runtimes like the one Go has (and the one Rust used to have pre-release).
I think it's worth questioning the fundamental assumptions that are causing us to do this. Can we identify why exactly 1:1 threading is uncompetitive with async/await schemes, and fix that?
All this, by the way, is not to say that Rust made the wrong decision in focusing on async/await. Rust generally follows the philosophy of "if the computing environment is difficult, it's Rust that has to adapt".
I think we can identify it, but fixing it is not easy, at least at the kernel level.
The easy part is the cost of scheduling. The kernel must schedule threads with very different behaviors -- say, encoding a video or serving requests over a socket -- so it uses a scheduling algorithm that's a compromise for all uses. But we can let the language's runtime take the scheduling of kernel threads with something like this: https://youtu.be/KXuZi9aeGTw
The harder part is managing the memory required for the stack. Even on 64-bit systems, the OS can't shrink and grow stacks finely enough. For example, I don't think there's a hard requirement that, say, memory below the sp is never accessed, so the OS can't even be sure about how much of the stack is used and can't uncommit pages. But even if there were such a requirement, or, that the language could tell the OS it follows such a requirement, still the OS can only manage memory at a page granularity, which is too much for lightweight concurrency. Any finer than that requires knowing about all pointers into the stack.
We can do it in languages that track the location of all pointers in all frames, though, which is what we're attempting to do in Java, and this allows us to move stacks around, grow them and shrink them as necessary, even at a word granularity.
> All this, by the way, is not to say that Rust made the wrong decision in focusing on async/await. Rust generally follows the philosophy of "if the computing environment is difficult, it's Rust that has to adapt".
... and it follows C++'s (horribly named) "zero-cost abstractions" philosophy which can reasonably be said to be a requirement of Rust's target domains. So I certainly don't think async/await is wrong for C++/Rust/Zig, but I think it's wrong for, say, JavaScript and C#, and we're going a different way in Java (more like Scheme's).
Another possible contributing factor is that in one respect Rust is a higher-level language than Java or JS: it compiles to a VM -- LLVM/WASM -- over which it doesn't have full control, so it doesn't have complete control over its backend. That, BTW, is why Kotlin adopted something similar to async/await.
I don't think that's been conclusively shown. 4kB is smaller than a lot of single stack frames. It would be interesting for someone to measure how large e.g. Go stacks are in practice—not during microbenchmarks.
> Another possible contributing factor is that perhaps ironically, in one respect Rust is a higher-level language than Java or JS, as it compiles to a VM -- LLVM/WASM -- over which it doesn't have full control, so it doesn't have complete control over its backend.
It's not really a question of "control"; we can and do land changes upstream in LLVM (though, admittedly, they sometimes get stuck in review black holes, like the noalias stuff did). The issue is more that LLVM is very large, monolithic, and hard to change. Upstream global ISel is years overdue, for example. That's one of the reasons we have Cranelift: it is much smaller and more flexible.
But in any case, the code generator isn't the main issue here. If GC metadata were a major priority, it could be done in Cranelift or with Azul's LLVM GC support. The bigger issue is that being able to relocate pointers into the stack may not even be possible. Certainly it seems incompatible with unsafe code, and even without unsafe code it may not be feasible due to pointer-to-integer casts and so forth. Never say never, but relocatable stacks in Rust seems very hard.
The upshot is that making threading in Rust competitive in performance to async I/O for heavy workloads would have involved a tremendous amount of work in the Linux kernel, LLVM, and language design—all for an uncertain payoff. It might have turned out that even after all that work, async/await was still faster. After all, even if you make stack growth fast, it's hard to compete with a system that has no stack growth at all! Ultimately, the choice was clear.
That would be interesting to study. We'll do it for Java pretty soon, I guess.
> After all, even if you make stack growth fast, it's hard to compete with a system that has no stack growth at all!
If the same Rust workload were to run on such a system there wouldn't be stack growth, either. The stack would stay at whatever size Rust now uses to store the async fn's state.
> a tremendous amount of work in the Linux kernel
I don't think any work in the kernel would have been required.
> Ultimately, the choice was clear.
I agree the choice is pretty clear for the "zero-cost abstractions" approach, and that Rust should follow it given its target domains. But for domains where the "zero-cost use" approach makes more sense, the increase in accidental complexity is probably not worth some gain in worst-case latency. But that's always the tradeoff the two approaches make.
There is no inherent reason why it would be any slower or faster. It also does not imply any particular scheduling mechanism.
How can it be wrong for JavaScript? Because JavaScript is a single threaded environment, it was the only possible choice and it's one of the most influential change in the language (in itself it changed more the language use than the whole ES6 bundle). It's a really great success and I'm not sure why you'd like to revert it.
I don't. The old way and async/await aren't the only two options.
> it was the only possible choice
Why?
Remember, the single threaded character of the js VM is not an implementation detail, it's part of the spec. Hate it or love it but the web works this way.
Also, if you think threads are a good concurrency abstraction, let's play a little game: consider you need to read 1M files on a spinning disk. How many threads do you need to run on to get the maximum read performance:
a) one per file
b) just one
c) one per core
d) a magic number which depends on your hard disk's firmware and your workload.
Convenient and intuitive right?
Threads are a dated concurrency primitive which would have died long ago if wasn't also a good parallelism primitive.
You could only have the first kind, but it's the same situation with async/await. Only async/await tracks the "good" kind of blocking, yet lets the "bad" kind go untracked.
> How many threads do you need to run on yo get the maximum read performance
The exact same number as you would for doing it with `await Promise.all` The same knowledge you have about the scheduling mechanism doesn't go away if you're no longer required to annotate functions with `async`.
> Threads are a dated concurrency primitive which would have died long ago if wasn't also a good parallelism primitive.
Maybe they are, but async/await are the exact same construct only that you have to annotate every blocking function with "async" and every blocking call with "await". If you had a language with threads but no async/await that had that requirement you would not have been able to tell the difference between it and one that has async/await.
The “bad kind” is indistinguishable from CPU intensive computation anyway (which cannot be tracked), but at least you have a guarantee when you are using the good kind. (Unfortunately, in JavaScript, promises are run as soon as you spawn them, so they can still contain a CPU heavy task that will block your event loop, Rust made the right call by not doing anything until the future is polled).
> The exact same number as you would for doing it with `await Promise.all`
From a user's perspective, when I'm using promises, I have no idea how it's run behind (at it can be nonblocking all the way down if you are using a kernel that supports nonblocking file IOs). This example was specifically about OS threads though, not about green ones (but it will still be less expensive to spawn 1M futures than 1M stackful coroutines).
> Maybe they are, but async/await are the exact same construct only that you have to annotate every blocking function with "async" and every blocking call with "await". If you had a language with threads but no async/await that had that requirement you would not have been able to tell the difference between it and one that has async/awaitof
I don't really understand your point. Async/await is syntax sugar on top of futures/promises, which itself is a concurrency tool on top of nonblocking syscalls. Of course you could add the same sugar on top of OS threads (this is even a classic exercise for people learning how the Future system works in Rust), that wouldn't make much sense to use such thing in practice though.
The question is whether the (green) threading model is a better abstraction on top of nonblocking syscalls than async/await is. For JavaScript the answer is obviously no, because all you have behind is a single threaded VM, so you lose the only winning point of green threading: the ability to use the same paradigm for concurrency and parallelism. In all other regards (performance, complexity from the user's perspective, from an implementation perspective, etc.) async/await is just a better option.
Of course it can be tracked. It's all a matter of choice, and things you've grown used to vs. not.
> but at least you have a guarantee when you are using the good kind
Guarantee of what? If you're talking about a guarantee that the event loop's kernel thread is never blocked, then there's another way of guaranteeing that: simply making sure that all IO calls use your concurrency mechanism. As no annotations are needed, it's a backward-compatible change. That's what we're trying to do in Java.
> but it will still be less expensive to spawn 1M futures than 1M stackful coroutines.
It would be exactly as expensive. The JS runtime could produce the exact same code as it does for async/await now without requiring async/await annotations.
> Async/await is syntax sugar on top of futures/promises, which itself is a concurrency tool on top of nonblocking syscalls.
You can say the exact same thing about threads (if you don't couple them with a particular implementation by the kernel), or, more precisely, delimited continuations, which are threads minus the scheduler. You've just grown accustomed to thinking about a particular implementation of threads.
> The question is whether the (green) threading model is a better abstraction on top of nonblocking syscalls than async/await is
That's not the question because both are the same abstraction: subroutines that block waiting for something, and then are resumed when that task completes. The question is whether you should make marking blocking methods and calls mandatory.
> In all other regards (performance, complexity from the user's perspective, from an implementation perspective, etc.) async/await is just a better option.
The only thing async/await does is force you to annotate blocking methods and calls. For better or worse, it has no other impact. A clear virtue of the approach is that it's the easiest for the language implementors to do, because if you have those annotations, you can do the entire implementation in the frontend; if you don't want the annotation, the implementors need to work harder.
Of course, you could argue that you personally like the annotation requirement and that you think forcing the programmer to annotate methods and calls that do something that is really indistinguishable from other things is somehow less "complex" than not, but I would argue the opposite.
I have been programming for about thirty years now, and have written extensively about the mathematical semantics of computer programs (https://pron.github.io/). I understand why a language like Haskell needs an IO type (although there are alternatives there as well), because that's one way to introduce nondeterminism to an otherwise deterministic model (I discuss that issue, as well as an alternative -- linear types -- plus async/await and continuations here: https://youtu.be/9vupFNsND6o). And yet, no one can give me an explanation as to why one subroutine that reads from a socket does not require an `async` while another one does even though they both have the exact same semantics (and the programming model is nondeterministic anyway). The only explanation invariably boils down to a certain implementation detail.
That is why I find the claim that even when two subroutines have the same program semantics, and yet the fact that they differ in an underlying implementation detail means that they should have a different syntactic representation, is somehow less complex than having a single syntactic representation to be very tenuous. Surfacing implementation details to the syntax level is the very opposite of abstraction and the very essence of accidental complexity.
Now, I don't know JS well, and there could be some backward compatibility arguments (e.g. having to do with promises maybe), but that's a very different claim from "it's less complex", which I can see no justification for.
From what I understand now, you are arguing about syntax: we should not need to write “async” or “await”. I'm not really going to discuss this, because as you said, I do like the extra verbosity and I actually like explicit typing for the same reason (Rust is my favorite, with just the right level of inference) and I'm not fond of dynamic typing or full type inference. This is a personal taste and that isn't worth arguing about.
On the other hand, there is also a semantic issue, and sorry I have to disagree, stack-ful and stack-less coroutines don't have the same semantic, they don't have the same performance characteristic nor they do have the same expressiveness (and associated complexity, for users and implementers). What I was arguing that if you want the full power of threads, you pay the price for it.
But from what I now understand, you just want a stackless coroutine system without the extra async/await” keywords, is that what you mean?
As to semantic differences, what is the difference between `await asyncFoo()` and `syncFoo()`?
BTW, I also like extra verbosity and type checking, so in the language I'm designing I'm forcing every subroutine to be annotated with `async` and every call to be annotated with `await` -- enforced by the type checker, of course -- because there is simply no semantic difference in existence that allows one to differentiate between subroutines that need it and those that don't, so I figured it would be both clearest to users and most correct to just do it always.
No matter the language, doing more work is always more costly than doing less… Because of the GC (and not the JIT) at least you can implement moving stacks in JS, but that doesn't mean it comes for free.
> Also, "stackless coroutines without async/await" would give you (stackful) delimited continuations (albeit not multi-prompt). The reason Rust needs stackless coroutines is because of its commitment to "zero-cost abstractions" and high accidental complexity (and partly because it runs on top of a VM it doesn't fully control); surely JS has a different philosophy -- and it also compiles to machine code, not to a VM
???
> As to semantic differences, what is the difference between `await asyncFoo()` and `syncFoo()`?
That's an easy one. Consider the following :
GlobalState.bar=1;
syncFoo();
assert(GlobalState.bar==1);
The assert is always true, because nothing could have run between line 2 and 3, you know for sure that the environment is the same in line 3 as in line 2.If you do this instead:
GlobalState.bar=1;
await asyncFoo();
assert(GlobalState.bar==1);
You cannot be sure that your environment in line 3 is still what it was in line 2, because a lot of other code could have run in between, mutating the world.You could say “global variables are a bad practice”, but the DOM is a global variable…
It comes at extra work for the language implementors, but the performance is the same, because the generated code is virtually the same. Or, to be more precise, it is the same within a margin of error for rare, worst-case work that JS does anyway.
> ???
Rust compiles to LLVM, and it's very hard to do delimited continuations at no cost without controlling the backend, but JS does. Also, because Rust follows the "zero-cost abstractions" philosophy, it must surface many implementation details to the caller, like memory allocation. This is not true for JS.
> The assert is always true
No, it isn't. JS isn't Haskell and doesn't track effects, and syncFoo can change GlobalState.bar. In fact, inside some `read` method the runtime could even run an entire event loop while it waits for the IO to complete, just as `await` effectively does.
Now, you could say that today's `read` method (or whatever it's called) doesn't do that, but that's already a backward compatibility argument. In general, JS doesn't give the programmer any protection from arbitrary side effects when it calls an arbitrary method. If you're interested in paradigms that control global effects and allow them only at certain times, take a look at synchronous programming and languages like Esterel or Céu. Now that's an interesting new concurrency paradigm, but JS doesn't give you any more assurances or control with async/await than it would without them.
JavaScript is much more constrained than Rust, because of the spec and the compatibility with existing code. The js VM has many constraints, like being single threaded or having the same GC for DOM nodes and js objects for instance. Rust could patch LLVM if they needed (and they do already, even if it takes time to merge) but you can't patch the whole web.
> fact, inside some `read` method the runtime could even run an entire event loop while it waits for the IO to complete, just as `await` effectively does
No it cannot without violating its own spec (and it would probably break half the web if it started doing that). Js is single threaded by design, and you can't change that without designing a completely different VM.
> No, it isn't. JS isn't Haskell and doesn't track effects, and syncFoo can change GlobalState.bar.
Of course, but if syncfoo is some function I wrote I know it doesn't. The guarantee is that nobody else (let say an analytics script) is going to mutate that between those two lines. If I use await, everybody's script can be run in between. That's a big difference.
> because the generated code is virtually the same.
You keep repeating that over and over again but that's nonsense. You can't implement stackless and stackful coroutines the same way. Stackless coroutines have no stack, a known size, and can be desugared into state machines. Sackful coroutines (AKA threads) have a stack, they are more versatile but you can't predict how big it will be (that would require solving the halting problem), so you can either have a big stack (that's what OS thread do) or start with a small stack and grow as needed. Either approach has a cost: big stack implies big memory consumption (but the OS can mitigate some of it) and small stack implies stack growth, which has a cost (even if small).
You don't need to patch the whole web. V8 could compile stackful continuations just as efficiently as it does stackless ones. It is not true for Rust without some pretty big changes to LLVM.
> No it cannot without violating its own spec (and it would probably break half the web if it started doing that).
Yes, that's a backward compatibility concern. But just as you can't change the existing `read` and need to introduce `asyncRead` for use with async/await, you could just as easily introduce `read2` that employs continuations.
> Js is single threaded by design, and you can't change that without designing a completely different VM.
A single-threaded VM could just as easily run an event loop inside `read` as a multi-threaded one.
> Of course, but if syncfoo is some function I wrote I know it doesn't.
First, it can't entirely be a function you wrote, because it's a blocking function. It must make some runtime call. Second, this argument works both ways. If it's a function you wrote, you know if it blocks (in which case any effect can happen) or not.
> You can't implement stackless and stackful coroutines the same way.
Your entire argument here is just factually wrong. For one, all subroutines are always compiled into suspendable state machines because that's exactly what a subroutine call does -- it suspends the current subroutine, runs another, and later resumes (but you need to know how it's compiled, something V8 knows and Rust doesn't, as Rust runs on top of a VM). But even if you want to compile them differently for some reason, a JIT can compile multiple versions for a single routine and pick the right one according to context without any noticeable peformance cost.
For another, it is true that you don't know how much memory you need in advance, but the same is true for stackelss coroutines: you allocate a well-known frame for each frame, but you don't know how many frames your `async` calls will need. All that is exactly the same for stackful continuations. In fact, you could use the exact same code, and represent the stack as a linked-list of frames if you like(that's conceptually how async/await does it). There is just no difference in average performance, but there is more work. The allocation patterns may not be exactly the same (so maybe not recommended for Rust), but the result would be just as likely to be faster as slower, and most likely just the same, as async/await in JS, a language where an array store can cause allocations.
> Of course it can be tracked. It's all a matter of choice, and things you've grown used to vs. not.
What? When has the halting problem become “something you've grown used to”?!
You can even track space complexity in the type system.
We were talking about JavaScript, remember?
As to JavaScript, there are much more commonplace things that it doesn't track in the type system, either, and they're all a matter of choice. There is no theory that says what should or shouldn't be tracked, and no definitive empirical results that can settle all those questions, either. At the end of the day, what you choose to track is a matter of preference.
And you implicitly acknowledged this fact by using TPL as example: you know, the class of language where all you can write is a provably halting program (that's the definition of Total programming languages!).
The halting problem being a property of Turing-complete languages, you just sidestepped the issue here.
Has Rust officially relinquished small systems to C?
I wonder if we should think of end-user machines as such a constrained environment. Certainly, complaining about software bloat and inefficiency is a popular pastime here on HN and on related message boards. Maybe we have an ethical obligation to our users to make the most efficient possible use of their resources, regardless of the extra complexity we have to deal with. Of course, economic realities prevent us from really doing that in many cases.
Maybe you will. It's an attempt to address shortcomings of shared memory, just like borrow checker. Once languages start moving away from shared memory to actors, there will be no need for any of it. But it could take a long time, the industry dug itself way too deep into the whole shared memory world.
This example from the article was to demonstrate a type that was more complex than would be acceptable in day-to-day use, so they agree with you here, it would be crazy to make that required to use the feature.
By mentioning the languages after that comment, was it your intention to imply that OOP were added to them? Because to my knowledge, all three were OOP based from the very beginning.
Please comment with a source if you feel I’m wrong on saying any of those were OOP from the beginning. I (as well as others I’m sure) would enjoy the learning opportunity.