Agree with the grandparent. Async/await is unnecessary complexity in all but the most performance-oriented languages (like Rust).
Very curious to see how Java Loom will pan out in practice.
It is about doing things in parallel. There are many ways on how to achieve that, I am just nit picking on performance.
Go is asynchronous-first, it just doesn't use the (flawed) async-await semantic to model asynchronicity.
The vast majority of projects don't use CGo, and of those that do, the majority doesn't call CGo in a tight loop, making the performance impact mostly irrelevant.
So it's really coming back to the performance-orientedness of a language.
.NET is the opposite - it was designed from the get-go from cross-language and cross-runtime interop. Back when Java folk still had to struggle with JNI to use anything native, .NET 1.0 had P/Invoke and COM Interop. And now that it has async/await, it's all mapped down to a C-compatible ABI for interop purposes, so you can easily write async code in C# that calls into async code written in C++ that calls back into C# etc.
assert P;
foo(); // blocking call
assert Q;
and assert P;
await foo(); // await on an async call
assert Q;
?There may be some (e.g. thread identity is preserved in the former but possibly not in the latter, depending on implementation, and that could be expressed with proper assertions), but they only highlight how the differences are merely technical and certainly not "inherent".
The question is rather: is it possible to remove the await and just automatically rewrite every program to use await where necessary?
And the answer is no, because with only sync functions you have the guarantee that if your thread crashes, every code immediately stops executing. Not so with async functions. Once they are run there is a chance that your own thread crashes but some other thread keeps executing more code, at least for some time, potentially executing side effects.
It cannot be automatically decided if this is okay or not. Or to turn it around: the behaviour of a program that automatically awaits async functions (which formerly where sync) can be different.
There's obviously a difference between launching an operation to run concurrently with the spawning code and launching an operation and awaiting its completion, but that difference does not require the introduction of an async/await feature, and is perfectly expressible with nothing but threads and synchronous code. In particular, there is no need for two separate syntactic domains between async and sync subroutines. Indeed, in Erlang, Go, and Java, all the same behaviour of async/await is achieved with no such need, and no need for different kinds of APIs.
Of course there is. The former constraints the type (or return-type) of the function it is used in, the latter does not.
> but that difference does not require the introduction of an async/await feature
I don't remember saying anything about async/await in my original post though.
That's not a semantic difference, but a syntactic one. Syntactic differences are expressed by which programs are accepted or rejected by the language; semantic differences are expressed by programs exhibiting different behaviours.
To show a semantic difference you need to find P and Q such that when the value of P is the same in the two programs, the value of Q could be different; that's what expresses the semantics of the code between P and Q. Again, there are some (e.g. P: `(t = currentThread(), true)`, Q: `currentThread() == t`), but they're purely technical.
> I don't remember saying anything about async/await in my original post though.
You said that the lack of need of a syntactic kind for asynchronous routines is a fallacy. It is not. Everything that you can describe with sync and async routines your can describe, with the same ease and semantics, with just sync routines and threads. There is no need for a special syntactic category of asynchronous subroutines.
I agree with that statement.
> semantic differences are expressed by programs exhibiting different behaviours.
It can cause different behaviour whether code runs on the same thread or on a different one (see one of my other posts here). That doesn't have to happen in every program and it doesn't have to happen on each execution, but that still makes it a different behaviour - at least in my terminology.
> You said that the lack of need of a syntactic kind for asynchronous routines is a fallacy. It is not.
That's not what I said (or at least not what I meant). I meant that it is doomed to fail to try to generally unify sync and async code on a language (or runtime) level.
Wrong. Languages that definitely got it right:
Erlang/Elixir (high level), especially elixir's Task.async
Zig (low level)
Very close runner-up: Go
Task.async/1 (or Task.async/3) at a high level just spawns another BEAM process links it to the currently running process, sets up a process monitor, and then runs. Task.await/2 just pulls from the mailbox of the current process for a returned message.
There is a semantic difference between just running a line of code or running ith with Task.async right? So then there is a distinct concept of sync vs async.
IOW, async/await in the BEAM is a high level thing (like goroutines) which is one of two reasonable choices, the other reasonable choice being "what zig did".
Task.async/3 and Task.await/2 are, from the thousand foot view, the equivalent of:
parent = self()
spawn(fn -> send(parent, do_something()) end)
receive do
result -> IO.puts result
end
In Elixir/Erlang parlance, I have just spawned a thread where the BEAM VM handles both the internal data structures to facilitate concurrent-safe communications between threads and the scheduling in userspace.Put more generally, the mentality of concurrency in Elixir/Erlang is the equivalent as that of traditional pthreads in C.
Is it (except for performance) always 100% equivalent whether a calculation is done "locally" or whether it is ran on some other, potentially just spawned, thread and then retrieved from this thread?
If your answer is "no it is not always 100% equivalent" then we agree (and I think we do).
An Elixir, Erlang, or any other developer using the BEAM VM (or at the very least myself since I don't want to generalize) doesn't make a semantic distinction between spawn and Task.async/3 because there is none. The significant differentiation is whether or not the caller is expecting a returned value, if not then Task.start_link can be used instead.
I don't need to worry about function colouring because there is none. I also don't need to worry about async vs Task.run like in C# because in Elixir there is none. To be asynchronous in Elixir/Erlang/etc... is to be synchronous. To equate it to modern backends, we basically have a microservice but every service can directly call each other.
I think I also might have been talking around you a bit, the BEAM VM sets itself apart from C#'s async/await or the TPL due to its implementation of the actor model. This video by Sasa Juric can explain better than I can the general overview of the BEAM: [1].
But a rough tl;dw, a BEAM process is to the BEAM, what an OS process is to the OS. We don't define a main function as the process supervisor in BEAM spawns processes we define instead and those are our "entrypoints."
Moreover, there is a context switch and it's more likely the code will run on a different core if you run it async than if you run it in the same process (which is very likely but not guaranteed) to run on the same core.
The way in which they are equivalent, is that the code that you write is identical, and the bytecode that gets run, and the existence of implicit yields is identical between the async and "sync" code.
> Moreover, there is a context switch and it's more likely the code will run on a different core if you run it async than if you run it in the same process (which is very likely but not guaranteed) to run on the same core.
Yes I agree, but the question as presented by valenterry asks us whether or not there is some existing semantic difference and not under the confines of performance (both gains or losses). Regardless what you have stated are all true.
But as an aside, and not directed towards you, Task.async/3 doesn't do anything that the developer cannot already do. Even in the tutorial for the learning Elixir, the fledgling developer is exposed to the different mechanisms that power Task.async/3, the source code [1] reflects this, although supervisors are covered much later down the line. The documentation for
> The way in which they are equivalent, is that the code that you write is identical, and the bytecode that gets run, and the existence of implicit yields is identical between the async and "sync" code.
And just to add on for those not familiar with the BIFs, receive, which is what Task.await/2 and Task.yield/2 use under the hood, yields execution. NIFs are another one.
[1] https://github.com/elixir-lang/elixir/blob/fb729784e5504f499...
Sure, but that was not my question. 100% equivalent means that there cannot be any observable change in the behaviour of the whole program/application (except for performance).
Or in other words: is the sole purpose of the existance of `Thread.async` (or spawn, pick whichever you want) to change performance characteristics? If not, what is the purpose of its existence then?
In a vacuum (i.e. via spawn/1 or spawn/3) there is no observable change in behaviour throughout the whole system.
If they're linked via spawn_link then a child process crashing means the parent process dies with it unless handled in some way.
So ultimately, no there isn't observable changes. My logging process doesn't care or even know that my WebSocket process crashed.
> Or in other words: is the sole purpose of the existance of `Thread.async` (or spawn, pick whichever you want) to change performance characteristics? If not, what is the purpose of its existence then?
As stated before Task.async/3 is just a convenience wrapper around low-level BEAM primitives, there isn't anything special about Task.async/3 that you couldn't do via spawn.
The reason being that the BEAM VM schedulers prioritizes overall latency of the system over throughput by enforcing reduction calls (parlance for a function call, but specifically for tracking looping in the BEAM) at around 4000 calls, then it moves the current process back into the queue; repeat ad infinitum.
So you don't get massive speedups just spawning a new BEAM process but your independent units of execution still enjoy the same overall latency as before. If you want performance increases in the BEAM you would need to jump to dirty schedulers and NIFs which carry their own dangers.
Okay, let me follow up on this one. So could we take arbitrary Erlang/Elixir programs and automatically/mechanically rewrite them to push previously inlined calculations to be run on a different process (using spawn) or the otherway around - and that without causing any observable difference in behaviour in any situation except for difference in performance?
> As stated before Task.async/3 is just a convenience wrapper around low-level BEAM primitives, there isn't anything special about Task.async/3 that you couldn't do via spawn.
You can likewise answer my question for the BEAM primitives (like processes/threads): is the sole purpose of the existance of those primitives to change performance characteristics? If not, what is the purpose of its existence then?
Also, just so that we understand each other: I'm a big fan of the BEAM. This is not a criticism or anything; I just want to explain why I believe that function coloring (or however you call it) can never be conceptionally be removed from a language if sync/async is to be supported.
We probably should've established this way earlier, sorry if I came across as pretentious or off putting.
I'll answer in a bit of reverse order.
> I just want to explain why I believe that function coloring (or however you call it) can never be conceptionally be removed from a language if sync/async is to be supported.
To respond in more of a roundabout way, why doesn't C have colouring issues if it too supports asynchronicity (I'm pretty sure I've made up a word but hope the point gets accross fine) via pthreads?
The colouring issue is, at least under my understanding of C#'s async/await system is to accommodate for the fact that Rosyln transforms the C# source into .NET IL with a state machine. It's, to me at least, just a consequence from what layer of the stack we talk at. Erlang doesn't need to worry about since the runtime was designed to run with this specific model of concurrency in mind; but if we look at C#, .NET only supports using OS threads (though David Fowler is investigating green threads in the, iirc, labs repository on the DotNet Github) the async/await is, for lack of a more better word, bolted on. It's to my understanding that IronPython (or maybe it was PythonNet) has difficulties calling into C# async for that reason; though if I'm wrong please correct me.
> is the sole purpose of the existance of those primitives to change performance characteristics? If not, what is the purpose of its existence then?
Fault tolerance, latency, scalability, perhaps minor throughput increases (yeah this is a bit of a contradiction to what I said earlier, but I've revised my opinion that you wouldn't see massive performance gains unless resorting to NIFs).
The Phoenix Framework's 2 million websocket connection challenge I think is a best demonstrator for the use case of BEAM processes. A specific websocket connection dies? No problem, the Supervisor respawns it. Two million connections receiving a Wikipedia length article on a topic without the entire system coming to a crawl? No problem. There isn't a need to worry about kqueue, poll/epoll, or IOCP here, just fire off a task for the purpose and let it do it's thing.
But overall you wouldn't spawn a BEAM process just so you could compute the matrices behind RAID6/RAIDZ2, you'd delegate those to NIFs. The biggest gains in overall performance came from BeamAsm instead.
> Okay, let me follow up on this one. So could we take arbitrary Erlang/Elixir programs and automatically/mechanically rewrite them to push previously inlined calculations to be run on a different process (using spawn) or the otherway around - and that without causing any observable difference in behaviour in any situation except for difference in performance?
Yes. Using the case of the 2 million websocket connection challenge earlier, it doesn't really matter if each of those connections spawned another process as part of its routines, the other processes don't know and don't care. Taking a more generic case of Elixir's GenServers (or gen_server in Erlang), when I perform a call (as in the GenServer behaviour) it doesn't matter to the calling process what happens behind the scenes, it just blocks and waits for a response. The GenServer could fire off any number of processes it wants but the calling process doesn't care about that.
But it does have those issues. In fact it is worse than in most other languages - just because there is no distinct support of async in a language does not mean it doesn't have async capabilities. C does have those, you just don't have any language support whatsoever, you have to take care yourself. Delegating the execution to the OS doesn't change that.
For example, if you want to transform elements in an array concurrently, you now have to use pthreads. You cannot do that while using the same code as before (which is would removing function coloring is all about).
Maybe we have a misunderstanding here: I'm talking abouyt sync/async not the async/await syntax that many languages have. The latter is to help deal with colored functions and the former is what enables the existance of concurrency. Any language that is - in one way or the other - capable of running code concurrently is suffering from that problem. It can be avoided by enforcing synchronous execution only, which some languages do (especially some DSLs), but most general purpose languages support concurrency.
And once they do, there's two choices: the programmer has to deal with it (using Thread.async, spawn, pthreads, go channels, promises, ...) or they don't. If they don't then sync and async code must look exactly the same (or be able to be converted forth and back automatically). And this goal (either making it look the same or converting it automatically) is just not solvable in general. It has been tried for decades.
For instance, check the paper "A Critique of the Remote Procedure Call Paradigm" from 1988. I think it was from Tanenbaum. It lists some of the problems that just cannot be solved in general. Some programs will always behave different after the automatic conversion, it's impossible to prevent that.
> > Okay, let me follow up on this one. So could we take arbitrary Erlang/Elixir programs and automatically/mechanically rewrite them to push previously inlined calculations to be run on a different process (using spawn) or the otherway around - and that without causing any observable difference in behaviour in any situation except for difference in performance?
> Yes.
What I'm trying to explain is that you are mistaken. It just is not possible. Maybe in specific cases, sure, but not in a safe way for every usecase that can be thought of.
And neither Erlang nor Elixir (nor the latest BEAM language, I forgot the name), luckily, try to do that. They give this controller to the developer and let them deal with the decision of how and when to use concurrency. They just try to make that as easy as possible, but they don't try to make sync and async code look and be the same.
vs-threading should be part of the .NET standard library but somehow isn't.
I knew Go can do it, so it doesn't "always fail"
What if you are reading the file and somewhere else in the code there is some logging happening. Should the logging flow stop until reading the file is done? Should it not? If not, what should happen if the "thing" (e.g. thread) that executes the code to read the file crashes?
Go does not treat sync and async the same. It specifically has syntax to deal with async, e.g. go routines.