I think you’re too quick to dismiss the utility of yield-point cancellations. You’re right that yield points are a coarse means of imposing suspension/cancellation. However, I think they’re often preferable to the alternatives. Erlang, for example, imposes significant overhead due to reduction counting and restrictions on copy-free data sharing in order to achieve (nearly: dirty NIFs are forever) complete arbitrary subroutine control from the outside along with progress guarantees.
From a certain perspective, the knowledge that suspension/cancellation can only occur at yield points shrinks (or at least makes somewhat more vislble) the risk surface of areas where arbitrary cancellation can leave the program in a corrupt state.
I do agree that more work is needed in this area to ensure automatic state reconciliation after cancellation (the mutex-held-during-await issue). Python makes a decent stab at this: its async/await concurrency implements cancellation as an exception raised at await points, allowing cleanup to occur before proceeding with cancel. Unfortunately, that comes with the consequence of allowing routines to choose to defer cancellations, or ignore them entirely. Perhaps that could be mitigated with an async/await system that makes cancellation a bidirectional exchange, changing observable behavior at the canceler when the cancellee does or doesn’t actually terminate? I think the Rust folks are moving in this direction with async-drop, though I personally do not like the idea of implicit custom destructors in the first place, and worry that allowing them to effect async cancellation control flow will hurt rather than help. That’s a separate topic, though.
All that said, I think that struggle is a somewhat inevitable outcome of a platform that supports shared memory between routines and C interop. If you’re willing to give up on those, then sure: you can preempt or cancel with much more flexible semantics. But I don’t think that’s innately better: first, C interop is critical for some platforms, as Boats mentions. Second, while you might be right in pooh-poohing some of my other points as “just efficiency”, adopting an Erlang memory sharing model is really expensive—impactfully so for a much wider variety of programs than the ones for which, say, stackful vs. stackless coroutines make a difference.
On to your next point: how important is yield-point cancellation, really? I think that
> the only case where task cancelation can be useful in a shared-memory system like this is when it's just an efficiency hack
isn’t correct.
Firstly, “efficiency” is kind of a moving target. Is it merely efficiency if we can, say, take advantage of the fact that memory used by cancelled coroutines is freed once they’re cancelled? Consider this code:
https://gist.github.com/zbentley/f45ea39843f533f5f54b23f8a49...
If cancellation of the first call to memory_hungry() did not trigger the drop of x (90% of system RAM), then we’d be unable to successfully call memory_hungry() a second time after the select!, since that would put us into OOM.
Like, sure, there are all sorts of ways to accidentally leak resources in the event of cancellation (we could mem::leak it, could have memory_hungry() spin and block the task, and so on). But this is one fairly happy path where a very important resource (memory) is correctly managed by the cancel-at-yield model.
Secondly, what about the cancellation of routines that keep something alive? For one example, consider routines that send heartbeats or liveness probes: cancelling those can be very meaningful to the program’s semantics. If your rejoinder is that that’s a bad idea because of the chance of a while-true loop blocking routine causing a spurious event loop hang, I’d respond that we’re either now discussing a whole different set of constraints (RTOS behavior), or that we’re back at the original tradeoff between a costly-and-C-incompatible thick runtime like Erlang’s, and some inescapable risks that come with the territory of using both coroutines and ordinary pthreads.
On to the ergonomics of heterogenous multiplex-waits. Consider this code that waits for the first of a timer, a background thread doing some work, and an IO event:
https://gist.github.com/zbentley/d05696b1f6268372cf858b157c6...
Regardless of what cancellation semantics are or are not present, how would you implement that code using threads? It’s certainly possible, but it’s tricky: once you have the background thread created, how do you wait for it with a timeout (so you can detect the other events occurring before that)? You probably put a condition variable/mutex in it and wait on that with a timeout, since the underlying pthread_join doesn’t take one. That’s extra work. Do you farm the timer out to another background thread? The cost is negligible unless you need a huge number of them, but it’s some extra work to code up if you do. Now how do you wait on either of two background threads’ conditions with a timeout? Perhaps you switch to a semaphore--make sure you use it right to avoid deadlock (extra work)! But the socket read is I/O, so perhaps it makes more sense to model the timer as a timerfd and poll/epoll/select over both FDs instead: tricky decisions, extra work. Or should the timer just be tracked in memory and implemented via the timeout to select/epoll (decisions, work)? Then, how to combine waiting on that IO multiplexer with waiting on the background thread? Perhaps we just pretend the thread is a process and give it one end of a pipe to write to when it finishes, and select/epoll on all of our file descriptors to detect what happens first: a nice consistent outcome, but lots of extra work (congratulations, you’ve written a rudimentary event loop! Hope you didn’t write any bugs!). Or maybe this is all in a request handler, and while we might be OK with thread-per-request, we aren’t OK with multiple threads per request, so we move the background job to a thread pool via a queue, or the timer to a pooled timer thread, or the IO to an epoll thread, so we need to bring in some kind of system that models channels … you get the idea.
Backing up: sure, all of those are solvable problems using non-async/await models. But that’s not “just efficiency”: that is a lot of decisions to make and more than a little code that you need to get right just to wait on three heterogenous events. We haven’t gotten into how to bookkeep contingent state machines (e.g. background-thread completion triggers a different continuation than timer completion) or how to add more types of event (signals, kqueue, inotify) into the multiplexer.
In other words, you are 100% correct, in that this is
> the kind of thing you could put into a library function in a system using multithreading
That’s what an async/await system is: a good utility library interface (that may or may not use threads) to code that wraps all of those concerns while providing a sane continuation and cancellation system.
That’s not nothing! Especially when you consider that it’s one of the few means of modelling that kind of concurrency without trading away runtime efficiency (no reduction counts, GC pauses, or added golang-style preempt points), resource conservation (it’s generally sensible and not as hard to understand as bare threads re: memory, and also does reasonable things re: file descriptor and background thread counts so those don’t get exhausted), and FFI models.
If those aren’t priorities for you, then yeah, it might make more sense to use an Erlang-ish or Golang-ish system, or just good ‘ol pthreads. But I don’t think it’s fair to say “efficiency is the only benefit”.
> you can totally get data races in a multithreaded system that is running on an in-order single-processor system
Fair enough. s/parallelism/preemptive concurrency models/.