> Preemption is costly, because you need to save the entire state of the execution you're preempting, in order to restore it later (which is also, as you should know, what makes context-switches slow).
In-userspace stackful context switches are not at all slow in a sane userspace-switching runtime. They are approximately a combination of setjmp() and longjmp(), but without saving all registers. Typically a few tens of nanoseconds or less on a modern CPU. Pre-emptive context switches are nearly always stackful, even if non-pre-emptive (async) context switches are stackless in the same runtime.
A pre-emptive context switch in a userspace co-operative scheduling system is slower than non-pre-emptive context switch if it is caused by a dedicated interrupt of some kind. In the case of userspace pre-emption, generally a signal and return, typically single digit microseconds.
However, the cost there is mainly the signal. If the pre-emptive context switch is caused by a signal that triggered anyway for another purpose, the actual context switch is, again, cheap.
In the model I described (the game environment), pre-emptive context switches are rare. It wouldn't matter if they took longer, because >= 99.99% of context switches are co-operative in that model. The important factor is that no task blows the frame budget or prevents other tasks from running at the full frame rate. In web services a similar target is tail latency, in the presence of diverse tasks you cannot predict in advance.
> meaning you can choose what works best for your workload.
Exactly. And as I've tried to explain, for some varying, complex workloads whatever you choose statically has poor metrics; it must be adaptive to maintain good timing metrics.
> Rust yield point are manually inserted by the developer
Indeed, and in the case of NPCs in a game engine, or a large program composed of many libraries written by hundreds of different authors (e.g. some web servers), that causes complex, interdependent performance characteristics, where timing behaviour in one of them ruins performance for everything, unless they are isolated using threads, in which case it's too slow. There is no "the developer" who can ensure this doesn't happen.
(Aside, if Rust's type system helped with this, that would be great, the same way types help with other "programming in the large" safety issues, but it doesn't address timing characteristics as far as I know.)
> leading to the best performance [..] Rust has both Async tasks and OS threads, meaning you can choose what works best for your workload.
You could summarise my point as: "For some dynamic workloads, especially in timing-sensitive large programs with many components working independently and unpredictably, neither async tasks or OS threads perform best for your workload (or even adequately sometimes). The optimal (or required) combination requires some dynamic responses, and cannot be achieved solely by static placement of yield points and thread initiations."