The rate of computation is obviously very similar for threads and async, but the rate of multiplexing is not. Switching between processes or threads takes more time than switching between fibers/coroutines. So the general observation that for massively multiplexed workloads fibers/coroutines are lower overhead is generally correct, but the more interesting question is of course "what qualifies as a massively multiplexed workload?" ... I'd argue most kinds of web application servers don't.