> I don't think this is true if you create have approximately 1 thread for each CPU core (which is the ideal scenario in a non-async app - why would you create more threads than CPUs?).
Right, so then how do you structure your program - particularly an I/O-bound program - to use that small pool of threads effectively?
Imagine you're building a REST frontend that queries a couple of datastores and aggregates the results - you get a web request, fire off a request to some coordinator datastore, get a response back from that, fire off a bunch of requests to other datastores based on that, then form the results into some JSON and send them back to the client. If you do blocking requests then you waste most of your threads most of the time (they'll be idle waiting for responses from the datastores).
If you use NIO then that solves that problem, but you've still got to get each thread to actually do the right kind of processing - you want each thread to run an event loop where it checks for returned results from the datastores, matches those up with the requests that they belong to, and does the next step of processing for that request whether it's firing off more requests to other the datastores or composing together the results and sending them back to the original client. And you've got to also handle timing out stale requests etc.
Now you can write a server literally like that, with a global buffers for each possible request type that hold the state that you need to pick up handling that request again. But it's a nightmare to maintain, as you're effectively forced to write unstructured programs where you do a GOTO for each I/O operation - there's no easy way to trace through the handling of a single request or run part of your program for testing.
So people generally prefer to use an async framework that will handle keeping track of each suspended request and what to do next when the result comes back, and multiplex those onto the native threads for you. That could be something where you call an I/O API and pass a callback to run when the result comes back (and the framework takes care of running the callback on some thread when the result is ready, without blocking a thread in the meantime), or something where there's a special operator to say "suspend this function here, package up the rest of it as a continuation, and continue with that continuation when the I/O result comes back". Or, as Loom is doing, it could be something similar that works completely invisibly whenever the programmer calls particular functions.
There are tradeoffs to all these approaches to async, but any of them is going to perform a lot better than the naive approach of just using a native thread for each request, and they're all (hopefully) going to be more maintainable than manually writing an event loop that does the right thing each time.