If you restrict yourself to 64 bit systems, you can essentially go back to using OS threads with blocking I/O because the address space is so large. You can fit thousands of 2MB call stacks into the processes address space and rely on the OS and MMU to manage the memory.
The problem that created the whole non-blocking, async domain was the memory needed for the thread's call stacks. You don't have enough address space to place a bunch of 2MB call stacks for each thread, if you're handling tens of thousands of connections.