If you manage your own memory/swap, at least you can use async IO and free up the thread while the IO request is being served by the OS.
If you manage your own memory/swap, at least you can use async IO and free up the thread while the IO request is being served by the OS.
In short, mmap() and friends were designed for a world where context-switching (e.g. multithread or multiprocess) is a great idea. Unfortunately, context-switching has become extraordinarily expensive for a lot of server software which makes multithreading a less palatable option.
These days, you want (1) native async I/O, (2) strict control of cache replacement behavior, (3) strict control of I/O scheduling, (4) minimal context-switching, and (5) memory locality control. On Linux, this means DIRECT_IO, io_submit, managing your own physical cache RAM, and locking one thread to every core (ignoring hyperthreads for the moment). This is more complicated to implement because there is not a simple, portable interface like mmap() widely available but it is also much more efficient and performant when done correctly. To make matters worse, some useful interfaces (like io_submit) are poorly documented.
But yes, if you build your own memory and swapping subsystem directly on top of the native interfaces, and use threads as an abstraction of a core rather than a swappable context, you can build very efficient server engines that stall minimally on I/O. (Note: even using io_submit and IO_DIRECT for async disk access, there are conditions that can cause blocking. They are just much rarer and easier to manage than mmap().)
If you want your program to take page faults as PHK suggests, it has to be multithreaded. If you choose event-driven concurrency you can't afford to take page faults in mmap() or read(). When you make the threads vs. events decision you're implicitly making a bunch of related decisions about I/O and scheduling as well; a hybrid approach (like using events and mmap) won't work well. http://news.ycombinator.com/item?id=1760642