No, Linux is rubbish. Seriously. FreeBSD does this properly.
Edit: FreeBSD, Windows, OSX, Solaris, AIX, HP-UX(?)...
No, Linux is rubbish. Seriously. FreeBSD does this properly.
Edit: FreeBSD, Windows, OSX, Solaris, AIX, HP-UX(?)...
FreeBSD does page faults just like everybody else. Your process blocks until the memory is read in.
And the bit from the paper that is relevant:
> A non kqueue-aware application using the asynchronous I/O (aio) facility starts an I/O request by issuing aio read() or aio write() The request then proceeds independently of the application, which must call aio error() repeatedly to check whether the request has completed, and then eventually call aio return() to collect the completion status of the request. The AIO filter replaces this polling model by allowing the user to register the aio request with a specified kqueue at the time the I/O request is issued, and an event is returned under the same conditions when aio error() would successfully return. This allows the application to issue an aio read() call, proceed with the main event loop, and then call aio return() when the kevent corresponding to the aio is returned from the kqueue, saving several system calls in the process.
Also, using a worker pool for I/O scales just fine and has no real disadvantages when talking to a disk or SSD.
Pretty simple stuff really.
2. kernel: hey thread, that page is still on SLOOOOW spinning disk, why don't you go off and do something else while I get it for you. I'll let you know with an event notification, so be sure to check in with kqueue regularly.
3. thread: OK then, i'll go off and serve these other people while you do that for me. kthanx.
4. kernel: hey DMA controller, I'd like you to get sectors 4, 5 and 6 from platter 3 on HDD 2 and load them into memory address 0x4fe6bb. Please send me an interrupt when done.
5. DMA controller: OK servo, please adjust read head to this offset. Read head, read me those bytes. Memory controller, please store bytes at address 0x4fe6bb. Hey CPU, here's an interrupt to store in your interrupt table, please wake up the kernel guy and let him know.
6. kernel: wow I just got interrupted. The interrupt seems to map to a request for data from this particular thread. Better let him know. (sends event up thread's kqueue).
7. thread: hey, I just got a kqueue notification that the file is now ready to be read. That means it must be in memory...cool!
Make sense now?
Yes, but the one thing they don't give you is notification the data is in memory. Do you really want to spin a hot loop calling mincore?
> However, once you get an actual page fault, the only option is to block the process.
See steps 1-3 again.
2. while it is being paged into memory, <DO OTHER STUFF>
3. now your data is in memory and you can update it (writes are async by default even on Linux, as the write just goes into memory and will get synced out by the kernels page sync mechanism but you can override that by setting the O_SYNC flag).
Writing different code is not the question I asked. That's not an OS feature. Do the OSes you like have the latter ability, or are you blowing smoke?
The POSIX semantics on the other hand are simply that the read won't block due to lack of input, so it does as much as it can straight away and then returns. If there's data, but the buffer is paged out, the non-blocking read call has to take more time, because the problem is a paged-out buffer and not lack of input.
(The Windows equivalents of man page section 1 are so awful that most POSIX fans just run a mile and install cygwin. More fool them; once you get to man page section 2, it's a lot better. MAXIMUM_WAIT_OBJECTS is lame, though.)
Well you are asking about doing something non-blocking. That means you must be in some kind of event-loop scenario, otherwise you would be happy with synchronous operations. And yes, you would need to add a read event on the file, and then trigger the socket.send() once the data is in memory.
I never said it was free. But at least it is viable.
In the Redis article, it is talking about in-memory data structures. The page fault happens from following a pointer to another location in memory, for example, a linked list, or a skip list. There is no read call to replace with an async read call. Your code evaluates the pointer to a page that has been swapped out, a page fault occurs, and the OS has to swap it back in for you to be able to read that memory.
You could potentially create a new signal, for page faults (maybe there is one I've never heard of), but that would still not let you continue executing from the previous location.