Understanding the code inside Tornado, the asynchronous web server
golubenco.org
golubenco.org
Let's say I'm using Java instead of Python. Lets say I use a lot more threads than 20. I will reach thousands of requests per second this way. There are real problems with that approach, and event-driven approaches have genuine advantages, but can people please stop straight-up lying and saying that you can't just throw a few thousand threads at a problem, because you can, and people do.
Much of the translation from blocking style to event-based style is moving the work of dispatching and looping from kernel to userland. Other ancillary benefits, like reduced address space usage by blocked threads, are in principle also achievable in a threading model - e.g. by storing stack frames on the heap and being more aggressive about collecting them (assuming GC).
Other benefits of async - such as overlapping work - are also fairly trivially possible with threading, though less deterministic.
Languages/frameworks where async IO is an implementation detail (like some web frameworks in haskell) don't have this issue.
One really cool one is http://thomas.pelletier.im/2010/08/websocket-tornado-redis/, which demonstrates how to use threading to support Redis pub/sub.
For heavy computational tasks that aren't time-critical, you could have an accessory worker thread that chugs queued computations (yay first-class functions) in computational downtime.
Still, this is a good general overview of Tornado's internals.
[1] https://github.com/facebook/tornado/commit/b6c4d6d20196fa4fe...
In an event based system the overhead for each connection is at least three orders of magnitude lower, sometimes four or five (hope I'm remembering this right). This translates into _considerable_ increase in number of connections that can be handled simultaneously, not quite the equivalent number of orders of magnitude due to other aspects of the system becoming more of the bottle-necks, but still dramatic.
This was a data-driven observation, not conjecture or assertion. I listened in during a break while the nodejs guys argued about approaches to getting TLS working better - they were worrying about the 1MB overhead for a TLS connection because that's a significant percentage of a connection handler. Think about that for a minute in the context of an apache threaded instance. 1MB matters? wow!
I'm new to this area and this was very interesting stuff and seemed related to the various discussions below about a threaded system can do what an event based system can do.
Wait, that is not correct. His IO is blocking. He gets a notification when data is ready, from the select/poll/epoll (the asynchronous part), but when IO is actually performed (read, write, recv, etc), the operation it is still happening in the main thread in user space, and it blocks it.
Currently only the file system (I think) has truly asynchronous and non-blocking IO. It is provided by the aio_* set of system calls and is has been sort of an exotic beast (it is not that popular).
Here is a good chart of the possible IO types and their combinations:
http://www.ibm.com/developerworks/linux/library/l-async/
And of course:
I dont think he is using aio_* which Linux does not really implement usefully.
In common usage, an I/O operation is considered "blocking" if the system call does not return until data is available. If on the other hand you set O_NONBLOCK on a fd, you will get EAGAIN when no data is available, so clearly that is the accepted meaning of "nonblocking."
Yes because there is _another_ mode of IO where it doesn't have to do that. So the reason that is called "blocking" is because there is another mode of operation (albeit it is exotic) where the process does not block while performing IO -- the kernel reads/writes the data on behalf of the process. If this second mode of performing IO did not ever exist you could just interchange arbitrarily asynchronous and non-blocking as synonyms.
Since you didn't bother reading the link I posted here is the basic matrix of IO operations:
Synchronous Asynchronous
Blocking read/write select/[e]poll
Non-Blocking O_NONBLOCK+E_AGAIN aio_*
There has been talk and various attempt at implementing aio_* style IO for network sockets on Linux, but nothing good so far.> Even aio_read() performs some work inline (namely enqueing a read request).
Enqueuing the read request is not the same as actually copying gigabytes or terabytes of data from disk where the actual work is performed.
I think you'll find these links more helpful than the misleading one you've been using.
http://www.circlemud.org/~jelson/software/fusd/docs/node36.h...
http://people.freebsd.org/~hmp/stuff/docs/freebsd_kse.pdf
I could always be wrong about this stuff; maybe I'm the one with the broken semantics. But I'm pretty sure I'm not.
Actually I did. You are misinterpreting what it is saying. The reason it is calling the select() model "blocking" is because you block during the select(), not because the read() itself is blocking.
This is also non-standard usage: most people wouldn't refer to a select()-based loop as blocking I/O, because a program generally only calls select() when it has nothing else to do but service I/O, so the select() does not "block" the application.
On a nonblocking socket, the I/O operations you're talking about are simple u/k and k/u buffer copies.
I think between aio_+ and the C10k page, you may have gotten a bit scrambled. Whether you have an event loop or not, disk I/O is often transparently blocking, even when you try to set descriptors nonblocking. But I/O operations on a nonblocking socket don't wait the process. If there's data in the buffer, you get the data; if there isn't, you get the error.
But because there is an IO mode where these copies are done by the kernel when data arrives not by the user process (after a kernel notification), there is now a differentiation between blocking and asynchronous. Asynchronous refers to IO readiness notifications, "blocking" refers to the copying of data (the actual IO if you will).
> A process blocks when it needs to be rescheduled, yielding back to the process scheduler pending I/O. That's what "block" means.
However when your process is copying data it is in the running state and preventing other processes from running. It could be doing something else or it let other processes run in the meantime.
> The I/O operations you're talking about are simple u/k and k/u buffer copies.
That is true, however if the data comes in very small chunks very fast you are doing a lot of switching to user space and a lot of small context switches to, say copy 1K of data. When you could just request that the kernel fill up your 10MB buffer with data from a socket and tell you when it is ready. If there is anything I learned is to never just say "it is a simple copy". Today's memory is not very fast compared to CPU speeds and copying is not something to be taking lightly. It is one thing when looking at a toy example, another thing when dealing with realtime systems or large data sets.
Yes, if aio_* system was never invented, we wouldn't be arguing, but because it exists it introduces a new possible way of doing IO. It is a general enough way of doing IO and at hardware level we have DMA but in the land of syscalls we have aio_.
Ok, just curious, what words would you use to describe what aio_ calls do?
Over the years many have thought of a way to bring that style of IO to networking. So far it works for disk IO very well. For example take a look at this benchmark from lighttpd for a sendfile:
http://blog.lighttpd.net/articles/2006/11/09/async-io-on-lin...
It looks like there is a consistent 50% improvement in speed when using aio with a 1MB block.
Wouldn't it be nice to have that kind of improvement for network sockets as well? I think it would be and there have been many attempts over the years but nothing good yet has happened.