A take on Tornado and "modern" web servers
audidude.com
audidude.com
Why does it matter if you use 4 or 8 event-handling processes on one machine, each with their own database connection? If that's your scaling bottleneck, you'd be running on many machines anyway, so a small constant multiplier doesn't change the story. Besides, you could use crap like ODBC to have one persistent DB connection per machine.
The other issue (a small number of processes accepting work) is a good idea though I think, to get the number of processes doing actual work closer to the number of cores.
After all, even if it is a 'small multiple' it is a divider when it comes to how much hardware you need and less hardware=better.
The DB may be a bottleneck, it may not be, that very much depends on the workload of the server.
The idea of rolling multiple interpreters into a single process is especially bad for python, as it has a real garbage collector and os.fork()-ed processes will share a lot of memory. It makes some sense for Ruby, since its assy mark-and-sweep will invalidate every single page on the next collection after a fork().
You can still separate it into two fundamental shifts. The first being making a fully asynchronous/event driven module/handler API. This should be done regardless of the runtime module.
The massively-scaled websites I've seen would tip with 4-8 times the sql connections. SQL Server is especially bad at this. It also makes it hard to simply throw additional hardware at the problem.
I guess you could yield on resources from a master process if you have the logic to know their status.
Without access to the original post there, this noob can only assume(hope) the discussion is about the Tornado Web Server (http://www.tornadoweb.org/). I'm interested in alternatives to Apache, so am looking forward to you being back online. Gracias.
That said, I think the big pain point is really the data store in any case. Yeah, it's nice to handle more with less, but in the end, adding django/rails/php/whatever machines is easy and a known quantity. It's tougher to scale up the data store.
Well like I mentioned I believe that is an implementation detail. For example, if you have the subprocess you have two fundamental designs. The first being where the master process still manages the client socket (so data is transfered from the worker back to the master). The second being where you pass the client socket (over a UNIX message with send_msg) and allow the worker to flush the buffers and close the socket. The problem with the first is that you increase the amount of wake-ups your event loop needs to do by 2x (since it needs to handle data in and out for the worker) which could increase your handling latency. This is a no-go in some applications. The problem with the second model is you lose the ability to have connections live longer their request (otherwise they are restricted to that workers affinity which will not be evenly balanced).
ODBC is an option (or SQL Relay) but it adds an increased latency without providing the ability to yield on the resource being ready. For example, even with the models I described above to reduce over subscription of resources, your container may not be correct in its assumption that the connection has little contention. So you effectively add latency and reduce correctness in your ability to run efficiently.
It seems like you're stuck on the idioms of Apache and 'real web-servers' (as you put it).
You're persistently parochial in how you're approaching these things.
Remove this problem by having the worker process pass the file descriptor back to the master process as an open (doesn't need to be accept()'ed) socket to be reprocessed or closed as necessary.