Go vs. Python for a simple web server
blog.kowalczyk.info
blog.kowalczyk.info
Yet another example of where I'd love to see some numbers.
kragen@inexorable:~$ python
Python 2.6.6 (r266:84292, Sep 15 2010, 15:52:39)
[GCC 4.4.5] on linux2
Type "help", "copyright", "credits" or "license" for more information.
>>> import os, commands
>>> print commands.getoutput('ps u %s' % os.getpid())
USER PID %CPU %MEM VSZ RSS TTY STAT START TIME COMMAND
kragen 20823 0.3 0.1 9636 3780 pts/4 S+ 14:00 0:00 python
>>> x = [{} for ii in range(1000*1000)]
>>> print commands.getoutput('ps u %s' % os.getpid())
USER PID %CPU %MEM VSZ RSS TTY STAT START TIME COMMAND
kragen 20823 10.8 7.7 168176 159756 pts/4 S+ 14:00 0:02 python
>>> (159756-3780)*1024/(1000.0*1000)
159.719424
So the baseline is about 160 bytes per dict. Objects, by default, are built on dicts. Objects with __slots__ can be a little more efficient.Plus, whatever the results were then, Go's compiler and runtime has improved in the past year, so the data would no longer be indicative of current performance. And the compiler is still being improved.
It seems these days that there is a confusion between the web server (often a reverse-proxy), the web application and the thing in the middle (gateway/pipe/container/app server). While I would certainly consider Go for the implementation of a reverse-proxy or an app server, it isn't clear that Go is a good choice for the web application itself, especially when you consider that most of the time is spent in the DB and in the cache...
Anway, the article mostly talked about Go and not about Python...
Go has an interesting combination of lightweight threading and a high quality performant HTTP server, so potentially it is good for implementing the http serving and application layers.
Last time I created a simple web service with Go it leaked like a hell.
Other than the issue with the 32bit compilers, I don't remember hearing of any leaks in Go.
This statement needs clarity. If the handler is CPU-intensive, the design of the server (threaded vs. AIO) is irrelevant. Both will perform equally, the only difference being that on an SMP system, the OS could schedule the thread on a different processor. This is why the typical deployment model for AIO servers is to start one server per processor and route incoming requests to a proxy to balance the requests among the available servers.
On the other hand, if the handler is dependent on slow network services (e.g. database or downstream API service), the AIO-based server will easily be able to handle additional incoming requests while waiting for the other tasks to complete, subject to memory. This is the power of non-blocking I/O function calls.
There's a bit more to it than that. The OS will frequently block and preempt the long-running thread even on a single CPU system, allowing other, shorter requests to be processed in the mean time.