But python is easy and fun to write code in, and "developer time is expensive, servers are cheap," so there are a lot of python servers which could benefit from a lack of GIL. Never mind that it's fast to write, but difficult to maintain since it is a dynamically typed language and one typo creates runtime errors that any statically typed language will catch at compile time. Or that it is slow. Or that a python process never really releases memory back to the system, just within itself, so the process slowly grows over the course of a few weeks. Or that the Twisted framework you're using for cooperative multitasking because of the GIL is really easy to block on a database query by accident, leading to uncooperative multitasking (= large lags), resulting to forced server restarts, and loss of players (= loss of revenue). So yeah, "developer time is cheap" but it's sort of an expensive cheap. I came to the conclusion that python is unsuitable for servers, but until Go came out, there wasn't a realistic alternative, since C++ and Java are too heavyweight, and Ruby suffers from similar problems (don't know about a GIL).
But with Twisted, and now with async def, it is straight-forward to write performant cooperative multitasking code. For some definition of "straight forward" -- I have been doing coopertive real time code since before Python existed, so I've learned to think that way. But it really is worth the effort to climb the learning curve on Twisted, and it really does make cooperative tasking painless.
The problem with the GIL and Python semantics is that collections imply a zillion fine-grained locks everywhere. The time spent aquiring/releasing locks, surprisingly, is not the issue. It is that every lock requires that every CPU cache has to sync on the lock. All that locking makes the process cache-invalidate bound.
1. For CPU-bound applications which use threading, performance is severely degraded due to only one thread being executed at a time regardless of number of cores/CPUs.
2. For all other threaded applications, performance can be slightly but measurably degraded by obtaining and releasing the GIL.
If you are an I/O-bound threaded application, the GIL is not something you're likely to really be bothered by, and probably will get lost among the noise of all the other things that will affect performance.
As a result, Python for "servers" -- assuming you mean network service daemons which are almost always I/O-bound -- is perfectly fine, including "at scale", as demonstrated by a number of large sites and services which get along just fine on Python.
You can also avoid the GIL entirely by using a parallelism construct other than threading.
And for completeness' sake, the oversimplified history of the GIL:
Python is an older language than people tend to realize. It predates Java. And Python came in part out of the Unix scripting-language tradition, where Unix approaches -- such as forking additional processes -- were the typical way to do things. Then along came Java, which had the limitation of being designed originally to run on set-top TV boxes which didn't have true multitasking. So Java imposed threading as the way to do multitasking, and Java became very popular.
Thus, Python was pressured to develop a story on threading. But since Python had been built in the Unix multi-process tradition, it wasn't implemented in a way that was friendly to threading. The GIL was the compromise that allowed Python to have threading: it would only seriously affect threaded CPU-bound applications on multi-CPU or multi-core hardware (since I/O-bound applications aren't affected nearly as much, and single-core, single-CPU hardware is only physically capable of running one thread at a time anyway).
Fast forward a couple decades and now we all have multi-CPU and/or multi-core computers, including literally carrying them in our pockets, and CPU-bound applications are more common. In retrospect, the GIL can look like the wrong tradeoff to make, which is why people want to get rid of it, but at the time it was quite reasonable.
Right, agreed. I can imagine some of the frustration you might experience using CPython for high throughput systems: kind of like NodeJS without the benefits of a standard library written with async/non-blocking I/O in mind.
A bit curious about a few things you mention here, though:
> Or that a python process never really releases memory back to the system, just within itself, so the process slowly grows over the course of a few weeks.
I'm not sure this is true in general, is it? Can you elaborate? It's been a while since I've dug around in Python innards, but if Py_DECREF(x) leads to a refcount of zero IIRC free(x) is ultimately called -- albeit in an indirect manner via a layer or six of tp_dealloc calls and tp_free. :) I suppose calling free(x) may only return the memory associated with x to (g)libc's free list and not necessarily back to the OS [0]. No different to C/C++ in that regard, I guess.
> I came to the conclusion that python is unsuitable for servers, but until Go came out, there wasn't a realistic alternative, since C++ and Java are too heavyweight, and Ruby suffers from similar problems (don't know about a GIL).
"Too heavyweight" in that they're relatively difficult to write in comparison? Maybe true of Java-the-language, but the JVM itself is an absolute workhorse when it comes to high performance. Plenty of languages to choose from there, typically without a GIL. Jython, for example, has no GIL [1].
And yep, Ruby/MRI has a GIL (but JRuby does not).
[0] https://www.gnu.org/software/libc/manual/html_node/Freeing-a... [1] https://stackoverflow.com/questions/1120354/does-jython-have...
This just isn't true anymore. Create a list with range(1000000) then delete the list and you can see the memory freed.
This distinction is only important if you do not have tests.
You need more tests if you are using a language that is weakly typed (e.g. C, js) and fewer if you are using a language that is largely type-safe (e.g. python) and even fewer if it is very, very type-safe (e.g. rust/haskell), but the compile-time/runtime distinction doesn't change the number of tests you need to achieve a requisite level of code quality - it only changes whether a behavioral test you should always be writing or a compiler picks up type errors.
Linting also catches variable misspellings.
If you're willing to go slightly outside the mainstream, there's stuff like erlang and haskell with kickass runtimes. And haskell, at least, is pretty strongly typed.
For that matter you should look into Theano, Tensorflow and Numba to bring Python code to the GPU (no GIL there). Or use Dask to scale to multiple cores or nodes.
Right - and one reason that it wouldn't be your first choice would be the GIL, so let's remove the GIL, and then the other barriers.
I don't think this is true. GIL means Global Interpreter Lock - it isn't held when you are in native code, which is how NumPy et al actually work - Python to marshal the data, then do the heavy lifting under the hood in C, FORTRAN and ASM, that the end user never needs to see. So this is a non-problem. As the above comment says, the problem is if you want to write a server that handles a lot of concurrency and shared state.
If you write a server and it uses threading as the concurrency model and its workload is primarily CPU-bound, the GIL will be a problem for you.
Take away either of those conditions -- use a model other than threading, or have an I/O-bound workload -- then the normal background overhead of the GIL is not something you'll notice.
People who deploy Python applications as network daemons tend to use pools of worker processes (not threads) and have I/O-bound workloads, which is why the assertion that Python is somehow "bad for servers" is not one you'll typically hear from people who actually deploy Python on servers. Much like the comment you're referring to, which is from someone who seems to primarily have type-system complaints about Python (i.e., they wouldn't touch a dynamically-typed language to begin with) and is throwing in "oh and the GIL probably makes it unsuitable for a server" as an additional (but factually incorrect) reason not to use Python.
Meanwhile, if you are writing a server which uses threading and lots of shared state which will be modified concurrently, you're in for a world of pain no matter what language you choose. There's a reason why there are multiple up-and-coming or even moderately-popular languages now which have as a design feature the inability to do such a thing.
multiple up-and-coming or even moderately-popular languages now which have as a design feature the inability to do such a thing
STM in Haskell seems quite promising, but the approach of just outsourcing that problem to Redis gets you pretty far.