Python GIL removal question
mail.python.org
mail.python.org
I dont' understand. Isn't this going to happen if you have multiple threads running even if the GIL is blocking them from running? I'm not a hardware expert, but I'm not sure how constant locking would prevent cache synchronization just because they weren't truly running in parallel.
I am fairly certain that constant synchronization(lock) because of the GIL would negatively impact cache performance, especially since well designed multithreaded applications avoid locking for as long as possible.
EDIT: Also, "stop whatever they are doing to synchronize the dirty cache lines with RAM," is not a very good way to describe what is going on, often times you don't have to hit RAM at all, the caches just synchronize between each other. It is still pretty bad for performance though.
Ah. Makes sense.
>just synchronize between each other
Yes, but that's bad because that cache line is 'stuck' for all processors while the synchronization is occurring, if I'm not mistaken...
I don't think it would. If there is a GIL (Global Interpreter Lock) only one thread of the process can be scheduled to run at any time. As the poster (Sturla) says, Python threads are native OS threads so they should be scheduled by the OS kernel (right?). A good scheduler would use affinity scheduling and schedule all threads of the Python program on the same processor/core every time to get benefits from cached data and code. I believe modern kernels (Linux, MacOS, Solaris, probably Windows as well) use this kind of affinity scheduling, so if we're lucky the Python process gets scheduled on the same processor every time and there will be no need any cache synchronization.
> I'm not a hardware expert, but I'm not sure how constant locking would prevent cache synchronization just because they weren't truly running in parallel.
I'm not sure if you misunderstood the mail. The constant locking would only be used if they were running in parallel.
Anyway, if you have a GIL you don't need that kind of locking described in the mail. You only need to do explicit locking on shared data structures when you read or update the contents of those data structures. If you have reference counting, threads that run in parallel and no GIL you would have to lock even if you are just assigning a reference to such a data structure to a new variable. If you have a GIL you are certain that only one thread at a time are updating the reference count. That is indeed what the GIL is, one coarse lock for all data (and the interpreter) instead of fine grained locks for every data structure.
(I don't know Python very well, I just answer from general knowledge of computer architecture and language implementation. But I've read about the Python GIL several times, since it's the most discussed GIL of any language.)
>I'm not sure if you misunderstood the mail. The constant locking would only be used if they were running in parallel.
No, I'm saying that the GIL is constant locking. You still have two threads being run concurrently on (possibly) two separate cores accessing the same cache lines. They just cannot actually run in parallel. I have no idea how the GIL time slices between the two threads, so what i'm saying is completely possible.
However, below my original post meastham correctly pointed out the GIL does prevent cache thrashing where updates to shared memory might go back and fourth multiple times unnecessarily. So it's not as bad as I was imagining.
Ah, I see now what you meant with constant locking. I interpreted your words as "constant" as in happening all the time as would be the case with fine grained locks instead of one long-lasting, global lock.
Java, .NET, various Lisps, and I'm sure other systems that I don't know off the top of my head have solved the problem of true threading.
http://morepypy.blogspot.com/2011/06/global-interpreter-lock...
Here are two examples of products that ship python. You can't just "swap out the interpreter".
http://www.thefoundry.co.uk/products/nuke/ http://usa.autodesk.com/maya/
According to this paper [1]: "Thus, in all cases, the single global lock semantics seem fundamentally compatible with both lock-based and transactional memory implementations."
[1] http://www.usenix.org/event/hotpar09/tech/full_papers/boehm/...
Removing the GIL might be useful for a faster implementation with JIT compiler though.
As long as you are not creating or destroying the Python objects you expect to give back to the Python side of your program, you don't have to care much about it.
Perhaps, but thanks to numpy, scipy and a host of other amazing libraries it still ties with matlab as the go to language for number crunching among everyone I know who crunches numbers for a living.
Ah, that old line again.
Translation: "We really don't like to even think about changing this crappy design that we started with in the first place, because we can just explain ourselves out of it by coming up with suitable language goals that don't actually require concurrent access to the interpreter. Not accessing the interpreter concurrently is one of our language goals because you can do everything else. So, if you think you still need to get rid of GIL then you're just a bad programmer and your programs are badly written because hey, we just defined the universe you're playing in."
Not that the case isn't well argued, but to claim that GIL isn't a fundamental limitation and a bad thing is silly.
A few years from now it will be like, "Oh, yeah.. that".
That's a quote. What more do you want?
The question is not and never has been "Does the GIL have undesirable characteristics?" It has always been "can someone produce an implementation that is missing the GIL and actually better, while meeting all the needs CPython has?" So far, the answer is no, despite rather a lot of smart people trying.
(Also note that many people have succeeded by dropping the second clause. Many non-CPython Python implementations don't have a GIL. But they aren't CPython, which in particular means that extensions written for CPython don't work in them, which is really the key thing that distinguishes CPython from just generic "Python".)
With PyPy the performance will get better, and they also have a GC, so that hinder is removed. I don't really know if PyPy has a GIL, I would guess that they don't.
PyPy still has GIL:
http://codespeak.net/pypy/dist/pypy/doc/faq.html#does-pypy-h...
For more information about it, check these sources:
Official PyPy Status Blog - http://morepypy.blogspot.com/2008/05/threads-and-gcs.html
Thinking about the GIL (read the whole thread, interesting stuff) - http://mail.python.org/pipermail/pypy-dev/2011-March/006991....
Is it really necessary for the GC to be re-entrant to run the interpreter in parallel? Couldn't you have the interpreter running in parallel and then when there is a need to run the GC you have a global GC lock that prevents all threads from running - a stop the world GC. The application runs for a longer time than the GC, right? So it would be a win and a step in the right direction? I believe the early Java mark and sweep GC was like that, and then later Sun developed several different kinds of concurrent and parallel GCs.
> Official PyPy Status Blog
Oh I read that every time they write something. :) But I started reading it in late 2010 and I haven't gone back to the archives, I guess it's time to do that. Thanks for the links.
Deleted comment
Removing GIL is massive work and will make the interpreter complex. Meanwhile, you have gevent, multiprocessing, c extensions... to work out the limitations.
Running Java threads on a single processor machines makes it not truly multi-threaded then? Python has multi-threading - due to GIL, only one of them execute at a time, which isn't very different from running multiple threads on a single processor machine. It facilitates concurrency, not parallelism.
> claiming that multi-process is the way to do parallel computation across the board is not a sane argument at all.
If you have n processing units, anything greater than n isn't parallel. Spawning 100 threads in a JVM doesn't give you 100 parallel workers(assuming JVM mapped those 100 JVM threads to 100 system threads).
Multi processes make perfect sense for parallel jobs. They do fine with nothing shared and message passing. They are problematic when the jobs need resource sharing.
> having to use proxy servers for database access to minimize connections across all the python process instances
That doesn't sound like pain. Some systems intentionally have kept db connection pooling outside the application server. Application server talks to the manager and manager delegates to the database.
> Using python for anything more serious, like a message queuing system for example is even more prohibitive.
Celery works great. Thank you.
For custom needs, there is gearman, then there is zeromq...and they are not written in Python, and I don't care, and that works for me.
> Meanwhile in the JVM world..
Then why not stick to the JVM world rather than crib and whine in Python world.
Yea it may be a lot of work to create a new GC implementation and change the threading model, but if you want the language to progress that's the way forward.
Yes.
> He and a number of others don't think it's one worth solving. Really.
Check your claims.
http://www.artima.com/weblogs/viewpost.jsp?thread=214235 http://docs.python.org/faq/library#can-t-we-get-rid-of-the-g...
The share-nothing multi-process approach Python encourages feels good enough for me.