Has the Python GIL Been Slain? Subinterpreters in Python 3.8
hackernoon.com
hackernoon.com
I get that the GIL is a very hard problem to solve, but this solution is so inelegant in my eyes that python would be better off without it. I'd feel better if this was a hidden implementation detail that coukd be improved transparently. Just my two cents.
So besides the unproven possibility of removing the GIL, subinterpreters are the best way forward, better than threads or the multiprocessing package.
That's not accurate.
> they have their place but aren't related to parallelisation
You can parallelize all sorts of things with Python threads--just not some things you'd expect to be able to parallelize, due to the GIL. Waiting on or buffering I/O, calling out to compiled code, doing cryptographic operations--all of those can be parallelized, as (in many cases) they entail releasing the GIL.
> But true multiprocessing is ugly when you have hundreds of cores
Why?
> There is no standard UI convention on most OSes to group those processes per app, in terms of signals or stats or whatever.
I have no idea what this means. What do UIs have to do with process groups? Do you know how many processes your Chrome instance is running on the operating system? There are very solid conventions regarding process management, at least on Unix-ish systems: process groups and parent-child relationships are well established and well understood, as is their relationship with signals and signal handling.
Green threads also allow for parallelism, if they're scheduled onto more than one executor.
I think threads are a bad abstraction for doing parallelism personally, though. Programs designed to run in parallel should be deterministic, unless they need concurrency for some other reason. I think trying to shoehorn parallel programming into Python's threads isn't necessarily the best approach.
As far as I'm concerned, if I'm using threads and they happen to get scheduled on to multiple cores, then that's a nice optimization, but isn't necessary for what I use threads for.
async/await are not only for asyncio.
But if you think you can write correct asyncio applications of any significant complexity, that's great. Doesn't change the fact that a lot of highly experienced developers struggle to do so.
To be fair, I've never written an asyncio application of any significant complexity (at least compared to the async JS application I've written). But nothing jumps out at me as horrific with asyncio. The worst thing is the split between futures and coroutines, but once you realize that there's no real need to use futures any more with asyncio (although you can integrate futures-based libraries if needed), you're good to go.
But fair criticism, I've only written small asyncio applications. The only sizable async-Python application I've made was made using Tornado and ZeroMQ async APIs(although even that eventually ran on the asyncio event loop and used an asyncio-based library).
Python already has exactly that, and has had that for ages.
>Instead, you should learn to use this weird contraption that is neither multiprocessing nor intuitive multithreading and comes with a cumbersome interface.
It also comes with performance improvements over multiprocess, so there's that.
Besides the "cumbersome interface" is irrelevant, as it would be easy to wrap and forget about it, the same way nobody really uses urllib directly.
And the point is that this should be fixed, instead of adding yet another clunky way of doing the same thing.
> Besides the "cumbersome interface" is irrelevant, as it would be easy to wrap and forget about it, the same way nobody really uses urllib directly.
Then why bother having it, if the only reason for it to exist is so that it can be wrapped? That's just stacking layers of turd.
This is assembly. Don’t want to wrap it yourself, fine, don’t complain that it exists.
Yes, but you couldn't do more fine-grained than requests things, unless you did them from excruciatingly lower level (e.g. sockets).
Well, unless we have a (1) capable (2) volunteer the point is moot.
Those who actually work on Python development accessed both scenarios and found adding the "yet another clunky way" if not optimal, a more feasible, easier and better use of their limited time.
>Then why bother having it, if the only reason for it to exist is so that it can be wrapped?
Because lower level primitives can be used in many useful ways than some constraining higher level construct, but the latter is still nice to have.
They are isolating the GIL into Guilds there, which are containers for language threads sharing the same GIL. They are providing two primitives for communication between threads in different guilds. Send, for immutable data (zero copy) and move, for mutable data (copy). They remove the need for the boiler plate code for marshalling and unmarshalling. However I bet that there will be some library to hide that code in Python too.
[1] http://www.atdot.net/%7Eko1/activities/2018_RubyElixirConfTa...
That might not be a bad idea because I am worried `move` will end up being problematic in ruby, but time will tell.
Such approach would cover many, if not most user cases for multithreading.
gsub! is a method of String that mutates the object (vs gsub which returns a new string)
freeze is a method that makes immutable the object it's called upon.
2.3.0 :001 > def replace_a_with_b(s)
2.3.0 :001?> s.gsub!("a", "b")
2.3.0 :001?> end
=> :replace_a_with_b
2.3.0 :002 > replace_a_with_b("abc")
=> "bbc"
2.3.0 :003 > replace_a_with_b("abc".freeze)
RuntimeError: can't modify frozen String
from (irb):20:in `gsub!'
from (irb):20:in `replace_a_with_b'
from (irb):23
from /home/me/.rvm/rubies/ruby-2.3.0/bin/irb:11:in `<main>'Hm. Who let the Rust folks in?
It’s much easier to map existing multi process code onto shared nothing sub interpreters.
I proposed something similar for Python 9 years ago.[1] Guido didn't like it.
Objects would be either thread-local, shared and locked, or immutable. Thread-local objects must be totally inaccessible from other threads, and not leakable across thread boundaries, for memory safety. (Python has "thread local" objects now, but it's just naming, and not airtight against leaks. You can assign a thread-local object to a global variable.) Shared and locked objects lock when you enter, unlock when you leave. Objects are thread-local by default, so single-thread programs work as before.
Minimize shared and locked, while using thread-local or immutable objects as much as possible. Locking is needed only for shared and locked objects.
This is almost conventional wisdom today, but 9 years ago, it was too radical.
Retrofitting concurrency is never pretty. But we have to. Individual CPUs are about the same speed per thread that they were a decade ago.
[1] http://animats.com/papers/languages/pythonconcurrency.html
Dangerous advice. Whether this is true or not depends on lots of things such as how many and which operations you're doing on those variables.
Sure, CPython might do lots of simple operations atomically, but this is not enough to avoid the need for all locks. Threads can still interleave their execution in many ways.
See also: https://blog.qqrs.us/blog/2016/05/01/which-python-operations...
Python's performance, in general, is a crappy[1] and is beaten even by PHP these days. All the people that suggest relying on multiprocessing probably haven't done anything that's CPU and Memory intensive because if you have a code that operates on a "world-state" each new process will have to copy that from a parent. If the state takes ~10GB each process will multiply that.
Others keep suggesting Cython. Well, guess what? If I am required to use another programming language to use threads, I might as well go with Go/Rust/Java instead and save the trouble of dabbling with two languages.
So where does that leave (pure-)Python? It can only be used in I/O bound applications where the performance of the VM itself doesn't matter. So it's basically only used by web/desktop applications that CRUD the databases.
It's really amazing that the machine learning community has managed to hack around that with C-based libraries like SciPy and NumPy. However, my suggestion would be to drop GIL and copy the whatever model has been working for Go/Java/C#. If you can't drop GIL because some esoteric features depend on that, then drop them as well.
[1] https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
Those recommending to use multiprocessing have probably never been in that bitter spot when serializing something and computing something takes exactly the same time.
Also forking didn't really work until Python 3.6.
And there's also Kotlin.
In POSIX there is such thing as copy-on-write memory during forks.. So if that state is mostly read-only, additional memory required by each slave should be minimal.
However it isn't required by POSIX for compliance.
http://pubs.opengroup.org/onlinepubs/9699919799/functions/fo...
[1] vs fully shared state of C-like, .NET, JVM, etc, etc. Rust’s no-shared-mutable state model allows it to do some fun stuff but python (and JS) don’t really have a strong concept of mutable vs immutable, let alone ownership so I don’t think it would be applicable?
SharedArrayBuffer is fundamentally the same as python’s C-only sharing - a bunch of raw bytes not tied to the host environment’s type system.
I wonder if they ever fixed the CPickle bug which broke it if you were using CPickle from multiple threads.
There is nothing preventing you from adding this support to an existing dynamic linker and using it for your program, though.
This is supported in ld-linux by using the dlmopen() glibc function and distinct namespaces. Loading the same .so file multiple times is one of the use cases explicitly mentioned in the manual page[1].
[1]: http://man7.org/linux/man-pages/man3/dlmopen.3.html#NOTES
One caveat seems to be:
> The glibc implementation supports a maximum of 16 namespaces.
my_foo = interpreterX.pass_object(my_foo)
(The assignment being required to delete the originating reference from the source interpreter.) The interface would be obligated to check that there are no references that escape to the current interpreter and then my_foo and all referenced objects could be handed off to the other interpreter in whole.I don't have any intuitions for if that would be cheaper than copying or not, and getting it right is certainly more difficult than serialization. (Because of the complexity, it's not worth having if it isn't cheaper.)
It's not a theoretical problem, just a very likely pragmatic observation. (Meaning, we'll see bugs in both the interpreter and the code using this feature.)
Plus if one of the threads crashes, the whole process aborts. (Sure, the interpreter can handle a lot of faults gracefully, but not all.)
I'm not claiming any originality here — the author of the article recognizes this problem and describes it, starting with:
> Because CPython has been implemented with a single interpreter for so long, many parts of the code base use the “Runtime State” instead of the “Interpreter State”, so if PEP554 were to be merged in it’s current form there would still be many issues.
https://en.wikipedia.org/wiki/Betteridge%27s_law_of_headline...
(1) https://perldoc.perl.org/threads.html
[0] https://docs.microsoft.com/en-us/windows/desktop/com/process...
I can think of maybe having network code run in their own process and the UI in another. That way there's no risk of bottle necks slowing down the UI and transfers are likewise protected. If you look at bottle.py it seems that this approach could add A LOT of performance for managing downloads / uploads if it's done right.
Wouldn't just using CLONE_FILES when forking off interpreters solve this problem?
How does this make sense? What's the point of having multiple threads then?
In practice, this doesn’t work particularly well, as you rarely have massively I/O bound things in Python.
Citation needed.
Only being able to run a single thread of "logic" basically makes that impossible, as you need to usually do some computation to figure out what bytes to send/receive, do something with the results, and so on.
So ... No.
The article also makes it clear that each sub-interpreter still has its own GIL, but two sub-intepreters can run at the same time without having to care about each other's GILs.
> Each of these [methods for communicating between subonterpreters] has pro’s and con’s, all of them have an overhead.
IPC-heavy programs with low concurrency may suffer worse under this model than threaded with traditional single GIL. As threads approach infinity, though, anything that scales beats the single GIL model.
It reminds me of how python advertising tricks people into thinking python has real parallelism by talking about ayncio and import multiprocessing; people then actually try those out only to discover the sad state of affairs.
Uh what?
After 10 years, the 2to3 transition is still ongoing, and lots of companies will be dealing with python2 code for a long, long time. Breaking backward compatibility was a terrible decision. Breaking it again now that most libraries and open source projects have finally moved to python3 would be suicide.
Programming languages should take backward compatibility as one of their most important features, because of the hundreds of millions of lines of code already written and that would need to be changed and validated. This is especially true in a dynamic language where you can't even rely on the compiler to point out the issues.
Won't python2 be EOLed sometime in 2020? Is someone expected to pick up security maintaince after that?
For starters:
If you get rid of CPython's atomic operations on lists and other simple objects, lots of Python code would need more locks than it needs now.
If, on the other hand, you wish to keep those operations atomic, without the GIL that means introducing fine-grained locks on specific objects, which would worsen even single-threaded performance. Imagine taking a pthread mutex each time you append an item to every list or store a number in every variable.
https://lwn.net/Articles/754577/
TLDR: simply replacing object usage counters with their atomic versions grinds the interpreter to the halt.
The question in my mind is how the Python design addresses the problems that made the perl equivalent so unusable. It's only had ~20 years of context to do better, so I'm hopeful
"The "interpreter-based threads" provided by Perl are not the fast, lightweight system for multitasking that one might expect or hope for. Threads are implemented in a way that make them easy to misuse. Few people know how to use them correctly or will be able to provide help.
The use of interpreter-based threads in perl is officially discouraged."
Perl5 also had a similar queue based scheme for sharing data across the interpreters: https://perldoc.perl.org/Thread/Queue.html