Faster CPython 3.12 Plan
github.com
github.com
> Python currently has a single global interpreter lock per process, which prevents multi-threaded parallelism. This work, described in PEP 684, is to make all global state thread safe and move to a global interpreter lock (GIL) per sub-interpreter. Additionally, PEP 554 will make it possible to create subinterpreters from Python (currently a C API-only feature), opening up true multi-threaded parallelism.
Very basic question: in a world where a Python program can spin up multiple subinterpreters, each of which can then execute on a separate CPU core (since they don't share a GIL), what will the best mechanisms be for passing data between those subinterpreters?
> Regarding the proposed solution, “channels”, it is a basic, opt-in data sharing mechanism that draws inspiration from pipes, queues, and CSP’s channels.
> As simply described earlier by the API summary, channels have two operations: send and receive. A key characteristic of those operations is that channels transmit data derived from Python objects rather than the objects themselves. When objects are sent, their data is extracted. When the “object” is received in the other interpreter, the data is converted back into an object owned by that interpreter.
> * None
> * bytes
> * str
> * int
> * channels
> Limiting the initial shareable types is a practical matter, reducing the potential complexity of the initial implementation.
That's a really interesting detail - presumably channels can be passed so you can do callbacks ("reply on this other channel").
I wonder why floats aren't on that list? I know they're more complex than ints, but I would expect they would still end up with a relatively simple binary representation.
Callbacks or notifications, yes. I use both patterns quite often.
hallelujah!
Global variables are evil. The fact that sub-interpreters aren’t currently possible in Python is one of my canonical examples of why they’re evil.
To apply to Python The subinterpreters could transfer ownership of the refcounts between subinterpreters as part of an enqueue and dequeue.
I believe the refcount locking approach has scalability problems between threads.
I implemented a multithreaded actor system with work stealing in Java and message passing can get to throughputs of around 50-60 million messages per second without blocking or mutexes. The only lock is not quite a spinlock. I use an algorithm I created but inspired by this whitepaper [1], which is simple but works. It's probably a known algorithm but I'm not sure of the name of it.
I have a multidimensional array of actor inboxes (each actor has multiple buffers for filling by other threads to lower contention to 0) then there is an integer stored for the thread that is trying to read or write to the critical section.
The threads all scan this multidimensional array forwardly and backwardly to see if another thread is in the critical section. If nobody is there, it marks the critical section. It then scans again to see if it is still valid. it's similar to going into a room and scanning the room left and scanning the room right. Surprisingly this leads to thread safely. I wrote a python model checker to verify the algorithm is correct.
Without message generation within threads, it can communicate and sum 1 billion integers in 1 second due to parallelism (it takes 2 seconds to do this with one thread) It takes advantage of the idea that variable assignment can transfer any amount of data in an assignment.
See Actor2.java (1 billion sums a second messages created in advance), Actor2MessageGeneration.java (20 million requests per second, messages created as we go) or Actor2ParallelMessageCreation.java (50-60 million requests per second, with parallel message creation)
There's also a Java multiconsumer multiproducer ringbuffer in this repository [3] too which I ported from Alexander Krizhanovsky [2]
[1]: https://lag.net/papers/content/leftright-extended.pdf
[2]: https://www.linuxjournal.com/content/lock-free-multi-produce...
[3]: https://github.com/samsquire/multiversion-concurrency-contro...
If it's performance, then, since subinterpreters run in the same process, it would be global shared state. You can't use Python objects across subinterpreters, but raw byte arrays will work just fine, provided you do your own locking correctly around all that.
(Brackets my own of course.)
Sharing data in concurrent programs is not trivial, especially in environments where data is mutable. The most trivial answer to the question is “message passing”, as in the SmallTalk notion of OOP or the Erlang/OTP Actor Model. Some solutions look much more like working with a database (Software Transactional Memory). Some models that seem entirely designed for a different problem space are also compelling (various state models common in UI and games like reactivity and Entity Component systems).
Shareable data is basically immutable data + classes/modules, and unshareable data can be transmitted via push (send+receive) or pull (yield+take). Transmission implies either deep copying (which "forks" the instances) or moving with ownership change (sender then loses access)
See here for details: https://docs.ruby-lang.org/en/master/ractor_md.html
Hah, I wonder what else Python 4 could have in it.
The Python 2 to 3 migration was hard enough and there were certain challenges along the way (mostly package availability and syntax changes, though the same is happening with new Vue versions) but it seems that in regards to most metrics Python 3 was indeed an improvement, apart from the startup time.
Hardly.
The majority of Python users don't use the typing module.
They might want it in (and helped add it), but it's not what made them use Python.
And we've even been to the same circle before: a lot of programming the 60s and 70s was untyped, then they switched to typed C++, Delphi and then Java.
I see some mild sense in this argument given how TypeScript has taken off and dispersed into audiences who wouldn't ordinarily be interested in such a thing. I'm not sure it works in the Python world though, since Python's latter day upward trajectory is probably more oriented around heavy use in education, science, ML, PyTorch, et al?
It was not at all obvious from your post.
The "GIL removal" umbrella proposal was two-fold. It included (a) removing the GIL, (b) several optimizations to handle some issues with GIL being removed and offset the GIL removal overhead (due to more frequent lock checks, etc).
The GIL-removal assisting changes and optimizations were merged, but the GIL removal was not.
> Implement PEP 554
> PEP 554 - Multiple Interpreters in the Stdlib
That's going to be fun. Why fight the GIL when multithreading, when you can just get around it with more interpreters?
But a weird "global state" (really more a global property) is the semantics between concurrent pieces of code and the expectations about things like setting variables, possibly interleavings etc.
The nice part of different interpreters isn't just getting around the gil, and maintaining similar isolation, but it's almost like a Terms and Services agreement: I opened this can of worms and it's my responsibility to learn what the consequences are.
Well, it depends on how it’s implemented.
If “made thread safe” means constantly grabbing locks around large blocks of data then the end result is concurrency (hopefully!) but not parallelism. Meaning you might only have one thread active at a time in practice.
Wrapping the universe in a mutex is thread safe. But it’s not a good solution.
Associate some shared memory with each subinterpreter (the same array or map)
You could have a rule that the refcount must be 1 when sending an object between subinterpreters.
In other words, you cannot use an object that was .send() to another subinterpreter.
Then you invalidate the reference in that subinterpreter when it calls send to the other subinterpreter which is transferred by assignment.
Can transfer any amount of data with zero copies.
Couldn't you separate the storage of the refcounts from the objects and use a map to get at them?
As for the identities between types being different.
To create an subinterpreter that can marshall between subinterpreters without copying the data structures requires a different data structure that is safe to use from any interpreter. We need to decouple the book keeping data structures from the underlying data.
We can regenerate book keeping data structures during a .send or .receive
Maintaining identity equivalence is an interesting problem. I think it's a more fundamental problem that probably has mathematical solutions.
If we think of objects as being a vector in mathematical space. We have pointers between vectors in this memory space.
For a data structure to be position independent. We need some way of intending references to be global. But we don't want to introduce a layer of indirection on the reference of object relationships. That would be slower. Could use an atomic counter to ensure that identifiers are globally unique.
Don't want to serialize the access to global types.
It sounds to me it is a many-to-many to many-to-many problem. Which even databases don't solve very well.
In other words, the code for a function is hashed and that is its identity that never changes while the program is running.
If we use the same approach with Python, each object could have a hash that corresponds to the code only, instead of the data. This is the objects identity even when added to the book keeping data of another subinterpreter.
This requires decoupling of book keeping information from actual object storage. But replaces pointers with lookups which could be inlined to pointer arithmetic.
Also, Go doesn't really solve the problem - sure, it has channels, but it still allows for mutable shared state, and unlike Rust, it doesn't make it hard to use.
Multiple interpreters with their own GIL keep all of that existing code working without any changes, and mean we can run a Python program on more than one CPU at the same time.
So you are just transforming the problem into a data sharing problem between interpreters, which requires careful thought on both the language side for abstractions, and the consumer side to use right.
It also makes the tooling and verification much harder in practice - for example, you aren't reasoning about deadlocks in a single process anymore, but both within a single process and across any processes it communicates with.
At an abstract level, they are transformable into each other. At a pragmatic level, well, there is a good reason you mostly see tooling for single-process multi-threaded programs :)
Absolutely, but is also the easiest to shoot yourself in the foot with. Trade-offs! I'm biased though, I'm a big fan of deep-copy channels (which for small shallow objects is still fast), though not having the option at all for shared memory here will be a bit of a pain for certain things of course.
interp = interpreters.create()
interp.run(tw.dedent("""
import some_lib
import an_expensive_module
some_lib.set_up()
"""))
wait_for_request()
interp.run(tw.dedent("""
some_lib.handle_request()
"""))
I'm actually shocked this is even being contemplated. We've regressed to evaling? import worker
worker.run()
The PEP explicitly mentions this, and that something like subinterpreter.run(func, ...) could be considered in the future: https://peps.python.org/pep-0554/#interpreter-callThe `interpreters` API is just the starting point. Compare it with `subprocess`, not with `multiprocessing`. Once subinterpreters are useful, people will build higher-level APIs for them.
To paraphrase Adam Savage from his excellent book, Every Tool's a Hammer, lists [of lists] are very powerful way to tame the inherit complexity of any project worth doing.
That's a pretty generic statement and likely true only in a very specific frame.
It rubs people the wrong way but I always call out blanket statements. Generally languages get faster with each version and theres a lot of numbers thrown around, it doesn't mean your apps will get anywhere near that boost.
If you're lucky that one loop that concats strings got a few ms shaven off while that ORM youre using continues to grind the whole thing down.
This will break many modules. Basically any that use static variables, which is done pretty much everywhere.
I believe you can do that with `dlmopen` in separate link maps. I have worked with multiple completely isolated Python interpreters in the same process that do not share a GIL using that approach.
There are a few cases where `dlmopen` has issues, for example, some libraries are written with the assumption that there will only be one of them in the process (their use of globals/thread local variables etc.) which may result in conflicts across namespaces.
Specifically, `libpthread` has one such issue [1] where `pthread_key_create` will create duplicate keys in separate namespaces. But these keys are later used to index into `THREAD_SELF->specific_1stblock` which is shared between all namespaces, which can cause all sorts of weird issues.
There is a (relatively old, unmerged) patch to glibc where you can specify some libraries to be shared across namespaces [2].
[1]: https://sourceware.org/bugzilla/show_bug.cgi?id=24776#c13
[2]: https://patchwork.ozlabs.org/project/glibc/patch/20211010163...
It's going to be a bit of a chicken and egg problem, core Python will need to prove it's worthwhile for extension devs to implement, core Python will struggle without support from extension devs. We shall see.
If extensions don't support it it means you just can't use that extension when trying to run multiple interpreters in the same process. Let's see if there's even a good use case for running multiple interpreters in the same process outside of embedded programming, it's not 100% clear yet.
Gotta admit. That sounds pretty official
I have no idea of the pay but based on my research of salaries for getting a job this year I would wildly speculate high 6 digits to low 7 digits.
I'll check back in a decade.
I mean 2.x is still in the wild and some companies provide support for it, still!
Python on the other hand has one real implementation, and the ecosystem has become extremely intertwined with that implementation. Implementing a new Python interpreter is great, but it rarely gains traction because most of the ecosystem is so reliant on CPython specific modules that don’t work in the new interpreter so they never really get off the ground.
[1]: https://github.com/python/cpython/blob/4b81139aac3fa11779f6e...
Yes, but doesn't the GIL currently serialise things?
Note that SQLite limits access to 1 writer at a time though.
A year or two ago I read up on the various efforts to make a fast, more parallel CPython, and one of the core underlying problems seemed to be the use of machine threads, resulting in a very high locking load as the large (potentially unlimited) number of threads attempted to defend against each other.
Letting an operating system run random fragments of your code at random times is very much a self-inflicted wound, so I was wondering if the python community has any plans to not do that any more ?
Removing fanatic compiler abuse is always a good thing. That said, I saw some assembler macro abuse (some assemblers out there have extremely powerful and complex macro pre-processors), then the hard part would be not to abuse the macro pre-processor of the assembler.
I know it is not to make python actually "faster", but to have a python implementation which does not require those grotesquely and absurdely massive compilers, then the SDK stack would be way more reasonable from a technical cost stand point.