> Implement PEP 554
> PEP 554 - Multiple Interpreters in the Stdlib
That's going to be fun. Why fight the GIL when multithreading, when you can just get around it with more interpreters?
> Implement PEP 554
> PEP 554 - Multiple Interpreters in the Stdlib
That's going to be fun. Why fight the GIL when multithreading, when you can just get around it with more interpreters?
interp = interpreters.create()
interp.run(tw.dedent("""
import some_lib
import an_expensive_module
some_lib.set_up()
"""))
wait_for_request()
interp.run(tw.dedent("""
some_lib.handle_request()
"""))
I'm actually shocked this is even being contemplated. We've regressed to evaling? import worker
worker.run()
The PEP explicitly mentions this, and that something like subinterpreter.run(func, ...) could be considered in the future: https://peps.python.org/pep-0554/#interpreter-callThe `interpreters` API is just the starting point. Compare it with `subprocess`, not with `multiprocessing`. Once subinterpreters are useful, people will build higher-level APIs for them.
Multiple interpreters with their own GIL keep all of that existing code working without any changes, and mean we can run a Python program on more than one CPU at the same time.
So you are just transforming the problem into a data sharing problem between interpreters, which requires careful thought on both the language side for abstractions, and the consumer side to use right.
It also makes the tooling and verification much harder in practice - for example, you aren't reasoning about deadlocks in a single process anymore, but both within a single process and across any processes it communicates with.
At an abstract level, they are transformable into each other. At a pragmatic level, well, there is a good reason you mostly see tooling for single-process multi-threaded programs :)
Absolutely, but is also the easiest to shoot yourself in the foot with. Trade-offs! I'm biased though, I'm a big fan of deep-copy channels (which for small shallow objects is still fast), though not having the option at all for shared memory here will be a bit of a pain for certain things of course.
But a weird "global state" (really more a global property) is the semantics between concurrent pieces of code and the expectations about things like setting variables, possibly interleavings etc.
The nice part of different interpreters isn't just getting around the gil, and maintaining similar isolation, but it's almost like a Terms and Services agreement: I opened this can of worms and it's my responsibility to learn what the consequences are.
Well, it depends on how it’s implemented.
If “made thread safe” means constantly grabbing locks around large blocks of data then the end result is concurrency (hopefully!) but not parallelism. Meaning you might only have one thread active at a time in practice.
Wrapping the universe in a mutex is thread safe. But it’s not a good solution.
Associate some shared memory with each subinterpreter (the same array or map)
You could have a rule that the refcount must be 1 when sending an object between subinterpreters.
In other words, you cannot use an object that was .send() to another subinterpreter.
Then you invalidate the reference in that subinterpreter when it calls send to the other subinterpreter which is transferred by assignment.
Can transfer any amount of data with zero copies.
Couldn't you separate the storage of the refcounts from the objects and use a map to get at them?
As for the identities between types being different.
To create an subinterpreter that can marshall between subinterpreters without copying the data structures requires a different data structure that is safe to use from any interpreter. We need to decouple the book keeping data structures from the underlying data.
We can regenerate book keeping data structures during a .send or .receive
Maintaining identity equivalence is an interesting problem. I think it's a more fundamental problem that probably has mathematical solutions.
If we think of objects as being a vector in mathematical space. We have pointers between vectors in this memory space.
For a data structure to be position independent. We need some way of intending references to be global. But we don't want to introduce a layer of indirection on the reference of object relationships. That would be slower. Could use an atomic counter to ensure that identifiers are globally unique.
Don't want to serialize the access to global types.
It sounds to me it is a many-to-many to many-to-many problem. Which even databases don't solve very well.
In other words, the code for a function is hashed and that is its identity that never changes while the program is running.
If we use the same approach with Python, each object could have a hash that corresponds to the code only, instead of the data. This is the objects identity even when added to the book keeping data of another subinterpreter.
This requires decoupling of book keeping information from actual object storage. But replaces pointers with lookups which could be inlined to pointer arithmetic.
Also, Go doesn't really solve the problem - sure, it has channels, but it still allows for mutable shared state, and unlike Rust, it doesn't make it hard to use.