A long time ago IronPython was released which showed that you could build a high performance Python interpreter that was GIL-free. It had thread-safe containers (so multiple threads could work against the same lists, dictionaries, etc) and in some cases was faster than CPython (it was implemented using CLR and .net)
When I saw IronPython I was immediately convinced that CPython should be the same way- already low-cost dual and quad core Intel machines were becoming available and it seemed like core counts were going to increase faster than clock rates. I figured that a small hit to serial performance would be more than acceptable if people could write multithreaded systems in Python, in much the same way as I wrote multithreaded C++.
Over time after watching nogil not going anywhere in CPython (the python leadership didn't want to do nogil), with concomitant speedups in the single-processor implementation, along with the increasing use of C++ code that releases the GIL, and seeing that many people just weren't good at multithreaded programming, I have started to conclude that the multithreading/multiprocess in Python today is about the best we can get. That is, instead of having threaded containers and multiple interpreter threads all banging on the same underlying data, it's a lot easier to just use threads as work queues that have minimal interaction with other threads.
So that's where I've ended up: some of my code using multiprocessing, typically the Pool or ThreadPool, with the concurrent future API to handle result-gathering, barriers, etc. Other code has an external system that starts many python processes from the command line and waits for those processes to complete. Other code is single-threaded in python and launches C++ cores that launch multiple threads just to return a computational result to python faster.
And I think trying to do both gil and nogil interpreters, rather than committing to one or the other, the python leadership will sign us up for untold inconveniences around packaging. We already see this in the move to async and it will only be worse with threading.
So sad to say I think sticking to GIL and using the approaches I mentioned above (along with others that work well for concurrent, rather than parallel, computing, like coroutines) is the best thing to do right now and I'm a bit bummed that the leadership signalled their intent to accomodate both.
The reason is that a single thread can invoke multicore c/c++ code (even calling into specialized accelerators if needed) but having python objects that are shared between threads is extremely clunky. And multiprocessing results in a lot of communication overhead. Worse-- python code cannot proceed in another thread. Mixing python and C++ is very common in ML and scientific computing workloads.
> And I think trying to do both gil and nogil interpreters, rather than committing to one or the other, the python leadership will sign us up for untold inconveniences around packaging. We already see this in the move to async and it will only be worse with threading.
How has packaging been affected by async? As long as you have a compatible python version what issue do you run into?
I personally think that pythons support for the massively parallel hardware we have is lacking, and the devs are being too slow and disparate to respond, combine that with what is basically subpar tooling around threads and you have people reaching for multiprocessing out of necessity.
Off the back of this, should Python just maybe give up on threads entirely? Should it relegate itself to simple-scripting and open the floor to something that can do these things?
I’ve personally stopped writing Python for these reasons- apart from “AI stuff” which isn’t something I dabble in anymore, there’s nothing that Python can do anymore that another language can’t do better, just as (if not more) easily, without giving anything up.
People think parallelism is a huge use case. Despite the fact that a ton of CPU cycles are spent on multicore-compatible code, in practice that's just not what the vast majority of programmers spend their days working on.
Everything that is a server is profiting massively off of true parallelism. Stuff written in Elixir, Golang and Rust scales much better than the same thing written in Python or Ruby. Me and many other colleagues have seen the before and after in our monitoring systems.
Maybe you should instead qualify your comments with "I have carefully and deliberately stayed away from the need to do parallel programming throughout my entire career" and I feel that way your comments would have the necessary context. Otherwise you are misleading less experienced people.
Some people would prefer a pure python connection pool to pgbouncer.
> More concretely, benchmarks show up to ~4× faster compile time, ~5× smaller binaries, and ~10× lower runtime overheads compared to pybind11.
The alternative road is to figure out how Python could automatically exploit parallelism in the underlying hardware, possibly in a way that would let it work on GPUs as well. The SIMT data-parallel way (seen in eg in ISPC, OpenCL, shader languages) is also more programmer friendly as it doesn't require the constant use of error probe synchronisation primitives. Or other HLL approaches in Jax, Futhark, etc.
The other option is you could allocate all the shared data off of the Python heap prior to the fork call.
It's been in Pyston for a while.
parallel -j96 'python -c "print('{}')"' ::: $(for i in {1..96};do echo $i;done)Subinterpreters are an answer (heck, so is multiprocessing).
Whether between them they are enough for Python's domains is another question. Probably not for the long term, but possibly for the neart term. But anything more is going to be a big lift, no-GIL is the obvious broader answer, so its good its being worked on, because by the time its betond question that its needed, it’ll be too late to start working in earnest.
* os.pipe and serialization (pickle or whatever): https://peps.python.org/pep-0554/#synchronize-using-an-os-pi...
* immortal object, but I don't see a way to create immortal object from Python (only from C). https://engineering.fb.com/2023/08/15/developer-tools/immort...
I guess it will more iteration to get a better way to communicate between the interpreters.
Not sure what the use case is.
This is exactly the use case. You can only parallelize in the native code if the boundary between Python and native code is absolute. But in practice people really do want to pass callbacks into the native code, inherit from native code interfaces in Python, even something as simple as forwarding the logging in their native code back to Python logging (e.g. all of the really useful behavior possible with binding tools like pybind11). All of these are impossible to parallelize effectively today.
There are caveats, but multiprocessing rarely gives you enough extra for the overhead.
Why are you better off that way?
also `import logging` acts funny as well.