Progress on No-GIL CPython
lwn.net
lwn.net
With modern computers I wonder if explicit parallelism is more fundamental to what our computer science will be than is in vogue in textbooks. Perhaps we should always be writing explicitly parallel code at this point.
eg `for` loops are being replaced by `foreach` loops,`map` and `filter` operations, etc. These tell the compiler/interpreter that you want to do some operation to all the items in your datastructure, leaving it up to the compiler/runtime whether and how to parallelize the work for you.
I've thought this way ever since MacOS added the Grand Central Dispatch [1]. Of course, I thought industry would follow quickly and that tooling would coalesce around this concept pretty quickly. Seems the industry wants to take its sweet time.
For basic parallelism, nothing beats OpenMP for ease of adapting existing code (often a single "#pragma omp parallel for" directive is enough). Even for more complex parallelism, particularly where per-thread resources need to be managed, OpenMP still provides a much simpler programming model than the alternatives.
I'd also say that most languages have something similar to OpenMP, parallel for loops, etc. Great if you have some read only data in arrays and wish to process it.
However, in my opinion, it doesn't really matter how convenient a parallel / async programming model is to use as the real work is ensuring that there isn't any shared mutable state being updated in parallel. The other issue is, once you have formulated / re-formulated a particular problem to this model, ensuring that it remains this way is pretty challenging on larger teams. Someone can easily unknowingly commit something that breaks such assumptions.
Parallelism introduces an additional class of bugs, but they are fundamentally addressed the same way as any other class of bugs - e.g. testing, tools, and code review. If some_one_ can unknowingly break a system, that means the tools and processes weren't good enough.
Also, the code introducing a race condition may get lucky when your CI system runs the tests and still make it into your main branch.
I agree that tooling (like static analysis, Rust's borrow-checker, etc) can play a big role here though.
Particularly for dynamic analysis, you need to have test cases that usefully cover the design behavior. E.g. if you design a component to be safely shared, you need tests that exercise that sharing where the static/dynamic analyzer(s) will identify unsafe sharing. Likewise, if you know something is unsafe, you should probably have tests that demonstrate that the static/dynamic analyzer(s) do detect the unsafe usage.
My HPC programming lectures were done on PVM, and I bet only grey beards know what it stands for.
We're inching towards an in vogue way to do what erlang had figured out in the 80's. We'll pick up the pace any day now. Surely.
> Humans are bad at reasoning about multiple threads simultaneously
I am not so sure this is true, I do believe that people are poorly practiced. My experiences have led me to believe Universities silo explicit parallel programming too much. It's generally it's own non-compulsory subject in a Comp-Sci major.
Humans are bad at reasoning about way too many things. I think mostly because many are lazy and do not want to learn. The ones who do have little problems. I do not find thread management particularly hard for the most parts (there are some exceptions but those are very uncommon).
> Reordering of reads and writes can be done both by the compiler and by the processor. Compilers and processors have done this reordering for years, but on single-processor machines it was less of an issue.
https://learn.microsoft.com/en-us/windows/win32/dxtecharts/l...
This is why we have things like WaitForSingleObject and many other that deal properly with the chance of reordering and other concurrency related issues. All is fine with the reasoning on CPU, OS, Compiler and my own level. One just have to understand what is going on and know the tools. Those who are setting boolean flag to indicate the data is ready should not be programming for modern CPU's and have a basic knowledge first.
But ensuring some parallelism while maintaining thread safety is straightforward in many contexts - an uncontended mutex is close to zero overhead. Languages like Rust make it harder to struggle with some of the thornier issues (data races which cannot be simulated by stepping threads).
E.g. in Java a typical system looks like a thread per request with some shared underlying data structures like caches or connection pools. Relatively easy to use these safely, or to guard some shared object with a synchronized.
Likewise with parallelism - a lot of problems just boil down to ‘do a few map reduces’ and the parallelism is pretty trivial.
Obviously, concurrent systems are fiendish to reason through - but there are a lot of cases where the complexity can be side stepped. Doesn’t seem to stop people writing scary code on the daily though.
How this has played out in my life gives me caution about making this standard in computing.
There's difference between doing it in order 1, 2, 3 and 3, 1, 2.
foreach will not be replaced behind the scenes into multithreaded version since it changes behaviour.
for is replaced with foreach because usually you dont need index and foreach is just handier and safer, that's it.
.NET's std lib has Parallel.ForEach for such a thing.
We really don't need magic to write multithreaded code. All we need is just really, really well designed APIs and primitives.
It only (meaningfully) changes behavior if you're both iterating over an odered datastructure and the body of your loop has direct or indirect side-effects. (like printing, writing to a file, making network requests, etc)
Right, and not always even then, because that depends on what the consumer is concerned with as well. But the fact that it can means it's not a safe automatic substitution.
So like... huge % of the real world code bases
Every statement in a HDL language runs in parallel but you can still write implicitly sequential code in VHDL processes.
Only now I can enjoy in modern hardware what I had to imagine when reading papers about Star Lisp and the Connection Machine, alongside other similar approaches.
The great thing about the disruptor it is that multiple threads can receive the same message without much more effort.
https://github.com/LMAX-Exchange/disruptor/wiki/Performance-...
https://gist.github.com/rmacy/2879257
I am dreaming of language that is similar to Smalltalk that stays single threaded until it makes sense to parallise.
I am looking for problems for parallelism that are not big data. Parallelism is like adding more cars to the road rather than increasing the speed of the car. But what does a desktop or mobile user need to do locally that could take advantage of the mathematical power of a computer? I'm still searching.
I am thoughtful of the Itanium and VLIW architecture for parallelism ideas.
The things we currently let servers do but it would mean we can keep user data local and not hand it over to service providers. I believe that is a worthy end goal.
Multithreading doesn’t always have to be around increasing speed, it can also reduce power
It's an implementation detail, and if we can abstract it away to make it easier to utilize, we should.
And there are things even further afield like what I work on [1][2][3].
I noticed the Regent code is inside the Legion repo. Is Legion the system, and Regent the language?
Can Legion be used without Regent, or vice versa?
Regent is a programming language. The compiler for Regent generates Legion code. Semantically, Regent is mostly a simplification of Legion. There are fewer moving pieces, so fewer things you need to worry about. Many of the "gotchas" that exist in Legion are taken care of by the language/compiler so idiomatic code usually "just works". It also does GPU code generation for you so you don't need to hand-write CUDA/HIP/etc. The tradeoff is that you're using a new programming language, so you have to be willing to take that risk.
"Safely" for a certain kind of definition of safety: https://github.com/search?q=repo%3Arayon-rs%2Frayon+unsafe&t...
When people say Rust is better for memory safety than e.g. C++, it's not because you can't write a provably-safe library with C++—you can do that with some effort. It's because the Rust language differentiates compiler-asserted safety from programmer-asserted safety via the `unsafe` keyword, an opt-in.
[1]: https://rust-lang.github.io/unsafe-code-guidelines/glossary....
(edit: Sniped, but I believe I've expanded upon the sibling comment.)
> ... would be disingenuous, or at least very ignorant.
That's quite literally what it means - there is a "safe" Rust code that never uses the "unsafe" but the example given by the parent comment is just not the example of that. I am not sure why would my comment come out as ignorant - it's factual state of things.
> it's not because you can't write a provably-safe library with C++
Hm, you really can't? Neither you can with safe or unsafe Rust. None of the compilers for those languages are formally verified.
For things like running a web service, requests are fast enough, and the real win from parallelism is in handling lots of requests side-by-side. This is where No-GIL comes in.
Within handling a single request, if there are a lot of sub-requests, that's usually handled by async code, but not so much for the async performance win as much as spinning up threads is either expensive or thread pools are a hassle. Remember that async is better for throughput, but worse for latency, and if you're parallelizing a service request, you're probably more worried about latency. Async won mostly on ergonomics.
The other place you see parallelism is large offline jobs. Things like Map-reduce and Presto. Those tend to look like divide-and-conquer problems. GPU model training looks something like this.
What never happened is local, highly parallel algorithms. For a web service, data size is too small to see a latency win, they're complicated, and coordination between threads become costly. The small exceptions are vectorized algorithms, but these run one one core, so there isn't coordination overhead, and online inference, but again, this is heavily vectorized.
GPUs maybe? Also, excellent answer.
There's interesting stuff going on in the VHLL world with languages like Futhark, Jax, Mojo, etc that would be a better peer group for Python and its high level of abstraction.
If you can find enough independent sequential problems in your programs then you can easily fill up cores, mostly because we don't have that many. I only have eight.
The problem is that this requires additional graph processing and there it makes your programs slower, which kind of defeats the point. The goal is to find the right tradeoff.
always good to have a pun.
Sharper teeth in that version. Same shaped creature though!
from future import nogil
It would hot swap interpreters at that point.
try:
import nogil
except ImportError as ie:
print(ie)
nogil = None
if nogil:
# gill free code here
pass
else:
# gil requring code here
passhttps://docs.python.org/3/reference/simple_stmts.html#future...
future statements are module specific and GIL/no-GIL doesn't fit easily into that model.
https://peps.python.org/pep-0397/
But for all platforms. I should have put that into the comment.
Specifically, talking about data races I’ve seen time and again in codebases across companies and OSS projects. The programs don’t break only because they implicitly rely on the GIL providing execution to a single thread at a time. If the GIL is gone, then these programs will break. And since Python is such a dynamically typed language, I seriously doubt that there exists a static analyzer that could identify these issues in existing Python programs. More likely, they’ll be insidious bugs that crop up at runtime in a non-deterministic fashion. Ideally leading to a crash, With this class of bugs, it’s likely to just result in incorrect operations being performed.
Perhaps this GIL-less proposal isn’t actually intended to be used on the overwhelming majority of programs? Maybe it’s just a hyper specialized tool for a very few number of circumstances where the programmer knows, there won’t be a GIL, and can program against that fact?
Not saying that's what they will do, just what they could.
When the count goes from zero to one, the thread attempts to upgrade its reader lock to a writer lock. When its count goes from one to zero, the thread downgrades its lock from writer to reader.
That way, there's at most one thread executing within GIL-dependent code at a time. Furthermore, if there is a thread executing within GIL-dependent code, all of the other threads are blocked waiting to acquire the GIL (in reader mode if they're nogil-safe, and writer mode if they're GIL-dependent.)
As now, any thread holding the GIL in writer mode would need to drop the GIL when attempting to acquire any other lock (and re-acquire immediately afterward).
To prevent starvation, one would presumably need a mechanism similar to periodic GC safepoints where nogil-safe threads still check if any thread is waiting to acquire the GIL in writer mode.
Python bombs before you can gain any control, because it parses the whole file. So you can separate things into modules to get around that, but it's not great when you want a simple one-file script.
Expecting every transitive module to add this marker is very optimistic though. It’s thousands of packages with hundreds of modules each, that need to add this everywhere.
Other languages have done similar journeys, like typescript “strict” added at top of each file. Except those are a lot more local, by not expecting all dependencies to follow.
Will look it up.
Which projects do you have in mind that have a significant user base, are still maintained, and would be too costly to port for someone to do the effort?
Multi-threading, concurrency, and parallelism are fraught with problems. Your precious ML/DL libraries may not even be upgraded, because writing neural network code is not the same as writing thread-safe code. If it comes from a CS lab, its authors have already left, and there's nothing worthy of a publication in adding thread-safety. Certainly not when you can simply stick to Python 3.last-gil-version.
PyTorch and sklearn won't stop being maintained though. I don't rely on unmaintained research code in production, I adapt what I need under MIT license. Any other way sounds crazy.
Plus, most research code is very high-level and uses the same facilities (from e.g. PyTorch again) that everyone else uses, the actual distributed and multithreaded work happens in the main libraries. You'll still be able to use the same neural network code that worked before.
I don't see a huge problem for people who already had their dependency list under control. If you had anything that's both hard to replace and not big enough to be upgraded though, I'd argue that it was always going to bite you at some point.
Even inside a C extension where the Python API feels like it gives you control over when you release the GIL (with functions you'd have to call explicitly to release the GIL), it turns out that:
* any operation that allocates new Python objects might trigger garbage collection
* garbage collection may run `__del__` of objects completely unrelated to the currently running C code
* `__del__` can be implemented in python, thus releasing the GIL between bytecode instructions
Thus there's a lot of (rarely exercised) potential for concurrency even in C extensions that don't explicitly release the GIL themselves. nogil will make it easier to trigger data race bugs, but many of them will already have been theoretically possible before.
Anyone who hasn’t upgraded to it by now is needlessly spending extra on compute.
3.8 is the oldest supported version, so I would hope not, but probably.
Not the whole world has the luxury to upgrade all their systems all the time.
Excuse me, would you mind stopping this factory for a few hours so we can replace this perfectly functioning system with an untested one that may or may not work in roughly the same way? is not a question that is generally met with wild enthusiasm.
Don't downvote me, answer the question.
"The point is that breaking legacy code doesn't have to be a show-stopper for eliminating the GIL, and backwards-compatibility does not have to be (indeed should not be) an overriding consideration in a GIL-free Python."
That has nothing to do with opt-in vs opt-out.
My understanding is the GIL does not protect against Python-side bugs, and bugs from GIL removal would only be introduced from C extensions.
This has been discussed extensively in the past (1), and my understanding of the take away was that the GIL doesn't protect you from arbitrary execution order; it protects you from undefined behavior due to concurrent write/read in parallel scopes and the resulting data corruption.
...which, as I understand it, there is no specific reason it would be restricted to native extensions.
Is there some more detail to the nogil proposal that addresses the type of UB you see in eg. c, with this? (Wouldn't that require that at some level the GIL still exists?)
Not really, no. The finer grained synchronization primitives are (a) already available in Python, and (b) necessary even with the GIL for reasons I've given elsewhere in this discussion.
What nogil does is enable multiple threads to run Python bytecode at the same time, so that CPU intensive operations in Python can be parallelized without having to use multiple processes. Python objects that are accessed from multiple threads will have to be guarded with synchronization primitives under some circumstances where, in principle, they don't need to be guarded now (operations that only take a single bytecode), but in practice I don't think that will make much difference. The big issue, as has been mentioned elsewhere in this discussion, is C extension modules.
The finer grained synchronization primitives does not already available in Python. Or, it should not be visible to Python code at all. What I'm talking about is the internal implementation of, e.g. PyDict. While from Python bytecode side setitem on it is already not atomic, it does guarantee that Python interpreter won't segfault if there are two Python threads manipulating one dict object concurrently. This is achieved via GIL and has to be replaced.
It's the same problem you mentioned above as "problems in C extension". But no, nogil is hard not only because of compatibility issues. People (especially those who insist on that their workload is inherently embarrassingly parallel) do NOT accept any regression in single-thread performance. If you only ever want to optimize for single thread one global lock is the optimal solution.
No, it doesn't, except for operations that complete in a single byte code. See my response to shrimpx downthread.
like l=[], then l.append(1) and l.append(2) running concurrently will not end up in some weird scenario where the length of the list is 1 yet you stored two items... anyways, I agree with the comment you posted higher up in the discussion, and that was my understanding.
Yes, and I weighed in on similar lines in that discussion:
a += b
takes four bytecodes: two LOAD bytecodes to put the values of a and b on the stack, an INPLACE_ADD bytecode to put the result at the top of the stack, and a STORE bytecode to store the result in the variable a. A thread switch could take place in between any two of these bytecodes, and if another thread mutated a or b or was reading the value of a or b, you would have data corruption. The GIL does not prevent this. The only way to prevent it would be to use one of the locking mechanisms provided to guard access to the a and b variables so that only one thread could access them at a time.
import threading
i = 0
def test():
global i
for x in range(1000000):
i += 1
threads = [threading.Thread(target=test) for t in range(20)]
for t in threads:
t.start()
for t in threads:
t.join()
assert i == 20000000, i
Presumably the assert should fail sometimes but I haven't been able to observe that.Even multi-process Python code is often broken. The "recommended" way to serve a Django app is to run multiple workers (processes) using gunicorn. If you point the default logs to a file, even with log rotation enabled, all workers will keep stepping over each other because nobody knows which file to use. Keep in mind that this is broken by default, and this is the recommended way to use all this.
a) It isn't a language maintainers job to make sure unsafe code written by someone else in that language runs correctly.
b) The GIL doesn't prevent data races. It keeps the internal running of the interpreter sane, that's it's only job. There is a reason the threading library has a plethora of lock-classes
So the existing threads API spawns threads in the current domain, which lets you isolate code that expects to take the lock, while new code can spawn new domains starting with one thread instead. You can also use both deliberately as a form of scheduling.
Python instead is trying to make the lock entirely optional, globally and outside the control of library writers. However, I think the Python lock is only guaranteed to protect the runtime itself, so most code depending on it is probably buggy anyway, so I think their plan is viable.
The only thing they may have in common is having to scan the entire codebase of their runtimes for unexpected shared state and fix that, as well as revising their C ABIs.
It's immature logic from people that lack a sense of perspective of what engineering compromises are and the fact that even substantial language warts don't move the needle as far as justifying language switches.
If you know about Python performance and what's the most you can squeeze from the language, its libraries implemented in C for high performance, and other things in the ecosystem, you already knew if the code you're writing was good for Python or not. If it was, between Cython, C bindings, and NumPy you should be covered. If not, you can drop down to C, Rust and C++ libraries and the slow performance of the glue code is an afterthought.
It is the opinion of an overwhelming majority of coders, who need high-performance code for maybe 5% of what they write (and most of that being code getting run relatively very little), that higher performance is not as high a priority for those users as some people think.
The rationalization in your response obfuscate the real reason you were triggered to down vote, which is that you are too emotionally vested in Python and afraid to try a better alternative.
An alternative language has to be far superior to its alternatives to justify the switching costs.
python4
python3-gilfoil
python3-gilfreeI of course respect that our needs may not represent everyone else's, and we are grateful for all the work that has gone into making Python a great language. But what am I missing?
I agree with you that if it was one or the other, most use cases are gonna be better met with straight up faster single threaded code. But why not have both?
For users, it'll be mixed. Users writing single-threaded code shouldn't have to change a single line of code, but they'll see (sans the concurrent efforts to speed up the underlying implementation) slower performance due to everything done to achieve no-GIL (the actual net effect will be a performance boost due to that concurrent effort over time). If they're writing multi-threaded code, then they should be writing it to be threadsafe now, not assuming that the GIL will protect them (because it doesn't guarantee it now anyways). So nothing should actually change for most users of Python if they're writing correct code today.
But that's before fastercpython improves single threaded performance closing that gap if not reversing it.
So yeah I’ve never stressed much about the GIL as a long time pythons dev
> So yeah I’ve never stressed much about the GIL as a long time pythons dev
You almost always have to use python objects at some point or another, even with very high perf libraries. Unless you are using python as nothing else than glue code, and loading and preprocessing the data entirely outside of your python code. At some point (and I'm not talking about extreme scales here, just feeding a single datqaloader can be enough) you hit a very hard performance bottleneck. Sure you can just use more external code at that point but I think the intention is to make python at least more suitable for the non compute intensive stuff.
Like, it is already suitable right now but it's not getting better, while everything around it is (I'm not talking about other languages, what I mean is that the tools and libraries are getting better so without improvement to the language the bottleneck will just get worse)
The other option is you could allocate all the shared data off of the Python heap prior to the fork call.
It's been in Pyston for a while.
parallel -j96 'python -c "print('{}')"' ::: $(for i in {1..96};do echo $i;done)A long time ago IronPython was released which showed that you could build a high performance Python interpreter that was GIL-free. It had thread-safe containers (so multiple threads could work against the same lists, dictionaries, etc) and in some cases was faster than CPython (it was implemented using CLR and .net)
When I saw IronPython I was immediately convinced that CPython should be the same way- already low-cost dual and quad core Intel machines were becoming available and it seemed like core counts were going to increase faster than clock rates. I figured that a small hit to serial performance would be more than acceptable if people could write multithreaded systems in Python, in much the same way as I wrote multithreaded C++.
Over time after watching nogil not going anywhere in CPython (the python leadership didn't want to do nogil), with concomitant speedups in the single-processor implementation, along with the increasing use of C++ code that releases the GIL, and seeing that many people just weren't good at multithreaded programming, I have started to conclude that the multithreading/multiprocess in Python today is about the best we can get. That is, instead of having threaded containers and multiple interpreter threads all banging on the same underlying data, it's a lot easier to just use threads as work queues that have minimal interaction with other threads.
So that's where I've ended up: some of my code using multiprocessing, typically the Pool or ThreadPool, with the concurrent future API to handle result-gathering, barriers, etc. Other code has an external system that starts many python processes from the command line and waits for those processes to complete. Other code is single-threaded in python and launches C++ cores that launch multiple threads just to return a computational result to python faster.
And I think trying to do both gil and nogil interpreters, rather than committing to one or the other, the python leadership will sign us up for untold inconveniences around packaging. We already see this in the move to async and it will only be worse with threading.
So sad to say I think sticking to GIL and using the approaches I mentioned above (along with others that work well for concurrent, rather than parallel, computing, like coroutines) is the best thing to do right now and I'm a bit bummed that the leadership signalled their intent to accomodate both.
The reason is that a single thread can invoke multicore c/c++ code (even calling into specialized accelerators if needed) but having python objects that are shared between threads is extremely clunky. And multiprocessing results in a lot of communication overhead. Worse-- python code cannot proceed in another thread. Mixing python and C++ is very common in ML and scientific computing workloads.
> And I think trying to do both gil and nogil interpreters, rather than committing to one or the other, the python leadership will sign us up for untold inconveniences around packaging. We already see this in the move to async and it will only be worse with threading.
How has packaging been affected by async? As long as you have a compatible python version what issue do you run into?
I personally think that pythons support for the massively parallel hardware we have is lacking, and the devs are being too slow and disparate to respond, combine that with what is basically subpar tooling around threads and you have people reaching for multiprocessing out of necessity.
Off the back of this, should Python just maybe give up on threads entirely? Should it relegate itself to simple-scripting and open the floor to something that can do these things?
I’ve personally stopped writing Python for these reasons- apart from “AI stuff” which isn’t something I dabble in anymore, there’s nothing that Python can do anymore that another language can’t do better, just as (if not more) easily, without giving anything up.
People think parallelism is a huge use case. Despite the fact that a ton of CPU cycles are spent on multicore-compatible code, in practice that's just not what the vast majority of programmers spend their days working on.
Everything that is a server is profiting massively off of true parallelism. Stuff written in Elixir, Golang and Rust scales much better than the same thing written in Python or Ruby. Me and many other colleagues have seen the before and after in our monitoring systems.
Maybe you should instead qualify your comments with "I have carefully and deliberately stayed away from the need to do parallel programming throughout my entire career" and I feel that way your comments would have the necessary context. Otherwise you are misleading less experienced people.
Some people would prefer a pure python connection pool to pgbouncer.
Not sure what the use case is.
This is exactly the use case. You can only parallelize in the native code if the boundary between Python and native code is absolute. But in practice people really do want to pass callbacks into the native code, inherit from native code interfaces in Python, even something as simple as forwarding the logging in their native code back to Python logging (e.g. all of the really useful behavior possible with binding tools like pybind11). All of these are impossible to parallelize effectively today.
The alternative road is to figure out how Python could automatically exploit parallelism in the underlying hardware, possibly in a way that would let it work on GPUs as well. The SIMT data-parallel way (seen in eg in ISPC, OpenCL, shader languages) is also more programmer friendly as it doesn't require the constant use of error probe synchronisation primitives. Or other HLL approaches in Jax, Futhark, etc.
There are caveats, but multiprocessing rarely gives you enough extra for the overhead.
Why are you better off that way?
also `import logging` acts funny as well.
Subinterpreters are an answer (heck, so is multiprocessing).
Whether between them they are enough for Python's domains is another question. Probably not for the long term, but possibly for the neart term. But anything more is going to be a big lift, no-GIL is the obvious broader answer, so its good its being worked on, because by the time its betond question that its needed, it’ll be too late to start working in earnest.
* os.pipe and serialization (pickle or whatever): https://peps.python.org/pep-0554/#synchronize-using-an-os-pi...
* immortal object, but I don't see a way to create immortal object from Python (only from C). https://engineering.fb.com/2023/08/15/developer-tools/immort...
I guess it will more iteration to get a better way to communicate between the interpreters.
> More concretely, benchmarks show up to ~4× faster compile time, ~5× smaller binaries, and ~10× lower runtime overheads compared to pybind11.
I get why the PEP exists. I don't get why it's receiving such priority.
In all, there might be reason to expect GILless python to be faster single core in certain scenario's.
With nogil python you can effectively have multiple threads that eg; call out to C code while having shared state accessible as python objects. This is pretty key for ML-- in fact this current incarnation of the PEP came from the PyTorch team.
Single threaded performance is important too but there have already been lots of decent workarounds for critical sections (eg; numba, Cython, and now things like Mojo).
The ordering is important too-- a lot of the faster cpython work would be thrown away completely if nogil came about. So the teams have had to coordinate.
In the ideal world that means both nogil mode + improvements to single threaded performance (Guido even hinted sophisticated JITing is being considered).
12 year transition and single threaded performance is still abysmal and it has a few painful transitions left to get to real multi-threading.
As kind as one should be with opensource development, at some point is it fair to call it a very poorly managed language?
A simple C implementation allows everyone and their mother to hack at it and interface with C libraries and add try features and evolve the language through endless PEPs.
Compare Python's language evolution to Java's abysmally slow language evolution because every new feature has to be implemented in a way that works with the JVM's JIT-compatible speed hacks. A _ton_ of very useful things Java could have done simply cannot be done because you can't work against the grain of the JVM. If Python's backward-compatibility is a pain, you have no idea how bad it is in the JVM (see how all the JVM's caveats have hobbled the semantics of generics as a good example).
Not providing the end users any guarantees about performant code means that coders offload the performant areas to libraries that use languages built for performance, keeping that complexity outside of the interpreter and language.
Shrug. Is it really that much worse than JavaScript or PHP?
I mean, maybe they are or have been abysmal, but at least Python is in popular company.
The simple fact is that performance is not a sufficiently important problem for the domain Python works it. At least not important enough to give up its dynamicism.
It was predictable and peopled commented at the time. Perl 5/6 was given as an example. And when it became apparent that nobody was switching, it still took about 5 years until they tried making it easier.
Worker [thread] is a sandboxed VM [context] with its own GIL.
Communication between these is done through messaging so no sync primitives are required.
Go's routines have similar concept.
I think that if it OK for Go it should be OK for Python, no?
[0] https://developer.mozilla.org/en-US/docs/Web/API/Web_Workers...
1. Threads are significantly cheaper than processes in terms of memory, setup and context switch overhead
2. Inter-thread-communication is simpler and more performant than inter-process-communication
3. C extensions can share data structures across (sub)interpreters
Not on Windows!
Then reality comes in and says, "nah".
Honestly at some point its better to likely start porting code to a language where you can control parallelism better. There's a few nice modern options that don't involve all the headaches of C++ or C. Could even be extensions as has been the case since the dawn of time.
But this nogil version is the first time we have an actually working GIL removal. All of the other ones were incomplete to the point of being non starters, and mostly served as discussion material. This is an actually working implementation which deals with the subtle issues that the other projects didn't even get to, and has gotten to the point that it's a technical possibility to commit it to main (though obviously with a huge migration to think about). So in this sense this is a very different discussion than all of the previous discussions about the GIL
* The PEP has been announced to be accepted (Steering Council are still working on details for final wording)
* Many things are already landing on CPython main in preparation for this
Unless something absolutely show stopping comes up in the next 12 months it will almost certainly released in 3.13 or 3.14 as an optional compiler flag.
There have been some proposals to add full shared memory constructs (SharedArrayBuffer) and synchronization (Atomics) mechanisms, but they are special constructs and don't work with normal javascript objects. Quite limited but provide full parallelism for things that usually need it (buffers).
One thing people often forget is that thread-safe data structures are usually a lot slower than single-threaded one, everything in JS is single threaded and in the event loop, but if you really need it there are some scape hatches.
I don't know, this just feels better and simpler? If you really need to you can go down to a lower level language for full memory sharing data structures.
Python has become essentially the primary language for scientific computing and deep learning, and for this the JavaScript approach is absolutely inadequate. Real shared memory parallelism is needed to avoid the unnecessary copying that the subprocess approach entails.
>I don't know, this just feels better and simpler? If you really need to you can go down to a lower level language for full memory sharing data structures
A low level language is already used in Numpy, PyTorch etc., but that's not enough; the Python glue code becomes unnecessarily painful when using multiprocessing, pain that would go away with proper threads.
However, the reason to use JS in particular is because when you have a lot of I/O blocking calls (DB calls, web requests) it really is a lot faster. But if you do anything with numbers, JS is the wrong language - it's so limited here that it's almost laughable. Same with dates. That makes it a very bad choice for data science.
That's pretty far from a root cause though. I'd say it has these because - get this - it's actually a quite nice language to work with, so some people prefer it, and other people can handle it, unlike more difficult languages. In particular, Python is fairly small / minimalist (though the cruft is accumulating...), straightforward and powerful.
That is somehow barely mentioned among all the performance complaints.
Me, usually writing C++, I like it for writing various little helper tools. No deep notebook tensors are involved in that work. It's mostly file and string processing, in one case also a GUI and Excel (unfortunately) tables.