Everything you need to know about Python 3.13 – JIT and GIL went up the hill
drew.silcock.dev
drew.silcock.dev
What is the real world benefit we will get in return?
In the rare case where I need to max out more than one CPU core, I usually implement that by having the OS run multiple instances of my program and put a bit of parallelization logic into the program itself. Like in the mandelbrot example the author gives, I would simply tell each instance of the program which part of the image it will calculate.
There are quite a few common cases where in process multi threading is useful. The main ones are where you have large inputs or large outputs to the work units. In process is nice because you can move the input or output state to the work units instead of having to copy it.
One very common case is almost all gui applications. Where you want to be able to do all work on background threads and just move data back and forth from the coordinating ui thread. JavaScript’s lack of support here, outside of a native language compiled into emscripten, is one reason web apps are so hard to make jankless. The copies of data across web workers or python processes are quite expensive as far as things go.
Once a week or so, I run into a high compute python scenario where the existing forms of multiprocessing fail me. Large shared inputs and or don’t want the multiprocess overhead; but GIL slows everything down.
I thought transferring array buffers through web workers didn’t involve any copies of you actually transferred ownership:
worker.postMessage(view.buffer, [view.buffer]);
I can understand that web workers might be more annoying to orchestrate than native threads and the like but I’m not sure that it lacks the primitives to make it possible. More likely it’s really hard to have a pauseless GC for JS (Python predominantly relies on reference counting and uses gc just to catch cycles).This assumption is incorrect. There are plenty of problems that consist entirely of business logic manipulating large and complex object graphs. “Just rewrite the hot function in rust, bro” and “just import multiprocessing, bro” are functionally identical to rewriting most of the application for these.
The performance work of the last few years, free threading and JIT are very valuable for these. All the rest is already written in C.
Isn't "just use threads, bro" likely to be equally difficult?
(random question, totally understand if you're not the right person to ask)
(Huh, people like hard OS design problems for marginal behavior? OSes had trouble adopting SMP but we also got to jettison a lot of deadlock discussions as soon as there was CPU 2. It only takes a few people not prioritizing 1 CPU testing at any layer to make your 1 CPU container much worse than a 2 VCPU container limited to a 1 CPU average.)
When you set "request: 1" in Kubernetes or another container manager, you're saying "give me 1 CPU worth of CPU time" but if the underlying Linux host has 16 logical cores your container will still see them.
Your container is free to use 1/16th of each of them, 100% of one of them, or anything in-between.
You might think this doesn't matter in the end but it can if you have a lot of workloads on that node and those cores are busy. Your single threaded throughout can become quite compromised as a result.
While yes this can cause a slowdown, wouldn't it still happen if each container thought it had a single core?
Also you said "a lot of workloads" so yes probably more containers than cores.
I don't think the scheduler picking a different core matters much unless your workload is super cache sensitive. My point is more about access to single threaded performance. If you have a single threaded workload (ex: an ffmpeg audio encode) and you want it to be able to access as many cycles from a single core as possible, it isn't always as simple as request: 1
On Docker, --cpuset-cpus=0 will pin the container to the first core.
K8s: https://kubernetes.io/docs/tasks/administer-cluster/cpu-mana...
CPU affinity and pinning is something I think you should be able to achieve without too much hassle.
Just let the CPU scheduler do its job. Unless you know better, in which case, by all means go ahead and allocate computational resources manually. I don't see a way to make that a sensible default, though.
Neat, I didn't know it was a single flag in Docker.
The k8s method you linked definitely has some caveats as it doesn't allow this scheduling at the pod level, and requires quite a bit of fiddling to get working (atleast on GKE). This isn't even available if you use a fully managed setup like autopilot.
Maybe my expectations just aren't realistic but "easy" to me would mean I put the affinity right next to my CPU request in the podSpec :/
I can't think of any actual computer outside of embedded that has been single core for at least a decade. The Core Duo and Athlon X2 were released almost 20 years ago now and within a few years basically everything was multicore.
(When did we get old?)
If you mean that single core workloads will be extinct, well, that's a harder sell.
* Most of the programs I write are not (trivially) parallelizable, and a the bottleneck is still a single core performance
* There is more than one process at any time, especially on servers. Other cores are also busy and have their own work to do.
1. Other people with different needs exist.
2. That's why we have schedulers.
Single core computers are already functionally extinct, but single-threaded programs are not.
Sometimes it’s also easier to split work using multiple threads. Other programming languages let you do that and actually use multiple threads efficiently. In Python, the benefit was just too limited due to the GIL.
This was the original reason for CPython to retain GIL for very long time, and probably true for most of that time. That's why the eventual GIL removal had to be paired with other important performance improvements like JIT, which was only implemented after some feasible paths were found and explicitly funded by a big sponsor.
If you have many CPU cores and an embarrassingly parallel algorithm, multi-threaded Python can now approach the performance of a single-threaded compiled language.
What's more I am now seeing in Julia that multithreading doesn't scale to larger core counts (like 128) due to the garbage collector. I had to revert to multithreaded again.
The real difference is the lower communication overhead between threads vs. processes thanks to a shared address space.
It took me less than an hour to add multiprocessing to analyze each file in its own process and merge the results together at the end. The runtime dropped to a couple seconds on my 24 thread machine.
It really was much easier than expected. Rewriting it in C++ would have probably taken a week.
let results = files |> Array.Parralel.map processFile
Literally that easy.Earlier this week, I used a ProcessPoolExecutor to run some things in their own process. I needed a bare minimum of synchronization, so I needed a queue. Well, multiprocessing has its own queue. But that queue is not joinable. So I chose the multiprocessing JoinableQueue. Well, it turns out that that queue can't be used across processes. For that, you need to get a queue from the launching process' manager. That Queue is the regular Python queue.
It is a gigantic mess. And yes, asyncio also has its own queue class. So in Python, you literally have a half a dozen or so queue classes that are all incompatible, have different interfaces, and have different limitations that are rarely documented.
That's just one highlight of the mess between threading, asyncio, and multiprocessing.
Here is the part of multiprocessing I used:
with Pool() as p:
results = p.map(calc_func, file_paths)
So, pretty easy too IMO.I agree though.
All this really means is that Python catches up on decades old language design.
However, it simply adds yet another design input. Python's threading, multiprocessing, and asyncio paradigms were all developed to get around the limitations of Python's performance issues and the lack of support for multicore. So my question is, how does this change affect the decision tree for selecting which paradigm(s) to use?
Threading is literally just Python's multithreading support, using standard OS threads, and async exists for the same reason it exists in a bunch of languages without even a GIL: OS threads have overhead, multiplexing IO-bound work over OS threads is useful.
Only multiprocessing can be construed as having been developed to get around the GIL.
In any other language, async is implemented on top of the threading model, both because the threading model is more efficient than Python's and because it actually supports multiple cores.
Multiprocessing isn't needed in other languages because, again, their threading models support multiple cores.
So the three, relatively incompatible paradigms of asyncio, threading, and multiprocessing specifically in Python are indeed separate attempts to account for Python's poor design. Other languages do not have this embedded complexity.
There are a lot of other languages. Javascript for example is a pretty popular language where async on a single threaded event loop has been the model since the beginning.
Async is useful even if you don't have an interpreter that introduces contention on a single "global interpreter lock." Just look at all the languages without this constraint that still work to implement async more naturally than just using callbacks.
Threads in Python are very useful even without removing the gil (performance critical sections have been written as extension modules for a long time, and often release the gil).
> are indeed separate attempts to account for Python's poor design
They all have tradeoffs. There are warts, but as designed it fits a particular use case very well.
Calling Python's design "poor" is hubris.
> So my question is, how does this change affect the decision tree for selecting which paradigm(s) to use?
The only effect I can see is that it reduces the chances that you'll reach for multiprocessing, unless you're using it with a process pool spread across multiple machines (so they can't share address space anyway)
Not in the least. Python is a poorly designed language by many accounts. Despite being the most popular language in the world, what language has it significantly influenced? None of note.
Hubris isn't rare.
> what language has it significantly influenced?
I can think of at least 1 language designer[1] who doesn't think it's "poorly designed," based on it's significant impact on what they're currently working on[2]
1. https://en.m.wikipedia.org/wiki/Chris_Lattner 2. https://www.modular.com/mojo
Oh? It is by far the fastest language for me. No languages comes close on the time from starting to write, to have code that runs. For me that time far outweighs the execution time, so it is a lot more important.
* PyTorch currently uses `multiprocessing` for that, but it is fraught with bugs and with less than ideal performance, which is sorely needed for ML training (it can starve the GPU).
* Tensorflow just discards Python for data loading. Its data loaders are actually in C++ so it has no performance problems. But it is so inflexible that it is always painful for me to load data in TF.
Given how hot ML is, and how Python is currently the major language for ML, it makes sense for them to optimize for this.
None. I've been using Python "in anger" for twenty years and the GIL has been a problem zero times. It seems to me that removing the GIL will only make for more difficulty in debugging.
One of the drawbacks of multi-processing versus multi-threading is that you cannot share memory (easily, cheaply) between processes. During model training, and even during inference, this becomes a problem.
For example, imagine a high volume, low latency, synchronous computer vision inference service. If you're handling each request in a different process, then you're going to have to jump through a bunch of hoops to make this performant. For example, you'll need to use shared memory to move data around, because images are large, and sockets are slow. Another issue is that each process will need a different copy of the model in GPU memory, which is a problem in a world where GPU memory is at a premium. You could of course have a single process for the GPU processing part of your model, and then automatically batch inputs into this process, etc. etc. (and people do) but all this is just to work around the lack of proper threading support in Python.
By the way, if anyone is struggling with these challenges today, I recommend taking a peek at nvidia's Triton inference server (https://github.com/triton-inference-server/server), which handles a lot of these details for you. It supports things like zero-copy sharing of tensors between parts of your model running in different processes/threads and does auto-batching between requests as well. Especially auto-batching gave us big throughput increase with a minor latency penalty!
I'm not in this space and this is probably too simplistic, but I would think pairing asyncio to do all IO (reading / decoding requests and preparing them for inference) coupled with asyncio.to_thread'd calls to do_inference_in_C_with_the_GIL_released(my_prepared_request), would get you nearly all of the performance benefit using current Python.
There is a lot of Python code that either explicitly (or implicitly) relies on the GIL for correctness in multithreaded programs.
I myself have even written such code, explicitly relying on the GIL as synchronization primitive.
Removing the GIL will break that code in subtle and difficult to track down ways.
The good news is that a large percentage of this code will stay running on older versions of python (2.7 even) and so will always have a GIL around.
Some of it however will end up running on no-GIL python and I don't envy the developers who will be tasked tracking down the bugs - but probably they will run on modern versions of python using --with-gil or whatever other flag is provided to enable the GIL.
The benefit to the rest of the world then is that future programs will be able to take advantage of multiple cores with shared memory, without needing to jump through the hoops of multi-process Python.
Python has been feeling the pain of the GIL in this area for many years already, and removing the GIL will make Python more viable for a whole host of applications.
Naturally I can easily compile my own Python 3.13 version, no biggie.
However from my experience, this makes many people that could potentially try it out and give feedback, don't care and rather wait.
[0] https://docs.python.org/3.13/whatsnew/3.13.html#free-threade...
A lot of langugage is still not optimized by tier 2 [1] and even less has copy and patch templates for JIT. And JIT itself currently has some memory management issues to iron out.
[1]: https://github.com/python/cpython/issues/118093
Talk by Brandt Butcher was there but it was made private:
Many of these kind of changes take time, and require multiple interactions.
See Go or .NET tooling bootstraping, all the years that took MaximeVM to evolve into GraalVM, Swift evolution versus Objective-C, Java/Kotlin AOT evolution story on Android, and so on.
If only people that really deeply care get to compile from source to try out the JIT and give feedback, it will have even less people trying it out than those that bother with PyPy.
Which itself needed to be compiled from source the first time I tried it. All the hours of Mandelbrot were worth the spectacular speedup.
As mentioned I don't have any issues compiling it myself, more of an adoption kind of remark.
You get both sides (yes, you might limit some who would otherwise try it out).
I think requiring people to compile to try out such a still-fraught, alpha-level feature isn't too onerous. (And that's only from official sources; third parties can offer compiled versions to their hearts' content!)
So much work, and so little recognition :-/ I was looking forward to trying it out but moved off python before the libraries I used were compatible (sqlalchemy I believe was the one I was really wanting...)
They should have just tried to go GIL-less instead of wasting time on trying transactional memory.
> STM was a research project that proved that the idea is possible. However, the amount of user effort that is required to make programs run in a parallelizable way is significant, and we never managed to develop tools that would help in doing so. At the moment we're not sure if more work spent on tooling would improve the situation or if the whole idea is really doomed. The approach also ended up adding significant overhead on single threaded programs, so in the end it is very easy to make your programs slower. (We have some money left in the donation pot for STM which we are not using; according to the rules, we could declare the STM attempt failed and channel that money towards the present GIL removal proposal.)
https://pypy.org/posts/2017/08/lets-remove-global-interprete...
Some of this complexity is also true for Windows, but Linux’s (good!) diversity makes it a bigger challenge.
Not end users but distro maintainers.
If you run almost any Linux distro imaginable, it has Python already in the distro.
If you want to try a version that's not yet even supported by the unstable branch of your distro, or not available on some PPA, etc, you need to be well enough versed in its dependencies to build it yourself.
Alpha-quality software, of course, benefits from more eyes, but it mostly needs certain kinds of eyes.
The problem is Linux ecosystem's fixation on "build environment = runtime environment" idea, making it incredibly difficult to build against older versions of glibc (e.g. Ubuntu 22.04) if you're using something new such as latest Arch. This is not a problem on macOS, you can use an SDK for a particular OS version on any other recent version.
Check out IndyGreg's portable Python builds. They're used by Rye and uv.
The JIT does not seem to help much. All in all a very disappointing release that may be a reflection of the social and corporate issues in CPython.
A couple of people have discovered that they can milk CPython by promising features, silencing those who are not 100% enthusiastic and then underdeliver. Marketing takes care of the rest.
You can find forum and Reddit posts going back 15-20 years of people attempting to remove the GIL, Guido van Rossum just made the requirement that single core performance cannot be hurt by removing it, this made ever previous attempt fail in the end
I.e. https://www.artima.com/weblogs/viewpost.jsp?thread=214235
The patches dropped some unrelated dead weight such that the effect is not as bad.
That is absolutely in the range of previous attempts, which were rejected! The difference here is that its goes in now to gratify Facebook.
I wonder if that’s something they could automate? I’m sure there are some weird risks with that. Maybe a small program ends up eating all your memory in some edge case?
It sounds like maybe you want GCs to be very tunable? That way, developers and operators can change how it runs for a given workload. That's actually one of the (few) reasons I like Java, its ability to monitor and tune its GC is awesome.
No one GC is going to be optimal for all workloads/usages. But it seems like the prevailing thought is to change the code to suit the GC where absolutely necessary, instead of tuning the GC to the workload. I'm not sure why that is?
Idea is to let people experiment with no-GIL to see what it breaks while maintainers and outside contractors improve the performance in future versions.
That was literally the official reason why it was accepted. Now we have slowdowns ranging from 20-50% compared to Python 3.9.
What outside contractors would fix the issue? The Python ruling class has chased away most people who actually have a clue about the Python C-API, which are now replaced by people pretending to understand the C-API and generating billable hours.
https://discuss.python.org/t/incremental-gc-and-pushing-back...
> What happens if multiple threads try to access / edit the same object at the same time? Imagine one thread is trying to add to a dict while another is trying to read from it. There are two options here
Why not just ignore this fact, like C and C++? Worst case this is a datarace, best case the programmer either puts the lock or writes a thread safe dict themselves? What am I missing here?
One could argue that he succeeded, considering how many members of the scientific community, who don’t primarily see themselves as programmers, use Python.
That the worst case being memory unsafety and a compromised VM is not acceptable? Especially for a language as open to low-skill developers as Python?
(Write single threaded code and have a compiler create multithreaded code)
https://en.m.wikipedia.org/wiki/Automatic_parallelization_to...
Not really automatic, but for some iterator chains you can just slap an adapter on.
Currently you have to think and benchmark, but for some scripting type applications the increased overhead of the heuristics might be justified, as long as it was a new mode.
I spent all day not knowing whether "up the hill" meant they shipped or didn't ship. So they shipped, right? Or they shipped a JIT but removed the GIL?
You’d think certain patterns could be probably safe and the interpreter could take the initiative.
Is there a term for this concept?
It is another one of those things that many programmers ask "Why hasn't this been tried?" and the answer is, it has, many times over. It just failed so hard and so fast you've never heard of the results. Toy problems speed up by less than you'd hope and real code gets basically nothing. Your intuition says your code is full of parallelization opportunities; your intuition turns out to be wrong. A subset of the general problem that even very experienced developers still never get great at knowing how code will perform without simply runninga profiler.
It has failed hard and fast in languages much friendlier to the process than Python. Proving something is truly parallel-safe in Fortran is hard enough. Proving it in Python is effectively impossible, you just don't know when something is going to dynamically dynamic the dynamics on you.
I'm going to use that!
Yeah right..
One problem about popularity of pypy is they don't do any advertisement , promotion which I had critically voiced about it in their community - they moved to Github finally.
Only other problem is CPython Ext , which is compatible but a little bit slower than CPython - that only pain point we have - which could be solved if there are more contributors and users using pypy. Actually Python , written in Python should be the main
Why ignore millions of dollars spent in a decade effort of fulltime PHD researchers' work and doing their own thing?
Yeah NIH is helluva drug.
On Heavy load (10k concurrent test) PyPy and Golang version are stable but Node version stop responding sometimes and packet losses occurs.
And if you use FastAPI then you're using Rust for serialisation and validation.
There’s probably a whole generation of programmers (if not two) who don’t know the feeling of shooting yourself in the foot with multithreading. You spend a month on a prototype, then some more to hack it all together for semi-real world situations, polish the edges, etc. And then it falls flat day 1 due to unexpected races. Not a bad thing on itself, transferrable experience is always valuable. And don’t worry, this one is. Enough ecos where it’s not “difficult to share data”.
Somehow, it’s perfectly fine in C#, Java, now Go, Rust and many other languages, with relatively low frequency of defects.
Of course changing the concurrency guarantees the code relies on and makes assumptions about is one of the most breaking changes that can be made to a language, with very unpleasant failure modes.
What I see from the development team and close community so far has been quite trust building for me. Slow and steady, gradual integration and testing with feature flags, improving related areas in preparation (like better/simplified C APIs), etc.
A similar set of tools and guardrails has yet to be seen for MT, at least in python-ic category. I think we can agree that relative success of MT in some X doesn’t automagically translate everywhere, because it depends on first principles which are different.
Python is absolutely the worst language to work in with respect to code formatters. In any other language I can write my code, pressing enter or skipping enter however I want, and then the auto formatter just fixes it and makes it look like normal code. But in python, a forgotten space or an extra space, and it just gives up.
It wouldn't even take much, just add a "end" keyword and the LSP's could just take care of the rest.
GIL and JIT are nice, but please give me end.
I am at the point where I prefer single quotes for strings, instead of double quotes, just because they feel cleaner. And unfortunately pep8 sometimes mandates double quotes for reasons unknown.
Are there any languages that use it, or is Python unique in using `:` to begin a block?
IIRC, Python's predecessor (ABC) didn't have the trailing colon but they did some experiments and found it increased readability.
>>> from __future__ import braces
SyntaxError: not a chanceAnd IMO for good reason. It makes the code so much cleaner and it's not like it particularly takes effort to indent your code correctly, especially since any moderately competent editor will basically do it for you. Tbh I actually find it much less effort than typing braces.
I've actively used Python for a quarter of a century (solo, with small teams, with large teams, and with whole dev orgs) and the number of times that not having redundant block delimiters has caused problems is vanishingly small and, interestingly, is on par with the number of times I've had problems with redundant block delimiters getting out of sync, i.e. some variation of
if (a > b)
i++;
j++;
Anytime I switch from Python to another language for awhile, one of the first annoying things is the need to add braces everywhere, and it really rubs because you are reminded how unnecessary they are.Anyway, you can always write #end if you'd like. ;-)
There are people that actually do this?
Most language design decisions involve tradeoffs, and for me this one has been a big positive, and the potential negatives have been - across decades of active usage - almost entirely theoretical.
Also, the real-world use case for importing code and auto formatting it is quite doable with Python - the main thing you can't do is blindly (unintelligently) merge two text buffers and expect it to work every time, but in general it takes only a smidgen of analysis on the incoming text to be able to auto reformat it to match your own. You could go all the way and parse and reformat it as needed, but much simpler/dumber approaches work just as well.
What you can't do is take invalid Python code and expect it to autoformat properly, but that borders on tautological, so... :)
(funny enough, that particular scenario is actually harder to miss in Python since it often produces an error that prevents the program from running at all)
some_function();
Python: some_function()
I am debugging how this function works so I actually only want it to run when some condition is true. C: if (some_condition) {
some_function();
}
Python: if some_condition:
some_function()
Oops. That's not indented correctly, so it won't run. To be fair neither actually looks good but that's 1. fine because this is for debugging and 2. for the C code I just format the file and it is instantly fixed. What if I have a loop? C: for (int i = 0; i < 5; ++i) {
iteratively_optimize();
}
Python: for i in range(5):
iteratively_optimize()
Unfortunately this breaks something so I want to step it through. In C I comment out the loop as follows: // for (int i = 0; i < 5; ++i) {
iteratively_optimize();
//}
In Python: // for i in range(5):
iteratively_optimize()
Nope, that's broken too. I can't just autoformat this code either because the formatter can't look at scope using anything else. I have to manually go and fix the indentation on that line too.These are actually very small cases. I can even imagine you saying, in C you have to fix two places if you want to comment out a loop or if: the opening brace and the closing one. In Python you need to comment out the control statement and the second thing is fixing the indentation, so what's the big deal? Well, as the number of lines in the block grows larger in C it's still just commenting out the two braces, while in Python you have to select the whole region line-perfectly and fix the indentation. As someone who writes both I always find this to be a lot more fiddly and annoying.
Python’s tradeoffs pay dividends every day, at the expense of few questions a year. Also code is read 10x more than written, where the extra delimiters lower signal to noise.
Wouldn't that make it behave pretty much as what you expect?
Maybe I've just lived a sheltered life, but I've never heard speed being used as a serious argument against Python. Well, maybe on silly discussions where someone really disliked Python, but anyone who actually cares about efficiency is using C.
With the excellent Python/Rust interop there is now another great alternative to rewriting the heavy parts in C. But sometimes the performance sensitive part spans most of your program
It is a slow interpreted language, but that isn't the only argument against it.
It has abandoned backwards compatiblity in the past and there are still annoying people harassing you with obsolete python versions.
The language and syntax still heavily lean towards imperative/procedural code styles and things like lambdas are a third class citizen syntax wise.
The strong reliance on C based extensions make CPython the only implementation that sees any usage.
CPython is a pain to deploy crossplatform, because you also need to get a C compiler to compile to all platforms.
The concept behind venv is another uniquely bad design decision. By default, python does the wrong thing and you have to go out of your way and learn a new tool to not mess up your system.
Then there are the countless half baked community attempts to fix python problems. Half baked, because they decide to randomly stop solving one crucial aspect and this gives room for dozens of other opportunistic developers to work on another incomplete solution.
It was always a mystery to me that there are people who would voluntarily subject themselves to python.
At least mine. Also because of the typing. It's probably improved, but I remember being very disappointed a few years ago when the bloody thing wouldn't correctly infer the type of zip(). And that's ignoring the things that'll violate the specified type when you interface with the outside world (APIs, databases).
> anyone who actually cares about efficiency is using C.
Python is so much slower than e.g. Go, Java, C#, etc. There's no need to use C to get a better performance. It's also very memory hungry, certainly in comparison to Go.
Not sure, what you mean by thread-safe language, but one of the nice things about Java is that it made (safe) multi-threading comparatively easy.