Notes from the Meeting on Python GIL Removal Between Python Core and Sam Gross
lukasz.langa.pl
lukasz.langa.pl
The lone geniuses can only go so far. Maybe the problem is that Python can't quite decide what it wants to be because it's too many things for too many people already.
Is the concept of Python the language, as opposed to Python the ecosystem, valuable enough so that a Python that broke backwards compatibility with all the C extensions would be useful as its own multicore-capable runtime? PyPy seemed to think so for a while and now has gone hard in the other direction, reimplementing (faking?) a bunch of the CPython extension API so maybe this approach would never work. I don't know, but seeing things like:
> there is a large number of “dark matter” Python (and C extension) code out there that isn’t open-source. We need to be careful not to break it since it might not be feasible for its users to make required changes, or to report problems back upstream to us. In particular, some C extensions protect their own internal state with the GIL. This is a big worry, and might be a big hindrance to adoption of a GIL-free Python.
really make me wonder if the community as a whole would conclude the same on these critical sorts of decisions which shape the future of the language if they were put forward and not just made by a couple people in a closed meeting.
Would you prefer to support some weird arbitrary nameless closed source extensions, or have a multicore Python? This obviously depends on who you are and what you're doing, which leads us back to Python being too much for too many, but even here we can get a feeling for how many people do what with the language.
There's nothing wrong with staying on an older LTS version of Python. Let the people with the nameless closed-source stuff stick with that. The beauty of open source is that they can fork the older, GIL-ful version of Python and maintain it, if they like.
Multicore would be a tremendous boon to the language.
Let's hope it becomes a teachable moment.
Also it's important to remember that a lot of of material contributions to the community (either to the foundation, via jobs, or even open-sourcing part of their internal stack) might be coming from closed source in some way. It's not wise to ignore that & I think the core team is rightfully cognizant of needing to balance that (balance - not tip to one extreme or the other).
https://www2.eecs.berkeley.edu/Pubs/TechRpts/2006/EECS-2006-...
I'm astounded that a change which will release untold heisenbugs into the wild is being considered. It changes my view of Python. In terms of inducing subtle, silent breakage in existing code, it reminds me of this horrifying change from PHP 8:
https://www.php.net/manual/en/migration80.incompatible.php#m...
Sounds like they’re going to make it opt-in (or opt-outable) if they do merge it.
The only thing horrifying to me is that that behaviour was there in the first place!
I would be 100% behind those fixes, I assume from your response that you would not do them in the sake of backwards compatibility.
What would be your solution? To always have these idiosyncrasies in the language, or did you have a problem with how the fixes were implemented or rolled out?
Not sure I understand that. The proposal is not to simply get rid of the GIL, but to have a two-tier mechanism that ensures correctness with all the C source that uses the macros it should use and doesn’t mess with refcounts behind Python’s back (doing sketchy stuff usually ends up in pain)
There was an older LTS version of Python called Python 2.7. Last I checked people hated the transition and were still bitching about it in 2021.
My chief complaint was that there was a decent syntactic change for 0 benefit for me, and many Python users. F-strings, swapping str<->unicode, print function. All white superficial stuff, at least as far as my domain is concerned (data science).
It felt like “hey other languages are getting breaking changes, we should too”.
This is completely different. Single core speeds have not increased for years (decades?), any language with performance vaguely on the list must have an answer to multicore computation. I’d put up with a fair amount of pain for this.
Perhaps, dunno, web devs would complain that this change doesn’t help them, and is only a pain. That’s what I disliked in 2->3. I was told that I’m a dinosaur and should put up and shut up. Which eventually I did. But this is my answer to the naysayers this time.
Of course this might still fail in technical grounds but I’m hopeful, sounds solid.
I for one couldn't care less if some proprietary binaries fail on Python 3.11 or so. That's why we keep multiple versions around (at last company, I could only use up to 3.6 because that was the version in the Sacred CentOS AMI)
And, of course, a very critical piece of code was depending on a bug in regex that was fixed in 3.8 or so, and decided to break during a demo (where I was using 3.9 instead of 3.6).
Sorry to hear about your demo. Sounds not fun! At least that's in the past :)
Probably doesn't work across minor versions anyway, most stuff isn't built against the limited API.
Taking a highly conservative approach to breaking changes is absolutely not the same thing as being indecisive. The Python team has learned from experience how disruptive breaking changes can be.
I think this might be a misunderstanding of the nature of the event that these notes are generated from, unless I'm misunderstanding your objection. The point of this Q&A as I saw it was to explore the feasibility of the idea and fully flesh out the costs and benefits so that we can make informed decisions about how to proceed.
The "random interjections" are notes of caution about what trade-offs need to be made. For example, it is very easy to overlook "dark matter" code because we don't have access to it, but it's almost certainly the majority of Python code out there. It is also not a complete deal-breaker to say that some change could break unknown proprietary extensions — otherwise we'd never be able to change anything; the key is that the changes have to be worth it. A lot of that depends on details — if it's easy to update C extensions for nogil mode (even if they were designed without parallelism in mind), then making breaking changes to remove the GIL might not be so bad. If nogil mode requires that most C extensions totally overhaul their reference counting and C API usage and the changes require restructuring code rather than something that can be done with automated search and replace, that's a much bigger cost and will probably come with a long term fork of the ecosystem (which is a huge pain to deal with) and it might not be worth it.
Avoiding this sort of criticism will not make the underlying problems go away, and I think everyone involved understood that this meeting was intended to bring to light any objections that might guide the work towards ultimate resolution.
"On a personal level, we are impressed by Sam’s work so far and invited him to join the CPython project. I’m happy to report he is interested, and to help him ramp up to become a core developer, I will be mentoring him. Guido and Neil Schemenauer will help me review code for the interpreter bits I’m unfamiliar with."
I'm not sure if there is a common name for this particular source of discomfort, but that quote definitely contains a lot of it. I'm a historical contributor to the Python source repository, but something about the social structure of the project has changed significantly in recent years that would dissuade me from submitting changes in future. The focus in the statement above no longer feels like it is on the actual productive output of the project itself, and in previous years it wasn't like that, nor needed to be like that.
Reminds me of something like the minutes of a professional schmoozer's business lunch, rather than a technical meeting, or something like that. If you have ever seen a stray engineer at an event like this (or had the misfortune of being that engineer), this feeling probably captures the problem well. Whatever it is, I'd love to see less of it.
They are good at self-promotion and corporate backed. Real work in the Python world is done by the quieter types who don't strive to "be in charge".
I'm rather interested in why you feel that way however. Do you think your work, or the work of some specific project/person hasn't received the recognition it deserves ? Or that the attention of the community diverted from important work toward subjects that are more eye catching ?
I'm struggling to find that in the quote. What is the fraternity in question?
The choice of the word "fraternity" by the poster above was deliberate and seems to be being used to imply a more severe exclusion of others.
We don't even have to argue anyway. From a dictionary:
> [treated as singular or plural] a group of people sharing a common profession or interests: e.g. “members of the hunting fraternity”.
Frankly it’s very easy to google.
However, I've just looked it up in Merriam Webster, and it looks like this term can include female members, too, even thought the word and the related adjective have strong masculine associations. Wikipedia basically says the same, "Although membership in fraternities was and mostly still is limited to men, ever since the development of orders of Catholic sisters and nuns in the Middle Ages and henceforth, this is not always the case. There are mixed male and female orders, as well as wholly female religious orders and societies, some of which are known as sororities in North America."
Much to the opposite, just like many other big and important open source projects, they are actively trying to recruit, mentor and urge people to come as far "in" as possible. It takes a lot of work and dedication to get acquainted with a code base.
I don’t believe you. Python was never like this in the past, and has a long history of rejecting objectively good performance enhancements because the BDFL and friends wanted the implementation to remain a simple teaching example for students.
I don't think anyone would deny Python has changed since then, I think most notably in the years following the release of Django, Python becoming a go-to web language, and the size of the community exploding.
The economic pressures surrounding the benefits of gross’s changes will likely influence this more than any tears shed over subtle backwards incompatibility.
I believe it was Dropbox that famously released their own private internal Python build a while back and included some concurrency patches.
Many teams might go the route of working from Sam Gross’ work and if we see subtle changes in underlying runtime concurrency semantics or something else backwards incompatible that’s it- either that adoption will roll downhill to a new standard or Python core will have to answer with a suitable GIL-less alternative.
I for one do not want to think about “ANSI Python” runtimes or give the MSFTs etc of the world an opening to divide the user base.
Google also had their Unladen Swallow version, but it seems they lost interest at some point.
Perhaps Google/Grumpy could be updated to compile Python 3.x+ to Go with e.g. the RustPython version of the CPython Python Standard Library modules?
"Inside cpyext: Why emulating CPython C API is so Hard" (2018) https://news.ycombinator.com/item?id=18040664
Today, conda-forge compiles CPython to relocatable platform+architecture-specific binaries with LLVM. https://github.com/conda-forge/python-feedstock/blob/master/...
conda-forge also compiles PyPy Python to relocatable platform+architecture-specific binaries with LLVM. conda-forge/pypy3.6-feedstock (3.7) https://github.com/conda-forge/pypy3.6-feedstock/blob/master...
https://github.com/conda-forge/pypy-meta-feedstock/blob/mast... :
> summary: Metapackage to select pypy as python implementation
Pyodide (JupyterLite) compiles CPython to WASM (or LLVM IR?) with LLVM/emscripten IIRC. Hopefully there's a clear way to implement the new GIL-less multithreading support with Web Workers in WASM, too?
The https://rapids.ai/ org has a bunch a fast Python for HPC and Cloud; with Dask and pick a scheduler. Less process overhead and less need for interprocess locking of memory handles that transgress contexts due to a new GIL removal approach would be even faster than debuggable one process per core Python.
The approach focuses on functional programming, does away with extensions completely.
For that approach to be successful, a pure python implementation of stdlib in the transpileable subset of python 3 would be super helpful.
However, I would speculate part of PSF's hesitancy is likely specifically around the perceived violence that gross' "GIL-less" changes may incur to the runtime semantics' backwards compatibility.
PSF in particular has a responsibility here as well I feel in that CPython is arguably the working spec or standard from which these other implementations work and are defined.
Do you also strongly prefer languages with different underlying concurrency semantics? While stackless and pypy etc. are around and available and this could suggest the answer could be "yes" we've been lucky that they haven't fundamentally changed the experience of writing Python.
The possibility that a ton of libraries might now be able to use efficient multi-threaded execution where they were previously constrained to multiprocessing will be a landslide of changes on its own, and likely reminiscent of python 2 -> 3 compatibility if we have to preserve "two ways of doing things."
The Python ecosystem once fell apart not too long ago because it was supporting two versions where one moved from “print “ to “print(“ and these were incompatible and broke things such as doc tests.
There’s a reason that ppl strongly started advocating to hard pivot to 3: there was a very real chance that Python 2 could fork the ecosystem.
Incompatible concurrency semantics would be a much worse can of worms.
What popular languages fall into this category ? I can only think of C/C++ and JavaScript - both seem like terrible examples of languages that took forever to evolve (people still compile down JS to ES5). I'm not sure what the Java story is but I would argue it has been terrible at evolving the language as well.
I much prefer languages that have one implementation as a de facto standard, worked on by core team (eg. C#, Rust, TypeScript). Sure they might be a few random implementations - but the language is basically what the main compiler supports. Standards and specifications add so much overhead and I really don't see the value.
Not all, but many. Going down the list of most popular languages from https://www.tiobe.com/tiobe-index/ -
Python - CPython, PyPy, MicroPython
C - numerous (gcc, clang, msvc, tcc, etc.)
Java - https://en.wikipedia.org/wiki/List_of_Java_virtual_machines
C++ - gcc, clang, msvc, etc.
C# - Microsoft's version and mono
Visual Basic - probably only one implementation
JavaScript - https://notes.eatonphil.com/javascript-implementations.html
SQL - assorted dialects, not sure that counts
PHP - probably only one implementation
ASM - assorted dialects, not sure that counts
Classic Visual Basic - probably only one implementation
Go - only one version that matters, AFAIK
MATLAB - probably only one implementation
R - probably only one implementation
Groovy - probably only one implementation
Ruby - https://opensourcelibs.com/lib/ruby-implementations (Why are there so many in Go?)
Swift - probably only one implementation
Fortran - https://fortran-lang.org/compilers/
Perl - probably only one implementation
Delphi/Object Pascal - there are a decent number of Pascals historically, and FPC and Delphi are the big modern options
It might not matter much if Canonical or IBM decided to port a critical mass of open source extensions/packages. Then they could ship the new CPython in place of the old one and mention the differences in the release notes. With one or both throwing their weight behind it, it would gain significant momentum above and beyond the original project.
The lifeblood of Canonical and Red Hat / IBM is long term platform support for companies that want their most critical code to not break underneath them.
Even if the open source libraries get ported, there are still plenty of proprietary C extensions out there for which this would be a breaking change - the "dark matter" referred to in this post.
It would make zero sense for them to "throw their weight around" and _unilaterally_ break their customers' code if even the upstream devs didn't want to go through with the changes.
If users are gaining performance, they will bend over backward porting their code to this new version.
At minimum, I predict, all FAANGMULA would jumped in the bandwagon and create a pretty big ripple effect.
And Python is not a business with customers. It’s an open source volunteer project.
(Archived through google cache, so two layers of cache.)
1. Supposing this method doesn't currently crash under GIL python, would it be true that this method will also run without crashing on the non-GIL python interpreter?
2. Would it be true that the non-GIL python interpreter will introduce a race to this method (resulting in a runtime error) that didn't exist under the GIL interpreter?
Not that it's bad but it should be mentioned that it's a corporate initiative.
At least the "Brand Guidelines" PDF makes it clear that “PyTorch, the PyTorch logo and any related marks are trademarks of Facebook, Inc.”: https://pytorch.org/assets/brand-guidelines/PyTorch-Brand-Gu...
Also a small hint in `CONTRIBUTING.md`: https://github.com/pytorch/pytorch/blob/master/CONTRIBUTING....
Different people want different things I guess...
It's one language, but why can't I tune my VM to my needs? I can't imagine Java not letting users tune their GC.
System python at least is generally only recommended to be used for system libs, and that's a relatively supportable set. Developers use virtualenv's and their own specific interpreter, but it would certainly move the needle on what language people were by default scripting and thinking in.
> It all depends on how well the community adapts C extensions so they don’t cause downright crashes of the interpreter. Then, the remaining long tail is community adopting free threads in their applications in a way that is both correct and scales well. Those two are the biggest challenges but we have to be optimistic.
Even if it's 10% of the mess the path py2->py3 was, it still worries me. I hope I'm wrong and it's much less than that (for the fatal cases ATL, and similar/non improved perf for the rest)
But on the other hand, isn't simple to just dedicate a script per core?
What python commitee considered as infeasible was almost done by a lone hero. Since the previous decision to change the format of the print function (that no one asked for) broke everybody's code for no reason and took ten years to be adopted, they will not push (for the one change everyone wants) for the foreseable future. Although it does not seem to be that impossible after all.
I am glad they will invite Sam, the lone gero we needed and hope he will be given some ownership of the task and not get him swamped with commiteeisms through a embrace, not extend and extinguish. He is on a success path, the commitee is in no path at all.
Just put a timeline and call it a fail if don't succeed and quit avoiding it through "discussions". We got it, it's not planned for the X.XX+1 version, each version.
Impressive work (2 years working full time knowing that it might never be merged is incredible)