The Changing "Guarantees" Given by Python's Global Interpreter Lock
stefan-marr.de
stefan-marr.de
One study [1] in the US released in 2020 found almost 90% of people admitted to speeding, but I don’t think anyone would say that speeding is now approved by the authorities and consequence-free.
[1] https://www.thezebra.com/resources/research/speeding-car-ins...
Depends on the definition of "correct".
One way to define "correct" is "as defined by the standard". However, standard serves a purpose - to maintain interoperability between components, so the other, deeper, way to define "correct" is "interoperable with the ecosystem".
The official standard can argue that relying on the implementation detail is incorrect because they didn't specify it explicitly, but they would go against the grain of the rest of the ecosystem. Which reminds me of an old joke:
An old woman is watching the news. She sees a news report saying there is a car driving in the wrong direction on the highway. So the old woman calls up her husband.
Old woman: be careful on the highway dear, there is a crazy driver on the highway driving the wrong way!
Husband: It’s not just one car, it’s hundreds of them!The "standard" is a combination of PEPs, the Python docs, and the CPython implementation.
Implementation details are language features, because implementation details are the standard. Python programs that rely on the GIL are not "wrong". Such programs are relying on clearly-documented features of the system that they use.
Is it "wrong" to use GCC-specific features in a C program, if you know that you only intend to target GCC?
The problem here is: Implementation details are not guaranteed to be stable. They can change with every release.
So if I write code that relies on any particular implementation detail, it may be correct today, and may be wrong two weeks from now, even though I didn't change anything.
On the other hand, Linux kernel developers have also been burned by GCC changing what it does regarding undefined behavior (the infamous not-so-corner cases of C/C++). They expect GCC to be sort of like an assembler, which does the most straightforward thing in case of ambiguities. Instead, GCC is an optimizing compiler that is expected to deliver high-performance code for general-purpose programs. It treats undefined behavior as opportunities to enact optimizations. What it actually does is also usually documented, but also here different people have different expectations about what that implies.
The GIL is an implementation choice that was put in place by Python's developers for simplicity's sake. It is important to be aware of it because it has huge performance implications, but it is unwise to rely on it for semantics. Especially since it has not exactly been a secret that various parties would eventually like it to be removed. Anyways, the GIL has very little to do with Python's semantics. User code is racy with or without the GIL, which is actually the point of TA.
Just hard as hell to change
Hyrum's Law says the opposite:
"With a sufficient number of users of an API, it does not matter what you promise in the contract: all observable behaviors of your system will be depended on by somebody."
Named after some guy called Hyrum who worked / works at Google.
I would say jaywalking in NYC is a much better example because breaking the law doesn't kill others except in rare cases. Jaywalking is illegal. It's also not enforced for all intents and purposes. Every attempt to enforce it has led to such loud public outcry that it was stopped.
e.g.; there are C-compilers that usually zero most allocated structures. That's an implementation detail however, not a feature of C. C doesn't guarantee you zeroed memory anywhere, and code that assumes otherwise is just one compilation away from a desaster.
This doesn't count?
> Consequently, if you were coming from Mars and tried to re-implement Python from this document alone, you might have to guess things and in fact you would probably end up implementing quite a different language. On the other hand, if you are using Python and wonder what the precise rules about a particular area of the language are, you should definitely be able to find them here. If you would like to see a more formal definition of the language, maybe you could volunteer your time — or invent a cloning machine :-).
Depends on the question. If your question is, say, "how should I implement the built-in types", yes, the language reference won't answer that question. But if your question is "what does my implementation have to do to count as an implementation of the Python language", then yes, the language reference does answer that question--since the language reference is what defines the Python language.
To me that is a "specification" of the Python language. It's not a specification of the implementation of the language, but why should that be required for something to be considered a language specification? The whole point is to specify what is required to define the language without specifying every detail of the implementation.
Yes, I know that statement is in the Introduction, but I think it's rather ill-considered. If my implementation is consistent with the language reference, on what grounds would someone claim it was not an implementation of Python but a "different language"?
I would like to see some indication from the people who say that the Python Language Reference I gave a link to does not qualify as "language specification", of what would qualify. Specific examples would be nice.
I don't see why not. The "Introduction" section specifically mentions different implementations and distinguishes implementation details, which can vary by implementation, from the language reference itself, which defines what every implementation has to meet to be considered an implementation of the Python language.
> my point in the GP was that something distinct called “the Python Language Specification” doesn't exist
And that's the point I'm disputing; AFAIK the language reference I linked to is that something distinct, even if it isn't called a "Language Specification" but instead a "Language Reference". Either way it defines what the Python language is.
Not sure why there is an argument here.
No, we don't. I have already said explicitly that I think the Python Language Reference is such a thing. (I am ignoring quibbles about it being called a "reference" instead of a "specification".) If you think it isn't, why? What does count as a "language specification" in your view? Do you have any specific examples that you can contrast with Python?
Yes, there is:
https://docs.python.org/3/reference/index.html
This specification does not mention the GIL anywhere, which means it is, as the GP said, an implementation detail. Other implementations of Python that do not have the GIL are still "Python" implementations because they meet this language specification.
I mean, I can see the point if there are multiple commercial competitors in the market, as there is with C/C++, or if the implementation is proprietary and the users want to avoid vendor lock-in.
But the Minimal BASIC of ANSI X3.60-1978 never did catch on for any of the BASICs I used in the 1990s, and the Full BASIC of ANSI X3.113-1987 was a flop, so clearly it's possible to put a lot of time into a standard only to have it be irrelevant.
Python does not have a specification or a standards document, it has a reference that describes how Python happens to work and the reference is a great resource for people to familiarize themselves with the language but it should simply be clear that its purpose is to reflect the existing state of the language rather than to specify how Python works.
I've already given my answer: they all meet the language reference I linked to (yes, the word "reference" appears in its title, not "specification"; that's just another quibble).
What is your answer?
Python is a lot more like LISP than it is C++. There are many flavors of Python, from MicroPython to GraalPython to Cyston and literally dozens of them. They most certainly do not all meet the language reference you linked to and they all have quirks here and there.
A language does not need a specification in order to exist or to have a name. What matters is that people use it and get work done with it, and ultimately if looks like a Python and quacks like a Python, then it's fine to call it a Python.
Here's another fact: Python doesn't have its own typeface either. Yet that fact is hardly germane to the thread.
A prescriptive document is a formal specification. The Python Language Reference is a less formal specification. It's still a specification.
It is also not a single document, but then again even a formal specification may incorporate other specifications by reference.
As my position is that the CPython implementation is the reference implementation for Python and as the GIL is an integral part of that implementation, then the GIL does form a part of Python's semantics regardless of whether the reference documentation mentions it or not.
The Python reference documentation is not a specification and it's not intended to be one.
You're quibbling. By your definition, no programming language has a "specification". Every language has implementations that do things that aren't explicitly described in any document.
Some languages do not, such as Python and Rust.
The ISO C++ committee even uses quite strong language about the C++ standard and makes it a point to differentiate between the C++ specification and C++ references:
>The standard is not intended to teach how to use C++. Rather, it is an international treaty – a formal, legal, and sometimes mind-numbingly detailed technical document intended primarily for people writing C++ compilers and standard library implementations.
Same question for Java and C#.
If the answer to all of these questions is "no", as I believe it is, on what grounds do you claim that these languages have a specification, while Python and Rust do not?
The grounds that I claim is that you can read them, here they are:
https://docs.oracle.com/javase/specs/
https://isocpp.org/files/papers/N4860.pdf
Note what the actual C++ standard states, and I quote:
>This document specifies requirements for implementations of the C++ programming language. The first such requirement is that they implement the language, so this document also defines C++. Other requirements and relaxations of the first requirement appear at various places within this document
The number of Google hits I get when I search on "C++ implementations that do not meet the language specification" does not seem to support this claim.
> Note what the actual C++ standard states
But your position with respect to Python is that it doesn't matter what the "standard" states because the actual definition of the language is in its reference implementation. Why are you now shifting your ground?
That is an incredibly weak and ill-informed method for vetting technical matters.
>Why are you now shifting your ground?
Python doesn't have a standard.
So is your claim that, because a "standard" says something and claims it's a "specification", that is sufficient to guarantee that every implementation does exactly what the "standard" or "specification" says.
> Python doesn't have a standard
Again, you're quibbling. It has a language reference, which I linked to, and which is what is used to determine what implementations count as Python implementations.
At this point I don't think we have enough common ground for a useful discussion.
It's the other way around, a standard specifies the requirements of a conforming implementation, and then an implementation can claim conformance with respect to that standard. With C++, many well-known implementations do claim conformance to the standard such as MSVC, GCC and Clang. With Python no one makes such a claim, not even the CPython reference implementation, because there is no standard to conform to. On the contrary what you do find are implementations that try to be compatible with CPython reference implementation specifically, such as PyPy and Pyston but no one claims to conform to the Python reference manual. The reference document simply describes how some Python implementations happen to work, but it does not prescribe how a Python implementation is required to work in order to be in compliance.
If tomorrow CPython decides to add a new feature to the language, then the reference will be updated to include this change because the reference is a reflection of the implementation.
If tomorrow the C++ specification changes, then it's GCC or any compiler that claims conformance that will update to include that change because with C++ it's the implementation that is a reflection of the specification.
That's the key difference between a descriptive document and a prescriptive document. This may seem like a quibble as you put it, but quibbles can lead to multimillion dollar lawsuits as Microsoft learned in the 90s when they claimed to have an implementation of the Java specification and then were sued by Sun Microsystems because Microsoft's implementation did not in fact conform to the Java standard.
Microsoft ended up paying 20 million dollars over this so called "quibble":
https://en.wikipedia.org/wiki/Microsoft_Java_Virtual_Machine...
Others object saying that the Python reference document is what specifies its semantics, not the reference implementation.
My position is that both the CPython reference implementation and the reference documentation are valid sources that document Python's semantics and that they both can and should be used.
Implementation details are documented to allow performance optimizations and to give insight into why certain things are how they are. Therefore, users can reasonably expect that the vendor won't cause performance regressions for existing code. However, it is unwise to derive semantics from them, even if they are technically documented in that way.
One of the biggest disadvantages of relying on implementation details is that it makes it way more troublesome for the vendor to maintain and improve the product.
Anyways, the GIL and the presence of possible concurrency bugs are completely orthogonal things as the GIL has always only served to prevent corruption of runtime data structures, not of user code.
And my position is that it is not because there are other implementations of Python that do not have it, and which everybody agrees are implementations of Python.
Not only that, but the very language reference that I linked to explicitly distinguishes CPython implementation details from the language itself. So the Python dev team does not appear to agree with your position.
What would it mean to have a language standard? A publication from ISO or ECMA?
I ask because the Python Language Reference at https://docs.python.org/3/reference/index.html seems to be a (terse) language standard. Among other things, it highlights some of the things which are implementation defined, rather than language defined.
It's an implementation detail, and relying on those for functionality is a great way of getting ones code to break.
Assuming sequential consistency is pretty broken. Assuming acquire/release atomicity is much more reasonable. Assuming "at least relaxed" is outright mandatory.
Writing to spec is very rare in my experience. You usually learn the spec once someone tells you your code doesn't work on XYZ.
That's a wild one
from FAQ:
Q: But isn't a language that deletes code crazy?
The library is documented, and the syntax is also documented, but the semantics themselves are not.
The language reference at https://docs.python.org/3/reference/index.html "describes the syntax and “core semantics” of the language.".
There has been a distinction between Python-the-language and CPython-the-implementation ever since JPython back in the 1990s. For example, reference counting is a CPython implementation details. The reference manual says only:
> Objects are never explicitly destroyed; however, when they become unreachable they may be garbage-collected. An implementation is allowed to postpone garbage collection or omit it altogether — it is a matter of implementation quality how garbage collection is implemented, as long as no objects are collected that are still reachable.
(Quoting https://docs.python.org/3/reference/datamodel.html )
A standards document or specification's purpose is to "prescribe" how a system shall work in order to be compliant as opposed to a reference which documents how systems currently happen to work.
For example, the following operations are all atomic (L, L1, L2
are lists, D, D1, D2 are dicts, x, y are objects, i, j are ints):
L.append(x)
L1.extend(L2)
x = L[i]
x = L.pop()
L1[i:j] = L2
L.sort()
x = y
x.field = y
D[x] = y
D1.update(D2)
D.keys()IOW, the python implementation is the standard, and therefore any code that relies on an implementation detail in the reference implementation is, by definition, correct.
That's not quite strictly true, Python does have a documented "Language Reference": https://docs.python.org/3/reference/index.html
If there is a contradiction between the Language Reference and CPython then one, or both, of them needs to be updated and it's treated on a case by case basis.
If an alternative Python implementation follows the Language Reference but chooses different details outside it, that doesn't stop it from being "Python". Of course practically speaking most alternative implementations are incentivized to closely follow CPython.
For a sufficiently popular project there are often cases where users (other projects) rely on existing observed behavior even if it not guaranteed by documentation/specification.
If major libraries and user code are constantly incorrect, but work because of an implementation detail, then removing that implementation detail becomes extraordinarily difficult, verging on impossible.
It would be like retrofitting your language to distinguish valid unicode text from arbitrary byte strings, when you'd previously treated them as equivalent.
a more interesting example is something like this:
# setup
l = []
# thread A
l.extend([1, 2, 3])
# thread B
l.extend([4, 5, 6])
is the resulting list always within the set of [1,2,3,4,5,6] or [4,5,6,1,2,3] ? or are the two sets of numbers randomly interleaved in the list? or if the GIL is removed does the interpreter segfault (I'm pretty sure this latter will not be the case for GIL removal but I don't understand the gil remove plan very much yet).Edit: before people jump in and correct how the above is a bad idea anyway, it's not like I'd ever do the above and expect anything but disaster. This is more of a thought experiment to understand what GIL removal is going to do.
> or if the GIL is removed does the interpreter segfault (I'm pretty sure this latter will not be the case for GIL removal but I don't understand the gil remove plan very much yet).
I haven't looked into the plan in detail either, but presumably not, that would be nuts. My understanding is that they're going to replace the GIL with locks on the objects themselves (your list `l` in this case). This is why in all the tests single-threaded performance suffer, you have to take and release more locks if you don't have a GIL, and the objects themselves grow larger as well.
Because the iterator can be pretty much arbitrary Python code it seems to be a bad idea to guarantee extend() to be atomic, as you don't want to hold the list mutex while calling out into user code.
A surprising amount of the CPython stdlib is just pure Python code.
Part of the integration work will be to better document the thread-safety guarantees, but there is still a lot of work to do before we get there.
request_id = self._next_id
self._next_id += 1
To think that is thread safe is just naive. Once you understand the potential problem you can look at the code and ask "why shouldn't this be unsafe?" Since there is nothing explicitly preventing the problem. After reading TFA (lazily I admit) I still don't know why that code is thread-safe with the GIL.
Python is fun and often forgiving, but a bunch of people who got lucky (because they were never taught about the hazards) are going to learn some new stuff with no-gil. I think it's a long overdue change and worth the (single thread) performance hit and bug surfacing phase.
On the other hand, it is safe from a memory-access perspective; the read from `self._next_id` will never dereference a partially-mutated invalid pointer or read a partially-mutated value.
I’m happy people are working on removing the GIL, but As a professional python dev for about 5 years now I have literally never had a problem where the GIL was a limiter. Although I just make web apps so maybe I’m not the target audience.
If you write a lot of code that parallelizes over data you will hurt all the time because of the GIL.
If you have worker processes that do something on the CPU, and the results need to be collected and processed further in some other process, you now need to pickle the data to copy it around. That can get really slow.
I understand a lot of python users don't do that kind of thing, but it's a real problem. I'm happy that the python community seems to be slowly beginning to take this seriously after decades of just claiming the GIL isn't a real issue.
I've added some text for clarity
With Python being the language of choice for ML workloads I guess it’s more common to have the CPU be a bottleneck. It seems cool they’re making an option to turn it off for those use cases.
It seems like Python could maintain a GIL compatibility option to preserve the current/old behavior for legacy code.
It shouldn't be anywhere near time-critical and or low-latency use-cases.
Python is functionally the modern BASIC, and included many of the same design trade-offs for usability.
Don't get mad, it is true... and we know it. =)
Thusly moving to multiprocessing and dealing with the lack of shared memory issues, with managers.
When/if the GIL goes away, good riddence.
Java 1.2 was branded as "Java 2", so you had the J2SE (Java 2, Standard Edition) and related J2ME and J2EE (Mobile and Enterprise, respectively) platforms. The "Java 2" moniker was dropped in Java 5, which was the largest rewrite of the language since, adding generics, sane memory model, annotations, etc., all in the same language revision.
3.10 and 3.9 are perfectly mechanically comparable (meaning one can write a program to deterministically compare them and return their relative order), just not with default numeric ordering (then again they're not numbers, they are composite values that are comprised by numbers) or naive string based ordering.
If we wanted trivially comparable with regular numeric ordering we could have incremental numbers as versions. 1, 2, 3, ...
And if we wanted string ordering (as with usual filesystem listing sorting with no extra flags to treat as numbers), we could have fixed length padded parts: 00001.00045.
Not sure if the latter is used, but some software does use the first.
> If we wanted trivially comparable with regular numeric ordering we could have incremental numbers as versions. 1, 2, 3, ...
Yes, and then we would be back to the day when the version number gave me zero information about what changed, and how that affects compatibility with existing code.
There is a reason semver is used across the industry by now.
Hence the whole "3.10 and 3.9 are perfectly mechanically comparable (meaning one can write a program to deterministically compare them and return their relative order)" part in my comment you perhaps missed.
>Yes, and then we would be back to the day when the version number gave me zero information about what changed, and how that affects compatibility with existing code.
Not that it's any better now with semver though: in practice the semver works 95% of the time, and give just a false sense of comfort at the other 5%. You update, and things still break, despite the semver promise.
I didn't miss any part of your comment. The line you quote was as a reaction to the alternatives presented immediately after that part.
Because why would I sacrifice the advantages of semver for a minor convenience to the programmer of a sorting function?
> in practice the semver works 95% of the time
Which, based on nothing but my gut feeling, is a lot better than the 30% of the time any other versioning scheme works, where the only way to be reasonably sure that an update would not break my code was to diff the library (if the thing is open source), or read all the documentation, and pray to Zeus that it's complete and accurate.
If it wasn't already common, it would probably become common shortly after the first time a big respectable company hit the "oops what comes after 9" problem and decided on dot-separated-integers rather than significant digits :)
It's basically semantic versioning, that is a hierarchical split based on levels of change (major = new release with possibly big breaking changes, minor = some incremental update version within the same release) and so on.
Who ever thought version numbers are decimals and why? The "." appears as a separator on all kinds of strings in software (filenames, domains, and IPs probably the most common ones).
Never assume that version numbers are decimal values. It’s more obvious when you see the full version triple (3.13.0 for example) that it’s not a single number, but the abbreviated version numbers can some times look like a decimal value. You should never compare version numbers as decimals in modern software unless you’re absolutely sure that’s how the project is structured.
Because otherwise it's perfectly comparable - and quite common.
Most FOSS uses a variant of <major>.<minor>.<patch> (here just major (3) + minor (13), minor meaning "release within the same major Python version").
Is there no great architect in the sky? Is there no software god after all, looking down, punishing sloppy engineers and granting blessings to thoughtful engineers? How else to explain this injustice of sloppy engineering eating the world (to say nothing of JavaScript)?
But consider that maybe it wasn't a sloppy design at all. For decades, the explicitly stated philosophy of the CPython development team was to prioritize simplicity of implementation over performance. I don't think anyone ever envisioned Python becoming the wild success that it is today.
That is, the GIL wasn't sloppy at all. It was perfectly reasonable and pragmatic decision that made sense given the tradeoffs of the time.
Watch people's commentary here when talking about the good and bad of various technologies.
There's no 'bad' tech, just tech with lots of users. Or the tech is good because you can hire for it. Because it has lots of users. Or it has a rich ecosystem, because it has lots of users.
Read the advice given by those who tell you to get to market first instead of polishing the tech.
We're the users of that tech which went to market unpolished and gathered all the users.
Especially when you consider this over the whole of Python's lifespan, which very, very firmly includes many years in which multicore was simply not a thing, followed by some years where it was a thing but it didn't work very well anyhow at the OS level so who cares what Python does with it.
It is not as if back when it was put it the choice was either to use a GIL or to correctly write a multithreaded interpreter and fix all the 3rd party libraries at the time for exactly the same cost. The latter option was orders of magnitude more expensive, and harder then than it is now, with better tooling and more collective developer experience. The choice of not using a GIL, rather than being some sort of nirvana that we could just be in if they hadn't chosen poorly 15 years ago, could well have killed the language. We don't really know. I do know that a programming language that just sort of breaks every so often when you use threads and there's absolutely nothing you can do about it from the Python level is not a very appealing proposition and it's hard to know how badly this could have hurt the language.
And Python of all the languages now has a well-justified fear of breaking everything and demanding that everyone upgrade.
So, to put this in a nutshell, if you believe the GIL is simply bad and should never have been an option, you have a very immature understanding of software engineering, especially in the light of being the leader of a very very large community who will be impacted by your decisions. It may not have been the only choice, but it was a good one, and regardless of what decision was made 15 years ago there would be some consequence to deal with now. No programming language community can be expected to get everything right in 2003 that the people of 2023 will want any more than we can expect any current programming language to be the perfect programming language of 2043 right this second.
What are you even complaining about? What is your point?
Python, JS, C, Bash aren't even particularly great at the problems they solve, but they succeed mostly on inertia (it's where all the libraries are, it's what people know) and occupying developer mindshare.
They are full of obvious design mistakes; things that not even the creators of the language (nor any of its users) can defend, yet those languages are used infinitely more than languages that eschew those mistakes. Why? Because they solve problems people have.
If this sounds terrible to you, the good news is that there is a tonne of low-hanging fruit in the programming language design space. Consider that most developers know nothing of sum-types, or eschew the idea of typing entirely. Consider that most developers see no fundamental problem behind having to venv or dockerise software lest it bitrot over a month. Consider that programmers actually use bash.
These terrible, obviously broken tools are somehow the most pragmatic things we actually have. The fruit is low-hanging; the door is wide open, if you wish to grab it.
I would argue the "problem" that Python really solves is the amount of engineering effort required to read and write code for common software use cases. Ie, it's purpose is to help developers write better code faster and easier than other languages, while execution speed has usually had a lower priority. In that framing, it's great at solving the problem of development time and old code being hard to maintain, and that's why so many engineers like myself love it.
This is basically the only part of your post I disagree with, for reasons you pointed out yourself.
Low-hanging-fruit would imply that it is easy to displace these flawed languages with something better. And you have made a beautiful argument for why that is very very very difficult indeed.
That's a good thing! Models are fit on data, data doesn't fit to models. This is like when people learn elementary music theory then go analyze some actual composer and it doesn't fit the model at all. Well kiddo the problem is that "music theory" is simply a model, a model people created after training some very limit set of musical data, everything outside of that data will probably behave different and you'll have to change your models.
If your software engineering model predicts Python would be unsuccessful, but there is evidence that Python is successful, this simply means your software engineering model is unpredictive and therefore must be revised.