Python stands to lose its GIL, and gain a lot of speed
infoworld.com
infoworld.com
Performance doesn't come from any one quality, but from the holistic goals at each level of the language. I think some of the most frustrating aspects from the history of Python have been when the team lost focus on why and how people used the language (i.e. the 2 -> 3 transition, though I have always loved 3). I hope that this is a sensible optimization and not an over-extension.
At a higher level, Python is getting serious about performance. But this gives both flexibility and performance.
Even if this will be a major breaking change to Python, it'll be worth it for us.
> The overall effect of this change, and a number of others with it, actually boosts single-threaded performance slightly—by around 10%
Then it sounds like having the cake and eating it too (optimism). Although my experience keeps nagging at me with, "there is not such thing as a free lunch" (skepticism).
https://lwn.net/Articles/872869/
The no-GIL version is actually about 8% slower on single-threaded performance than the GIL version, but the author bundled in some unrelated performance improvements that make the no-GIL version overall 10% faster than today's Python.
I would love to be proven wrong but I am skeptical.
> The resulting interpreter is about 9% faster than the no-GIL proof-of-concept (or ~19% faster than CPython 3.9.0a3). That 9% difference between the “nogil” interpreter and the stripped-down “nogil” interpreter can be thought of as the “cost” of the major GIL-removal changes.
[1] https://docs.google.com/document/d/18CXhDb1ygxg-YXNBJNzfzZsD...
It’s interesting because I think many people (myself included) would be far more interested in the perf patches than the GILectomy.
> though, as Guido van Rossum noted, the Python developers could always just take the performance improvements without the concurrency work and be even faster yet.
Why be 10% faster single threaded when you can be 20% faster single threaded!
When you carry a heavy suitcase filled with lead and you drop it, things get lighter for free. You paid for it by carrying the damn thing around with you for the whole time.
Well, yeah, someone had to make the changes. That's the cost that was paid.
You can get a mass-produced machete that is cheaper and higher-quality than a 7th-century sword. It's easy for one thing to be better than another thing across several dimensions simultaneously. That's why certain technologies go out of use -- they have negative value compared to other technologies. But that has nothing to do with the principle that there's no such thing as a free lunch.
The crux of the issue (as I understand it) is that the GIL absolves the Python interpreter of downstream memory access control. You can replace the GIL with memory access controls of various strategies, but the overhead of that access control is just that: overhead. In a multi-threaded program the concurrency gains should outweigh that overhead, but in a single-threaded one it's just extra work that wasn't being done before.
Which brings us back to no free lunch. It turns out that the claim "10%" faster without the GIL is actually a result of Gross (GIL removal author) doing a multitude of unrelated performance improvements. These performance improvements increase performance enough that the performance of single-threaded no GIL code (with overhead) is ~10% higher than today. But as Guido pointed out, the core developers could upstream the performance improvements without the GIL removal:
> To be clear, Sam’s basic approach is a bit slower for single-threaded code, and he admits that. But to sweeten the pot he has also applied a bunch of unrelated speedups that make it faster in general, so that overall it’s always a win. But presumably we could upstream the latter easily, separately from the GIL-freeing part. [2]
[1] https://docs.python.org/3/faq/library.html#can-t-we-get-rid-...
[2] https://lwn.net/ml/python-dev/CAP7+vJJ1hzXiyDwVs6-eXed+DtodH...
Tell me how "faster in general" doesn't make you suspicious about a free lunch.
Put another way, if the performance improvements were upstreamed without removing the GIL the resulting performance increase would be ~20% instead of just ~10%. Which is what Guido was getting at in the quote I cited. Assuming the benchmarks to be true for the moment, this means that removing the GIL on this PoC branch is a 10% performance hit to single-threaded workloads.
Can't see Python getting there unless we go to Python 4 which, given the fiasco that was Python 2->3 is probably never gonna happen.
Might as well wait for Julia to improve its TTFP then to hope for a Python 4.
The GraalPython implementation of Python 3 is built on the JVM, which is a fully thread safe high performance runtime, and Graal/Truffle provide support for speculation on many things. For pure Python it provides a 5-7x speedup already and the implementation is not really mature. Although at the moment they're working on compatibility, in future it might be possible to speculatively remove GIL locks because you have support for things like forcing JITd code to a safepoint and discarding it, if you want to change the basic semantics of the language.
It resolves a few big problems Jython had:
- GraalPython is Python 3, not Python 2
- It can use native extensions that plug into the CPython interpreter like NumPy, SciPy etc. The C code is itself virtualized and compiled by the JVM!
But last time I checked, pypy had much better performance than Graal, even though TruffleJS (javascript interpreter built on the same model as graalpython) has comparable performance to the v8 engine for long running code. Though the latter is the most actively developed truffle language, let me add that.
And it is also about flexibility. What I love about Python is the simplicity, and let's be honest, multiprocess anything but. Especially if you fall into one of the gotchas (unpickable data for example).
As for multiprocess, I currently have 150 python process running on the work cluster. Each doing their bit of a large task. The heavy lifting is in a python library but it’s C code. but it’s actually not bad performance wise and frankly wasn’t to bad to code up. I think for my use case threads would make it harder.
Maybe Im liking python more over time..
JavaScript interpreters that people actually use, have a JIT inbox.
Both Java and Javascript are ultimately compiled into machine code, through JIT. And this matters because Python doesn't have JIT.
Even Ruby, that is historically a way slower language, is gaining JIT nowadays. Python has got no excuses.
It starts to become an issue when you have built a few well-performing subsystems and now want them to run together and interact. With the GIL, your subsystems are suddenly not performing as well anymore. Without the GIL, you can still get good performance (within limits of course).
Performance referring here to throughput and/or latency (responsiveness).
What does it mean? how is python different here than Java/C#?
Mypy and other static analysis tools are becoming more common in part because IMO they basically require you to stop and think about your "Pythonic" dynamic patterns (containers of mixed element types, duck typing of function arguments, mutable OOP etc.), and often realize that they are a bad idea.
So in some way we are hampering multithreading to support programming constructs that are mostly used to make python flavoured spaghetti, especially in the hands of beginners and non-programmers who are encouraged to learn Python...
I'm not sure Python is fixable at this point. Oh well.
Is it truly worth it just to avoid some memory overhead? Or is there some other windows specific thing that I'm missing here?
#2 need not be true; e.g., the approach proposed here is transparent to most Python code and even minimized impact on C extensions, still exposing the same GIL hook functions which C code would use in the same circumstances, though it has slightly different effect.
The idea of a process-local cache (or other data) shared among all worker threads is a different story. Along with reduced memory consumption, I see this as one of the bigger advantages of threaded app servers. However, preforking multiprocess servers can always use shmget(2) to share memory directly with a bit more work.
lol, you're so deep into python stockholm-syndrome "don't share anything between threads because we don't support that at all even a little bit" that you don't even realize that connection pools exist. Instead of holding a connection open per process, you can have one connection pool with 30 connections that services 200 threads (exact ratio depends on how many are actually using connections, of course). literally everybody "shares a single DB/RPC connection across multiple threads" (or at least shares a number of connections across a number of threads), except python.
and yeah you can turn that into yet another standalone service that you gotta deliver in your docker-compose setup, but everybody else just builds that into the application itself.
The GP mentions connection pooling literally three sentences later.
> literally everybody "shares a single DB/RPC connection across multiple threads" (or at least shares a number of connections across a number of threads), except python.
Right, but multiple ≠ many. You're discussing the former. GP is discussing the latter.
A tangent but I find it amusing to contrast the perpetual Python GIL debate with all the new computation platforms that claim to be focused on scalability. Those are mostly single threaded or max out at a few virtual CPUs (eg "serverless" platforms) and there people applaud it. There people view the isolation as supporting scalability.
Moreover, unless thread pinning is enforced, a given thread will bounce around between different cores during execution, so the number of distinct L3 caches in action will not be constant.
Of course you have the same story with memory, accessing another thread's memory is slower if that thread is on another CCD.
TL;DR NUMA makes life hard if you want to get consistent performance from parallelism.
like processing web requests?
A lot of people will argue that its not worth optimising for, that programmer time is more important and expensive. This may be true at the start, and is probably even worthwhile when you need to get a product out fast to test the waters in a startup, but I've worked in multiple companies that have spent significant developer time to try to reduce their cloud infrastructure costs. At this point, having your language and famework be able to make good use of the available hardware really can make a real difference to the overall performance and therefore cost, both hardware cost and engineering time to optimise it later.
The productivity gap between languages that optimise for development and those that try to at least somewhat optimise for runtime really isn't that large anymore nowadays. Even modern Java is quite productive now, compared to 10 or 15 year ago. Outside of a startup trying to find product market fit building their MVP's super fast to see what works, I think its usually worth spending a bit more on up front development time, a one off cost, to reduce the recurring infrastructure cost.
Of course, it takes a lot more than just a language that supports multithreading to do this, but everything the language and the libraries/frameworks you use do to help you is helpful. I'd rather have a tool that won't get in my way later, that gives me lots of room to grow when performance starts to become an issue, than one where I need to invest in significant and painful development time (based on personal experiences at least) later. This is one area where Go seems to shine, perhaps Rust too, although I have not tried any web backend dev in Rust yet so don't know how productive it would be.
Rust got it pretty good, but they designed a significant part of their language so that multithreading could be done well. Python did almost the opposite.
Is my PHP knowledge out of date? Is the this the current state the art: ? https://www.php.net/manual/en/parallel.setup.php
https://github.com/php/php-src/tree/master/TSRM
I have written about this recently
> Swoole is a complete PHP async solution that has built-in support for async programming via fibers/coroutines, a range of multi-threaded I/O modules (HTTP Server, WebSockets, TaskWorkers, Process Pools) and support for popular PHP clients like PDO for MySQL, Redis and CURL.
Basically I love peanut butter ice cream (Python) I’d just like it even more with sprinkles.
I think a JIT would be the best possible improvement for CPython as far as speed is concerned. Though I can imagine there are plenty of people doing processor heavy stuff with c extensions that would benefit from sharing memory. So from their perspective removing the GIL would be a better improvement.
So basically a JIT would help every Python program, and removing the GIL would only help a small subset of Python programs. Though I'm just happy I get to make a living using Python.
Edit: This was in the back of my head, but I didn't mention it, and it would be unfair to dismiss. A JIT does slowdown startup time, so for short programs that are finished running quickly it may make things worse. Though I suspect it would be easy enough to have a value to turn off the JIT at the start of the program.
What if the "Global Interpreter Lock" needs to be removed for JIT? I put that in quotes to highlight it because AFAICT no compiled (or JITed) language has such a thing. I think it functions differently than regular stuff like critical sections.
It doesn't. PyPy has a JIT and a GIL. The JIT just compiles the byte code to native code before running, along with some tricks that are over my head.
Edit: Here is PyPy's FAQ about it
https://doc.pypy.org/en/latest/faq.html#does-pypy-have-a-gil...
The compiled code polls a global or per-thread variable as it runs (but in a very optimized way). When one thread tries to change something that might break another thread, the other threads are brought to a clean halt at specific points in the program where the full state of the abstract interpreter can be reconstructed from the stack and register state. Then the thread stacks are rewritten to force the thread back into the interpreter and the compiled code is deallocated.
The result is that if you need to change something that is in practice only changed very rarely, instead of constantly locking/unlocking a global lock (very, very slow) you replace it with a polling operation (can be very fast as the CPU will execute it speculatively).
However, this requires a lot of very sophisticated low level virtual machinery. The JVM has it. V8 has it. CLR has a limited form of it. Maybe PyPy does, I'm not sure? Most other runtimes do not. For the Python community, very likely the best way to upgrade performance would be to start treating CPython as stable/legacy, then support and encourage efforts like GraalPython. That way the community can re-use all the effort put into the JVM.
This gives you a unusually fast Python that is also GIL-less. It doesn't seem to be used much, so there may be some compatibility problems or similar, but for a trivial test it worked just as described many years ago.
It also tells me that the GIL isn't terribly important for most things Python is used for. It certainly isn't for me.
https://stackoverflow.com/questions/1050222/what-is-the-diff...
In this market you really want to shuffle tons of data quickly, and that’s usually achieved through parallelism.
Python multiprocess library does a poor job at that.
The timing was critical, and although Python may be the Facebook of languages, one can't discount that it was extraordinarily lucky for Guido to be in the precise time and place to capitalize on good design choices.
This can be a problem for a lot of treading use cases. If I'm working on an ETL app that parses large amounts of data, either the related CPU bound tasks need to either run sequentially, call out to C extensions, or use multiple processes which incurs an overhead.
It's a pain when you know threads would suit your use case well, but the threading implementation in the language you're working in isn't up to the task.
I see this as being a nice stopgap solution for those who are too big for single-threaded but not big enough to need Spark.
I've had to rewrite my Python code in another language in 3 different projects already (multi-processing wasn't an option) and I'm not even a heavy user. Removing GIL would be very welcome.
That Python and other languages in its speed class are still used for new projects in production demonstrates that you can, often, get away with 100× slower code. But, sure, you can get away with 10× slower more often.
A lot of people learned Python and now want to use it for something more serious even if Python was not designed for it.
I work on high volume backends (like million TPS on a single node) and Python for me is just a toy language.
But I see no reason to purposefully inconvenience these people by keeping GIL if it can be removed.
Threads will allow you to split your I/O in multiple procedures, so you might start computations as soon as possible when the data is ready. They will also allow you to massively speed up aggregates without having to create a new process each time (which doesn't allow you to share memory). Threads are a BIG issue when you don't want to rely on asyncio[1].
Note [1]: Asyncio pools are single-threaded because of the GIL. This is already bad enough in practice, but they also perform very badly in CPU-bound contexts. This make them an absolute no-go when dealing with datascience code: a cpu-bounded core in an i/o-bounded wrapper.
This means that provided the code is operating on sufficiently large amounts of data (such that calls to numpy are each of sufficient duration), the multithreading in BLAS / Lapack within numpy usually give you weak scaling wrt to thread count without any tricks.
The issue however is that this require by hand making everything into structs of arrays from arrays of structs, removing as many iterations from python as possible, potentially balancing thread usage within python and within numpy etc. By this point IMO your "python" code looks more like Fortran or SQL with better string IO...
We're not interested in scaling up the parts that are already fast, but the rather mundane, uninteresting work that come before.
Assuming now you have a big dictionary, which contains tens of millions of records to be filtered upon, another say TB of data. In current Python land, you either:
1. Using multi-processing, but each process needs to create its own dictionary
2. Or create an external DB, and using the DB's client to retrieve data in certain way.
This pattern has occurred again and again in my use case, and it is always messy to solve in Python. Had Python has true multi-threading, then sharing some big but read-only object among real threads would be a possibility, and believe me, a lot of people would be really happy.
Could be that the impact of this change is far broader than just a few key libraries.
Most of the issues with multi-threading come from concurrency, not parallelism. The GIL allows concurrency, you just don’t get any of the advantages of parallelism, which is normally the reason for putting up with the complexity concurrency creates.
That thread behavior is enough to reduce the likelihood of races and collisions; particularly if the critical sections are narrow.
There's already a term for that: not thread-safe.
The definition of thread safety does not include theoretical or practical assessments regarding how frequent a problem can occurr. It only assesses whether a specific class of problems is eliminated or not.
Well, obviously.
The challenge I am putting forth on HN is to meaningfully describe _usable_ thread-unsafe software. If you've spent enough time outside university, you'll be aware that there are all kinds of theoretical race conditions that are not triggered in practical use.
Nothing obvious changed (it was still running a decade old JRE), perhaps it was a kernel security patch, perhaps a RAM was replaced or even just the runtime data increased/changed in some way which woke up this monster.
Fun fact, I actually do! It's from that perspective I wrote that: every time you perturb the software environment, a new set of bugs that didn't happen in the old env before arises.
Also, rare on one computer (or today's computer) might not be rare on another (tomorrows faster one for example).
These types of bugs are also very hard to detect. You might not know your data is corrupted. Reminds me of how bad calculations in excel has cost companies billions of dollars, except now, the calculations could be "correct" and the error sitting dormant, just waiting for the right timings to happen. Much better to not make assumptions about the safety and think about it up front: if you are using multiple threads, you need to carefully consider your thread safety.
There is no such thing as "rare enough". Random or probabilistic bugs are one of the worst things software can have.
Why not just set it to auto reboot every week, that seems to fix it - right?
I don't know how the point of the comment could be missed, but what I am saying is, it is a mistake, a rookie baby not-a-programmer not even any kind of engineer in any field, to even think in those sorts of terms at all. At least not in the platonic ideal worlds of math or code or protocol or systems design or legal documents, etc.
Physical events have probability that is unavoidable. How fast does the gas burn? "Probably this fast"
There is no excuse for any coder to even utter the word "likely".
The ONLY answers to "Is this operation atomic?" or "Is this function correct?" or "Does this cpu perform division correctly?" Is either yes or no. There is no freaking "Most of the time."
"Likely" only exists in the realm of user data and where it is explicitly created as part of an algorythm.
You cannot guarantee your public key algorithm is impossible to break, but you can use keys long enough that an attacker has an arbitrarily low chance of succes with the best known methods.
You cannot prove your program is bug free, outside of highly specialized fields like aircraft control, but you can build a multi-layered architecture that can reduce the likelihood of successful intrusion. You cannot prevent a EMP bomb from wiping all your hard-drives at once, but you will likely maintain integrity of your database for uncorrelated hardware errors.
"Likely" is a tool that works in the real world. If you will chase mathematic certainty, your competition will likely eat your lunch.
Where you might be correct is that "unlikely" is very close to "likely" in the particular topic of thread safety, you just need a sufficiently large userbase with workloads and environments sufficiently different from your test setup.
Thread1: a = 0xFFFFFFFF00000000
Thread2: a = 0x00000000FFFFFFFF
One might think that the two possible values of a if those are run concurrently are 0xFFFFFFFF00000000 and 0x00000000FFFFFFFF. But actually 0x0000000000000000 and 0xFFFFFFFFFFFFFFFF are also possible because the load itself isnt atomic.
The GIL (AFAICT) will prevent the latter two possibilities.
The GIL exists to protect the interpreters internal data, not your applications data. If you access mutable data from more than one thread, you still need to your own synchronisation.
One example would be incrementing a counter for statistics purposes. If the counter is atomic, and the reader of the value is ok with a slightly out of date value, it's fine. If code is doing this in GIL Python, it's working now, and will break after the GIL is removed.
Atomics are isolated to where they are needed, the GIL is global though C (but not Python) code can release it.
https://stackoverflow.com/a/1717514
The Python documentation seems misleading to me on this:
> In theory, this means an exact accounting requires an exact understanding of the PVM bytecode implementation. In practice, it means that operations on shared variables of built-in data types (ints, lists, dicts, etc) that “look atomic” really are.
count = 0
def inc():
count += 1
sure "looks atomic" to me, so according to the documentation should be, but isn't.On the other hand, I think you could build a horribly inefficient actual atomic counter with
count = []
def inc():
count.append(None)
def get_count():
return len(count)
I think it's quite likely there's correct and non-horrible code relying on list append and length being atomic. Although it sounds like this might continue working without the GIL:> A lot of work has gone into the list and dict implementations to make them thread-safe. And so on.
So removing the GIL might not be a problem.
Regarding list, it sounds like it might actually keep working atomically without the GIL:
> A lot of work has gone into the list and dict implementations to make them thread-safe. And so on.
if
I know you came to the same conclusion in another comments, but here's a look at it using the `dis` module:
a += 1
turns into LOAD_FAST 1 (loads b)
LOAD_CONST 1 (loads the constant 1)
INPLACE_ADD (perform the addition)
STORE_FAST 1 (store back into b)
So if the interpreter switches threads between LOAD_CONST and STORE_FAST (so either before or after the INPLACE_ADD), you could clobber the value another thread wrote to `b`.From your other comment:
> so even though an increment on a loaded int is atomic, loading it and storing it aren't
Its always the loads and stores that are the problem.
But that's the problem with relying on the GIL: it has lock in the name, but it does not protect you, unless you understand the internals and know what you're doing. It protects the interpreter. This isn't much different from programming in other languages without a GIL: if you understand the internals what you're doing you may or may not need locks, because you will know what is and isn't atomic (and even when things are atomic, its still difficult to write thread-safe code! lock-free algorithms are much harder than using mutexes).
Thread safe code requires thinking hard about your code, the GIL does not protect you from that.
For instance, in your example, I know that the call to the C code to do the list operations is atomic (a single bytecode instruction), but I can't assume that all such calls to C code are safe because of this, unless I know for sure that the C code doesn't itself release the GIL. I assume that simple calls like list append/pop wouldn't have any reason to do this, but I can't assume this for any given function/method call that delegates to C, since some calls do release the GIL.
So, with or without GIL, you either really need to understand what's going on under the hood so you can avoid using locks in your code (GIL or atomics-based lock-free programming), or you use locks (mutex, semaphore, condition variables etc). No matter what you do, to write thread safe programs, you need to understand what you're doing and how things are synchronizing and operating. The GIL doesn't remove that need.
Of course, removing the GIL removes one possible implementation option, I just don't believe the GIL really makes it any easier. Once you know enough internals to know what is and isn't safe with the GIL, you could just as easily do your own syncrhonization.
So with Python concurrency, you can get unpredictable behavior (such as two threads losing values when incrementing a counter), but not undefined behavior in the C sense, such as use-after-free.
The compilers also take care to align most variables.
So while your scenario is not impossible, it would take some effort to force "a" to be not aligned, e.g. by being a member in a structure with inefficient layout.
Normally in a multithreaded program all shared variables should be aligned, which would guarantee atomic loads and stores.
Real life bugs have come from misapplication of correct parameters for memory barriers, even on x86. Python GIL removes a whole class of potential errors.
Not that I'm against getting rid of the GIL, but I'm more sceptical that it won't trigger bugs.
Though in my opinion python just isn't a good language for large programs for other reasons. But it'd be nice to be able to multithread some 50 line scripts.
The current ARM memory model is more relaxed regarding the ordering of loads and stores, but not regarding the atomicity of single loads and stores.
Some more discussion here: https://stackoverflow.com/questions/1717393/is-the-operator-...
Presumably this patch changes the list implementation in some way so that the extend operation remains thread-safe without the GIL.
with gil:
call_a_method()
print(some_debugging_info)
with all the sit-ups you’d have to do in a “real” concurrent language.It only protects the state of the Python interpreter and that of C/Cython extension modules. Though even there, you can have unexpected thread switches, e.g. in Cython `self.obj = None` can result in a thread switch if the value previously stored in `self.obj` had a `__del__` method implemented in Python.
And AFAIK pretty much any Python object allocation can trigger the cycle collector which can trigger `__del__` on (completely unrelated) objects in reference cycles, so it's pretty much impossible to rely on the GIL to keep any non-trivial code block atomic.
Porting the thing to a multi-core system revealed that there were a lot of nasty concurrency bugs, like dead-locks or crashes which happened after a day of operation. And this wasn't a toy system - it was in use for a long time in an industrial application, and the customer was not too happy about the intermittent dead-locks. I commiserated with the poor engineer who had the quite stressful task to debug this, equipped with a lot of dedication but an insufficient background.
Frankly, while it would be nice to be able to write parallel code in pure Python, I think that Clojure with its purely-functional approach has the better concepts for this. And moreover, actually improving performance by parallel computation (using several CPUs to work in parallel on the same thing) is damn hard and unsolved in many cases (just come up with an efficient parallel Fast Fourier Transform and you might get a Turing award). What is mostly needed (outside of massive data processing pipelines) is concurrency for event-driven systems. Python can handle that, Clojure does handle it in a much more elegant way.
Maybe time to rethink? https://www.techrepublic.com/article/programming-languages-w...
If this is as promising as it sounds, it seems Python 4 now has its "thing" and is on the horizon. Or at least may become a serious thing to talk about
What struck me as most significant was the opportunistic breakage of things not related to the unicode transition. In the many years it took to win people over to v3, they could have marched over all the breaking changes a year at a time. Given that side-by-side installs of python3.x point versions are very functional, with or without venvs, this would have been much more palatable. Perhaps harder than it sounds though.
I attempted a couple of 2to3 translations of open source libraries over the years, with varying degrees of success. Every time I found that most of the changes were easy, but debugging the broken bits was hard due to the sheer volume of source changes. If instead I could have done conversions where there was only a single major semantic change at a time, it would be so much easier to figure out what was going wrong at any given step. Furthermore, I imagine that a single-breaking-change mentality would lead to better documentation on how to transition for each version.
For this reason, I have become rather suspicious of yearly release schedules. Swift is even more frustrating: the version changes are really just dictated by Apple's yearly PR calendar. Some big things get rushed out for WWDC before they are ready, and smaller fixes can get held back until the next year. I would much rather that the language teams just prioritize one thing at a time, release it when it is ready, and foster a community where staying up-to-date on the latest version is easy and desirable (a more complicated story for Apple than for Python I think, due to ABI, OS version, etc).
From past discussions on HN I've gathered that there is such a thing as release fatigue, where developers get irritated when libraries release breaking changes too often. Nevertheless I often wonder if languages and libraries could improve faster by making more breaking changes, one at a time, with robust side-by-side installs to facilitate testing across versions. I wish side-by-side library versions were possible in Python, just to facilitate regression testing.
Bringing this all back to the post, I sincerely hope that if Python 4 is a breaking change to the GIL, that it will be only that.
I'm curious what others think about all this. Thoughts?
That was the point of the "from __future__" imports. You could get most of the way toward Python 3 so that 2to3 would be easier to work with and the new semantics could be gradually baked into the code prior to migration.
Python 3 had 25 years of cruft to clean up. They won't have to do that again.
This.
It makes absolutely no sense to claim that having to deal with a single non-backwards compatible release is somehow worse than having to deal with a sequence of non-backwards compatible releases.
Even though the migration from Python2 to Python3 faced some resistence, if anything the decision was totally vindicated.
Corporate developers, who have taken over Python and other people's work, like unnecessary changes, because they get many billable hours of seemingly complex work that can be done on autopilot.
Corporations might even take over more C extensions whose developers are no longer willing to put up with the churn and who have moved to C++ or Java.
In the long run, this is bad for Python. But many developers want to milk the snake until their retirement and don't care what happens afterwards.
In my 20+ year career I have never worked with a programmer that matches this description.
Maybe I've got lucky.
I want the python that Guido promised me, with 2021 performance. I don't want some abhorrent committee-designed piece of middle-of-the-road shitware glue language that I must use because everyone uses it.
I want a language that doesn't spin it's single-threaded wheel in a sea of CPU cores, and I want a language that has one obvious way of doing things without needing to grok and parse dumb """clever""" hacks that will only be abused by midlevel programmers to show off hoe they saved typing a few lines of additional code.
To me, speed + simplicity = ergonomy = joy. I want a new python 4 to focus exclusively and intensely on performance improvements and ergonomy.
Gone are the days when you invest in a platform like python, and they make crazy decisions that kill the platform's future (e.g. perl5). Ignore small syntax stuff like := and focus on the big stuff.
That says nothing about their quality. It just says you like them. If you gave me unhealthy food I'd probably eat it immediately too. Doesn't mean I think it's good for me.
> Ignore small syntax stuff like := and focus on the big stuff.
They're not "small" when you immediately start using them in a "large fraction of your code". And a simple syntax that's easy to understand is practically Python's raison d'être. They added constructs with some pretty darn unexpected meanings into what was supposed to be an accessible language, and you want people to ignore them? I would ignore them in a language like C++ (heck, I would ignore syntax complications in C++ to a large degree), but ignoring features that make Python harder to read? To me that's like putting performance-killing features in C++ and asking people to ignore them. It's not that I can't ignore them—it's that that's not the point.
my_match = regex.match(foo)
if my_match:
return my_match.groups()
# continues with the now useless my_match in scope
Versus if my_match := regex.match(foo):
return my_match.groups()
# continues without useless my_match in scope
How is the second one less readable? Have you ever heard of a real world example of a beginner or literally anyone ever actually expressing confusion over this?The more major problem with the walrus operator is more complicated expressions they made legal with it. Like, could you explain to me why making these legal was a good thing?
def foo()
return ...
def bar():
yield ...
while foo() or (w := bar()) < 10:
# w is in-scope here, but possibly nonexistent!
# Even in C++ it would at least *exist*!
print(w)
# The variable is still in-scope here, and still *nonexistent*
# Ditto as above, but even worse outside the loop
print(w := w + 1)
If they just wanted your use case, they could've made only expressions of the form 'if var := val' legal, and maybe the same with 'while', not full-blown assignments in arbitrary expressions, which they had (very wisely) prohibited for decades for the sake of readability. And they would've scoped the variable to the 'if', not made it accessible after the conditional. But nope, they went ahead and just did what '=' does in any language, and to add insult to injury, they didn't even keep the existing syntax when it has exactly the same meaning. And it's not like they even added += and -= and all those along with it (or +:= and -:= because apparently that's their taste) to make it more useful in that direction, if they really felt in-expression assignments were useful, so it's not like you get those benefits either.I am baffled to learn that it's kept in scope outside of the statement it's assigned, and I agree it would have a negative impact on readability if used outside of the if statement.
No, I'm pretty sure that's intentional. You want the left-hand side of an assignment to be crystal clear, which "foo() or w := bar()" is not. It looks like it's assigning to (foo() or w).
def thing(): return True
if thing() or w:= "ok": # SyntaxError: cannot use assignment expressions with operator
pass
print(w)
. . .
if thing() or (w := "ok"):
pass
print(w) # NameError: name 'w' is not defined
The first error makes me think your concern (that w is conditionally undefined) was anticipated and supposed to be guarded against with the SyntaxError. I believe the fact you can bypass it with parentheses is a bug and not an intentional design decision.> The motivation for this special case is twofold. First, it allows us to conveniently capture a "witness" for an any() expression, or a counterexample for all(), for example:
if any((comment := line).startswith('#') for line in lines):
print("First comment:", comment)
else:
print("There are no comments")
I have a hard time believing even the authors (let alone you) could tell me with a straight face that that's easy to read. If they really believe that, I... have questions about their experiences.The beauty of Python...
From the PEP:
> An assignment expression does not introduce a new scope. In most cases the scope in which the target will be bound is self-explanatory: it is the current scope. If this scope contains a nonlocal or global declaration for the target, the assignment expression honors that. A lambda (being an explicit, if anonymous, function definition) counts as a scope for this purpose.
I find this particularly strange and inconsistent:
lines = ["1"]
[(comment := line).startswith('#') for line in lines]
print(comment) # 1
[x for x in range(3)]
print(x) # NameError: name 'x' is not definedI'm pretty sure it's what I explained here: https://news.ycombinator.com/item?id=28899404
False or w := 1:
Is grouped like so: (False or w) := 1
Which is a SyntaxError. That's... not a smart place for it to be in the operator hierarchy. I expected it to be near the very top, like await.1. https://docs.python.org/3/reference/expressions.html#operato...
Edit: 20 minutes later, can't respond.
There are two aspects I have been thinking about while looking at this: Introduction of non-obvious behavior (foot-guns) and readability. Readability is important, but I have been thinking primarily about the foot-gun bits, and you have been emphasizing the readability bits. I can't really accurately assess readability of something until I encounter it in the wild.
x := 1 if cond else 2
never resulting in x := 2 which is pretty unintuitive.And you have to realize, even if the precedence works out, nobody is going to remember the full ordering for every language they use. People mostly remember a partial order that they're comfortable with, and the rest they either avoid or look up as needed. Like in C++, I couldn't tell you exactly how (a << b = x ? c : d) groups (though I could make an educated guess), and I don't have any interest in remembering it either.
Ultimately, this isn't about the actual precedence. Even if the precedence was magically "right", it's about readability. It's just not readable to assign to a compound expression, even if the language has perfect precedence.
Here's another way to trigger the same NameError, via "global":
import random
def foo():
return random.randrange(2)
def bar():
global w
w = return random.randrange(20)
return w
while foo() or (bar() < 10):
print(w)
For even more Python-is-not-C++-fun: import re
def parse_str(s):
def m(pattern): # I <3 Perl!
nonlocal _
_ = re.match(pattern, s)
return _ is not None
if m("Name: (.*)$"):
return ("name", _[1])
if m("State: (..) City: (.*)$"):
return ("city", (_[2], _[1]))
if m(r"ZIP: (\d{5})(-(\d{4}))?$"):
return ("zip", _[1] + (_[2] if _[2] else ""))
return ("Unknown", s)
del _ # Remove this line and the function isn't valid Python(!)
for line in (
"Name: Ernest Hemingway",
"State: FL City: Key West",
"ZIP: 33040",
):
print(parse_str(line)) # w is in-scope here, but possibly nonexistent!
# Even in C++ it would at least *exist*!
because I don't see how bringing up C++'s semantics is relevant when Python has long raised an UnboundLocalError for similar circumstances.If I understand you correctly, you believe Python should have introduced scoping so the "w" would be valid only in the if, elif, and else clauses, and not after the 'if' ends.
This would be similar to how the error object works in the 'except' clause:
>>> try:
... 1/0
... except Exception as err:
... err = "Hello"
...
>>> err
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
NameError: name 'err' is not defined
If so, I do not have the experience or insight to say anything meaningful.Python’s `if` statements do not introduce a new scope.
I will readily admit that the walrus operator doesn't do what I thought it did and I have no interest in whatever utility it provides as it exists today.
Definitely. You would think if they're going to undermine decades of their own philosophy, they would instead introduce variable declarations and actually help mitigate some bugs in the process.
if (percentage := ((student["marks"]/700)*100)) > 70:
As a non-Python programmer it is usually pretty easy for me to correctly guess what a piece of Python code does. (And once in a while I need to take a look at some Python code).
Walrus operator got me. I tried to guess what it did, but even having simple code examples I could not. My guesses were along the lines of binding versus plain assignment, or some such. None of my guesses were even close. I had to google it to find out (of course I could also read the documentation).
class C(object):
A = 1
B = 2
x = 3
y = 10
print(x - B)
match y:
case C.A:
print('A')
case B:
print(y - B)And no, this is not really a scoping issue. Match is literally writing to a variable in one pattern but not the other. A conditional write is just a plain inconsistency.
The sad part is both of these features are stumbling over the fact that Python doesn't have variable declarations/initialization. If they'd only introduced a different syntax for initializations, both of these could have been much clearer.
I guess I'm not sure where "design" ends and "implementation" begins? To me, how to handle matching on variables that already exists is both, because "pattern matching and destructuring" are the features and how that must work in the context of the actual language is "implementation". It being written in a design doc and having real world consequences in the resulting code doesn't make it not part of the implementation.
Instead of quibbling over terms, I was much more interested in whether you like the idea of pattern matching.
I think not liking the final form a feature takes in the language is fundamentally different from wholesale disliking the direction the language design is going.
But I mean, you can call it that if you prefer. It's just as terrible and inexcusable regardless of its name. And yes, as I mentioned, I would have loved to have a good pattern matching system, but so far the "direction" they're going is actively damaging the language by introducing more pitfalls instead of fixing the existing ones (scopes, declarations, etc.). Just because pattern matching in the abstract could be a feature, that doesn't mean they're going in a good direction by implementing it in a broken way.
I guess like they say, the road to hell is paved with good intentions.
By this definition, bugs and other unintended consequences that the user encounters are "Design".
> Calling that an implementation detail is like calling your car's steering wheel an implementation detail.
Yes, if there weren't so many important decisions behind the outcome of a car being steered with a steering wheel, it could be a steering handle, or a steering joystick, or just about anything else that allows you to orient the front wheels of the car. The same is true of the pedals on the floor. Those could be implemented as controls on the steering wheel instead. Whether it's an implementation detail depends on the specificity of the feature in question. When I asked you about an "implementation detail", it was scoped to "the feature is pattern matching" (can I steer the car?) and you scoped it to "the feature is pattern matching without overwriting variables conditionally in surprising ways" (can I steer the car with failure modes that aren't fatal?).
and then all that?
match status:
case 404:
return "Not found"
not_found = 404
match status:
case not_found:
return "Not found"
The first checks for equality (`status == 404`) and the second performs an assignment (`not_found = status`).`not_found` behaving differently from the literal `404` breaks an important principle: “if you see an undocumented constant, you can always name it without changing the code’s meaning” [0].
[0] https://twitter.com/brandon_rhodes/status/136022610839909990...
Worst of all, though, it’s really just another way to write if-elif-else chains that call `isinstance` - a pattern which Python traditionally discouraged in favour of duck-typing.
1. https://pythonsimplified.com/the-most-controversial-python-w...
But for the sake of maximum pendatry let me paste some nitpicky little detail from a somewhat recent syntactic addition:
>>> def f(a, b, /, **kwargs):
... print(a, b, kwargs)
...
>>> f(10, 20, a=1, b=2, c=3)
10 20 {'a': 1, 'b': 2, 'c': 3}
a and b are used in two ways.
Since the parameters to the left of / are not exposed as possible keywords, the parameters names remain available for use in **kwargs
Jesus fucking hell on a tricylce so now i have *'s and /'s showing up in function signatures so someone can prematurely optimize the re-use of variable names without breaking backwards comparability?!Python is becoming a mockery, dying a death through a thousand little cuts to its ergonomics.
Avoiding tedious boilerplate by adding nice features like the walrus operator is precisely what lets us avoid "death through a thousand little cuts to its ergonomics", in my view.
Sure, maybe writing
m = re.match("^foo", s)
if m != None:
...
isn't so bad, but in that case maybe writing i = 0
while i < len(stuff):
element = stuff[i]
...
i += 1
wouldn't be so bad, and we could get rid of Python's iterator protocol?The iterator protocol is way more general than what you have; it's not remotely comparable.
if foo := data.get("foo"):
handle_foo(foo) foo = data.get("foo")
if foo:
handle_foo(foo)
The only place I have found the walrus operator useful is in similar `while` loops, where the equivalent code would be: while True:
foo = data.get("foo")
if not foo: break
handle_foo(foo) for foo in iter(partial(data.get, "foo"), None):
handle_foo(foo)
So I would only use the walrus operator for the first example (the if statement), which even though it is exactly the same as doing it in two steps just feels nicer as a single step.Python has always been strongly typed (Python has strong dynamic typing). Adding typechecks moves it towards being gradually/statically-typed.
On that list: the walrus operator and the new switch thing. If I understand them fully and correctly, those two things don't enable developers to do things that were impossible before, instead they add new ways to do things that were possible prior.
That's the Python I know and love.
Of course, this doesn't mean I'll love Python any less, just that I wished there were more focus on staff that matters like the topic of this article. Or maybe getting type hinting better.
Again, this is just my opinion.
After using pattern matching in Rust and switch statements in JavaScript, I personally am very excited for that addition to Python, but I understand the feature is divisive and will concede it as a matter of opinion.
Edit: turns out the walrus operator does not cause the variable to move out of scope after the if block, which is disappointing. IMO the worse anti-pattern has already been part of the language, which is not creating new scopes for if statements.
Python always had at least half a dozen ways of doing trivial things such as packaging dependencies or even doing http requests or looping over stuff.
But yes, of course can one also argue that they add friction on the global scale, because it's yet another syntax-element to know about, and the benefit is rather small on surface. But that's the problem with syntax, it's always a trade-off between overhead and benefit.
.. but. Let's not count our chickens until they are home. I'm wondering if the Python dev community will take on this challenge. I hope so, Sam seems to really have put in a lot of effort!
Python is there now, if you ask me. It should slow down and focus more on "maintenance" stuff with little to no impact on its interface. And maybe work on big projects like multithreading or stronger typing on the background and them when they're fully ready.
Please, don't. There are wonderful strongly types languages out there, so if one wants or needs a strongly types language, use that, and not Python.
Perhaps, without the GIL and with typing information included, additional performance gains with be on offer.
But the "have it your way" nature of Python is a bigger win than either end of the data typing spectrum.
TMTOWTDI - There's more than one way to do it - is Perl's motto :)
Zen of Python suggests there's just one correct way.
We won't get data race free guarantees but if built into pandas or Vaex we can have a near transparent API.
It'll really open things up for those apps running on 32 core machines. They're out there, I deploy these things frequently (Plotly Dash framework for large Enterprise customers).
IMO, removing the GIL is a major mistake. The GIL is what allows you easy concurrency and to keep the language's 'magic' while ensuring correctness. If you need parallelism, there's processes and probably other tactics (I'm not super up to date on Python things). If you simply remove the GIL you have a bunch of race conditions, so you need a bunch of new language constructs, and it just adds a bunch of complexity to solve problems that don't really need solving.
IMO they should just do what Ruby did with Ractors; basically a cheap alternative to spawning more processes. Rewriting absolutely everything that uses threads to be thread-safe is a waste of time.
if x in d:
del d[x]
else:
d[x] = True
Is a classic example -- if two threads execute that, you can't predict the outcome (but a KeyError is quite likely)The GIL only protects the CPython virtual machine; it doesn't protect user code. Concurrent code with shared mutable state already needs explicit mutexes.
I already have to be careful to only write to a shared object from one thread, since I have no guarantees on order of execution.
The main benefit of the GIL, from my recent reading is that it makes ref counting fast and thread safe. The meat of the proposal is changing ref counting so that it's almost as fast and atomic without the GIL.
In Java you would have to worry about safe publication to make the change visible to the other thread, but thanks to the GIL changes in Python are always (I think?) made visible to other threads.
current_boolean_value = global_cancelled
global_cancelled = new_boolean_value(current_boolean_value)
The race conditions a Gil program and a no Gil program should have should be same. A Gil is not the only way to keep certain operations safe.
Only if you are writing multithreaded coded. Otherwise it is just a loss.
The nice thing about everything being single threaded is that nothing will break if you remove the GIL. It will still be single threaded. It's only when you actively start using multiple threads that things might break. So, you won't have race conditions until you do that and then only if you do things that you shouldn't be doing like sharing things across threads that you should not be sharing because they aren't thread safe.
Removing the GIL will simply enable people to start gradually fixing things and give them the option to use threads instead of forcing them to use completely different languages.
Ractor style asynchronous programming might be a good idea for python as well. One does not exclude the other.
Interesting. When I write Python, and I write Python most of the time, I'm not chasing performance. But when it comes to speed, a 18x speedup gets Python up to par with currently much faster languages like Java. At least if you are willing to spam threads like your life depends on it.
conda install numba & conda install cudatoolkit comes to mind?
And you should data orientated, so all that remains from objects is arrays of structures with the index being the object indicator. Im honestly puzzled about the use-case..
But what would be nice currently, is to create map reduce API w/o shared data or just readonly data.
Something like ProcessPoolExecutor but instead of spawning new process and pickling input/output data, create threads w/o GIL w/o pickling input/output data.
<grumble> Please escape Discord links so people don't accidentally click a live one and thereby expose data to Hammer & Chisel, Inc. </grumble>
It's a constant slap-in-the-face how prevalent Discord is in the tech community. Need support with an obscure library? Discord. Want to talk to contributors for a project? Discord.
I honestly don't know why people so greatly prefer Discord to IRC. The interface? Network effects? Whatever it is, Discord and H&C are disgusting, and I pray the tech community finds a way to escape it.
can you fill me in? a couple google searches didn't turn up much.
Also, it's so disheartening to see people tout Discord's interface as its 'killer feature.' As far as I'm concerned, if I can't mold it into a decent TUI for my terminals, it's okay at best. If an app actively resists its interface being molded, however, that's just evil.
Why I Hate Discord (A Manifesto) [Without Sources] Discord is aggressively proprietary.
People ought to own their data, but Discord's architecture ensures everything gets hoovered into the mother-ship. At first, this was for nothing more than to facilitate their client-server communications model; recently, all of the hoovered data gets submitted to their AI moderation platform. (searched for sources on this but it was very hard to find anything. I remember talk about this ~c. Nov 2020, might be wrong)
I should be able to modify an interface to my tastes. Modifying Discord is explicitly against their ToS, including their interface; attempting to do so will lead to a ban. Don't like their painted whore? Prefer to chat from your bespoke terminal? GTFO, Hammer & Chisel knows what's best for We Peons.
Addenda: I find Hammer & Chisel developers and Discord admins to be disgusting. This is hearsay & personal opinion from my time hanging around the developer chats, but they all came across as nasty people. I, personally, believe many of the news reports surrounding the grooming controversies; searching "Discord admin controversy," "Discord allthefoxes controversy," "Discord cub policy," &c. turn up some relevant articles. Most of these articles are from low-quality reporting shops, but I buy in to the narrative.
Except IRC doesn't even save your community's history in the first place.
I am praying for CPython to become faster. But I need faster singlethreaded performance, so web applications benefit from it.
Typical ways around that include mmap’ng things or various types of shared memory IPC, but that is a lot of work.
While performance may not matter, productivity very much suffers from multi-process programming.
I’m saying that the penalties (including development effort) for work co-ordinated between processes compared to threads can vary from nearly zero (sometimes even net better) to terrible depending on the nature of the workload. Threads are very convenient programmatically, but also have some serious drawbacks in certain scenarios.
For instance fork()’d sub-processes do a good job of avoiding segfaulting /crashing/dead locking everything all at once if there is a memory issue or other problem when doing the work, it’s very difficult in native threading models to do per-work resource management or quotas (like maximum memory footprint, or max number of open sockets, or max number of open files) since everything is grouped together at the OS level and it’s a lot of work to reinvent that yourself (which you’d need to do with threads). Also, the same shared memory convenience can cause some pretty crazy memory corruption in otherwise isolated areas of your system if you’re not careful, which is not possible with separate processes.
I do wish the python concurrency model was better. Even with it’s warts, it is still possible to do a LOT with it in a performant way. Some workloads are definitely not worth the trouble now however.
When I execute this:
python3 -c 'import time; time.sleep(60)'
And then pmap the process id of that process, I get 26144K. That is 26MB.As for timing, when I execute this:
time python3 -c ''
I get 0.02sAlso, since the copy on write generally is memory page by memory page, even if you were doing a lot of that, if most of those ref counts are in a small number of pages, it’s not likely to really change much.
It would be good to get real numbers here of course. I couldn’t find anyone obviously complaining about obvious issues with it in Python after a cursory search though.
That said, many such databases are (or already have) rolled out connection pool proxies for reasons like this, so meh.
Like anything, it’s possible to split this work out to a separate process but the IPC overhead is a lot.
That was a really sad moment, and I've never felt good about python since.
I wanted "as you wait for the GPU to churn through this batch, start reading the next batch from disk & preprocessing it on the CPU"
Getting this to work turned out so ass backwards it made me sad
Also I pity the fool who tries to connect a debugger to code using multiprocessing.Pool()...
PyCharm works fine
I solved it by using python threads and C function (cv2.resize)
Friend: I saw this interesting tech article ...
Me: Site?
Friend: mumble corp IT site mumble
Me: Bye
.. but, as I said, removing the GIL will almost certainly be opt-in.
I assume that a change to relax the GIL will both allow you to opt-out of it, and allow you to use locking versions of primitive data-structures, anyway; it's not like it's going to just vanish overnight with no guardrails.
Rust has a great rule: Sharing XOR Mutation.
Python is higher level, so message passing and passing "owned" values between threads is all the more feasible and sensible.
`dict.update()` will call methods like `__eq__`, and those methods (if implemented in Python) may temporarily release the GIL.
If you’re relying on whatever python version is distributed with whatever machine it happens to be on, there are a huge number of problems you’re already going to have.
Even if those changes are better for other ways of solving problems.
I'm probably biased because I think Python is a hacked together mess but I just don't see what the point is in dragging it around.
I think Python is shit but so is most everything else, I'm interested in whether people jump ship or just work around it's issues at scale. I work on a programming language designed to avoid messy python scripts internally, so I am sincerely interested in these decisions.
I think Go benefited a lot from that.