Instrumenting Python GIL with eBPF
coroot.com
coroot.com
But the author literally starts by explaining that that's a typical argument for webservers because they're mostly I/O bound. Anyone working with code that's more CPU bound will have very different numbers, and interpret them differently.
All of this is why the GIL wasn't removed 20 years ago. There are real trade-offs here.
This is the reason why previous attempts were rejected. But those attempts came from single individuals and not from a photo sharing website.
This matters if --disable-gil becomes the default in the future and is forced on everyone.
https://dabeaz.blogspot.com/2011/08/inside-look-at-gil-remov...
However, PEP 703 specifically points out that performance-critical container operations (__getitem__/iteration) avoid locking, so I'm still highly skeptical that those locks are the cause of the 30-50%.
https://peps.python.org/pep-0703/#optimistically-avoiding-lo...
But I think you're right be sceptical that somehow this is to blame for the Python perf leak.
If a Linux process wants a million locks that's fine, that's just 4MB of RAM now.
I'm not a web person, but can't you also gain I/O from parallelization? The I/O bound is waiting on responses right? So parallelization should increase I/O because you can make multiple asynchronous requests and even if there is dependence you can often stage or do partial computation in the mean time (at least this is common in scientific computing). (And if disk, well reading/writing to disk in parallel is far faster than serial but idk why you'd read/write to disk with pure python. Though it seems people do). So wouldn't this have significant effects that are more than the 36ms that we see in TFA? Or am I missing something and can someone explain why my guess is wrong?
It's "small", sure, but in production performance issues are often "death by a thousand cuts" situations, so a 3.6% reduction is a big win compared to optimizations that are often in the 1% range.
> only 3.6%
> GIL is then a pretty overblown problem, as this is not very significant
How do you conclude this? 3.6% of the time spent being locked seems pretty significant. As another user notes, this is 3-4 fewer machines you need to pay for out of every 100 you currently use.But the problem is worse than that. That's 3.6% of the time you're locked, not 3.6% of the time your program takes vs the time it would take without GIL. During that same time other processes could be run. I've seen plenty of mundane and ordinary tasks be 10x to 1000x faster when comparing serial to even unoptimized parallel. I'd be really careful to take this conclusion unless you actually have experience in writing parallel and/or optimized code.
But I think an interesting one is the also literally the most common example that's in many stats textbooks, at least the ones that cover Bayesian stats. That being that if a test is 95% accurate, that it does not mean that that scoring positive on a test means that there's a 95% chance you have whatever the test is testing for.
I think there's a great irony in the latter, as I find in the tech and engineering communities a lot of confidence, especially in understanding numbers yet very often make this mistake[0], and frequently want to establish meritocracies (advocating for standardized testing and leet code). But understanding the former should result in thinking the latter is ineffective.[1]
--------------------------------
[0] I particularly like 3Blue1Brown's introduction to Bayes theorem as it includes the Kahneman and Tversky questions AND (most importantly) he discusses how when the questions are reframed that peoples accuracy can dramatically swing from overwhelmingly wrong to overwhelmingly correct. I like this because the hardest thing in doing statistics (and much of science) is actually down to this. It is all about understanding the assumptions you are making, and more specifically, the hidden ones. https://www.youtube.com/watch?v=HZGCoVF3YvM
[1] I think this one is worth working out a little and so let's apply Bayes to a proxy problem (for illustration). It's good to start to understand how metrics can undermine meritocracies and it will illustrate how people actually lie with statistics despite never saying anything that is actually untrue (and that is why it is so common and why often lying with stats/data is unintentional).
Let's let monetary wealth represent our measure of success, let's assume that college entrance exams (and the ability to graduate) is purely meritocratic (i.e. "I went to Standford" -> "I'm smart"), let's define $100m net worth as ultra wealthy, let's assume that graduation from a college gives a person just intelligence and the connections you make there are not significantly influential to success (this one is actually key, but "left to reader as an exercise"), and let's assume your final net worth is purely due to your own work/efforts. We're using this because the frequency of which awarding institution is used to imply one's value/worth/skill (your cousin/friend/boss/coworker who brags about being an "x" graduate or parents showing off their kid's intelligence, etc) and how admittance to these institutions has a very strong correlation to standardized test scores.
Top 8 schools here are Harvard, MIT, Stanford, UPenn, Columbia, Yale, Cornell, Princeton which represents 7%, 5%, 5%, 4%, 4%, 4%, 3%, and 3% of US ultra wealthy, respectively. And these represent 35% of all of the US's ultra wealthy, which there are 10,660 today (~28.5k globally). So 746, 533, 426, and 320 people, respectively. (Note we now have to assume there isn't double/triple counting: e.g. undergraduate Harvard, masters MIT, PhD Stanford) Each year these schools graduate 10k (Harvard is a bit confusing), 3.5k, 5.2k, and so on. Now I don't have data for when those people graduated from those schools, so we'll describe a function to help us estimate. Using Harvard we have 746 people / (10k graduates * number of years). I was able to find that the median age of this group is >55 but also evidence of a bimodal distribution). So let's naively assume a 10 year period, as this still gets us a pretty conservative/lower bound estimate (as well as assuming no dupes). So now we have that Harvard represents 7% of US centimillionaires, but only (and this is an absurd overestimate) 0.7% of Harvard graduates are centimillionaires.
Now we see a framing issue. Harvard tries to show its value by advertising that it is the highest producer of centimillioaires, citing this 7% number. But when we consider context (i.e. the number of graduating people) we find that far fewer than 1% of graduates make it to that level of success, which suggests that a Harvard education has little to nothing to do with said success. Even if we try a clearer cut case like with MIT where there are 2k engineering degrees awarded a year (200 arch, 100 hum/art/soc. sci, 1k management, 500 sci, etc) then we have 533/(2k * 10) or 2.6% and this still does not imply that MIT has a large effect on you becoming ultra wealthy, because this all relates to the exact same problem as the intro medical test problem: the base population size is large and the event you're testing for is rare. In fact, you can do this type of analysis for any test and mark of success and you'll probably end up with very similar results. The underlying reasoning being that there are many factors that actually lead to success and any single one of them is not a significant indicator, but it is the agglomeration of weakly influencing events that accumulate over time. Of course these numbers still have utility and this doesn't suggest you shouldn't get a degree or that you shouldn't aim for a top school degree, but they have to be understood under context and that context is often complex. Which removing the context is how one lies with statistics/data and the most common person that lie is told to is yourself.
Sadly, that is not the world we live in.
I've cleaned up dozens of applications written by people with flawed understandings of threads, multiprocessing, and asyncio. I don't even blame the developers for this; it's a glaring language design problem.
If you need parallelism, Python is not the language you should reach for. Nobody ever takes my word for it until it's release day and the product is a broken pile of spaghetti code.
Maybe someday this will make parallelism in Python at least as sane as other languages, but in the meantime, I still want to use a compiled language any time performance matters and wait for the kinks in no-GIL to be ironed out.
[0] https://docs.python.org/3.13/whatsnew/3.13.html#free-threade...