A similar misunderstanding exists about SQLite and concurrency.. but that's a topic for another time.
I am perfectly aware that threads are not free.
Just as I am perfectly aware that a context switch between threads is less expensive that switching a new process onto the core, and that IPC requires kernel involvement.
But the overhead of a context switch for a thread and process is very similar. The main difference is whether memory is shared by default.
Except that the threads share the exact same virtual address space, and processes do not, which makes the thread context switch faster.
And that is to say nothing about the setup and teardown process, which for a process involves copy-on-demand'ing the entire memory, but for a thread merely setting up its own stack.
That's what I said. But it's really not much. I'm afraid we will need numbers now to continue the conversation. If I measured would you be open to changing your opinion? Or are you committed to this topic, so that it would have no bearing?
Spin up and tear down a million pthreads in C, and see how long that takes and how much memory it takes. Then spin up and tear down a million processes in C and see your computer grind to a halt until you kill the process that is starting the processes, if you can even get your computer to do that without power-cycling.
It's <50 lines of code for each, so I'm eagerly waiting for your response!
Notably, my confidence here comes from the fact that I don't generally get into performance arguments without having actually tested what I'm saying. I've written this code before--it's what I do whenever I'm checking out a new programming language or threading library. Given the complexity of modern computers, nobody really can predict how a program will behave without testing it (except maybe in assembly) there's just too many variables. So you should stop doing that.
If you decide to try the same thing in Java (the other language mentioned), probably drop the number of threads/processes down to 100,000, since Java's lightweight threads aren't quite as efficient. 100,000 processes will probably still be enough to crash your computer.
I'm sure you can find some language/library which implements threads particularly inefficiently, so let's stick to pthreads/C and avoid that straw man.
EDIT: Here ya go, I had ChatGPT write this one for ya:
#include <stdio.h>
#include <pthread.h>
#include <unistd.h>
void* threadFunction(void* arg) {
// Sleep for 10 seconds
sleep(10);
pthread_exit(NULL);
}
int main() {
int numThreads = 1000000;
pthread_t threads[numThreads];
// Create threads
for (int i = 0; i < numThreads; i++) {
int result = pthread_create(&threads[i], NULL, threadFunction, NULL);
if (result != 0) {
printf("Failed to create thread %d\n", i);
return 1;
}
}
// Join threads
for (int i = 0; i < numThreads; i++) {
int result = pthread_join(threads[i], NULL);
if (result != 0) {
printf("Failed to join thread %d\n", i);
return 1;
}
}
return 0;
}
And... #include <stdio.h>
#include <sys/types.h>
#include <sys/wait.h>
#include <unistd.h>
int main() {
int numProcesses = 1000000;
pid_t childPID;
// Create processes
for (int i = 0; i < numProcesses; i++) {
childPID = fork();
if (childPID < 0) {
printf("Failed to create process %d\n", i);
return 1;
} else if (childPID == 0) {
// Child process
sleep(10);
return 0;
}
}
// Wait for all child processes to finish
int status;
pid_t pid;
while ((pid = wait(&status)) > 0);
return 0;
}
It looks like the latter just crashes the program without taking down my whole machine now, which is an improvement over the last time I tried this with processes. $ gcc threads.c
$ time ./a.out
real 0m10.097s
user 0m0.035s
sys 0m0.239s
$ gcc process.c
$ time ./a.out
real 0m10.168s
user 0m0.579s
sys 0m0.347s
Were you running on something besides linux, or not natively? Is this something that degrades with the large numbers?Also, spawning is not context switching. That's the overhead that matters. But according to your own test, spawning in reasonable numbers will be about the same.
Running on MacOS, but I've run this in Linux.
> Is this something that degrades with the large numbers?
The concern here is memory--once you push into pagefile your processes will become extremely slow.
> Also, spawning is not context switching.
Thank you obviousman.
> That's the overhead that matters.
Why do you think you know every use case? You don't. There are tons of use cases where having to be concerned about creating and destroying threads places a large burden on the developer.
> But according to your own test, spawning in reasonable numbers will be about the same.
You didn't run my test.
Running 3000 processes is a few orders of magnitude less than running 1000000, and you don't get to determine what "reasonable" is for every application that exists.
> Running 3000 processes is a few orders of magnitude less than running 1000000
I didn't decide on the limit, my OS did. So they must think it's unreasonable.
> Running on MacOS, but I've run this in Linux.
Well your test works fine on linux. MacOS is not designed to run large multi-process server loads. Linux has specifically optimized forking and context switching for processes.
> Why do you think you know every use case?
Why do you think you can't handle most cases by using processes? The goal isn't for a tool to handle every use case, it's to handle a specified set of use cases well.
Python itself doesn't work for every use case.
Looks like we are safe to ignore overhead of launching processes for programs with fewer than 3000 threads.
> Thank you obviousman.
I wanted to compare context switch, you decided to measure something else. I am pleasantly surprised anyway.
> Well your test works fine on linux. MacOS is not designed to run large multi-process server loads. Linux has specifically optimized forking and context switching for processes.
Oh JFC, stop wildly speculating and pretending it's the truth. The test was a million threads/processes, and you ran 3000. By your own description you didn't run the test, period. I just ran it on Debian on a VPS and, you know, it did exactly what I said it was going to do, because gosh, this isn't the first time I've run this test on Linux.
And you seriously want to attribute this to Linux having optimized forking and context switching as if you have any idea what that means? Please do tell, which optimizations did they apply that somehow they've hidden from BSD and Apple?
Given your propensity to make things up when you don't know something, I'm beginning to think whatever problem you ran into with a million threads was fixable and you just didn't know how to, so you made up a new straw man test to try and win an argument. Have you measured the memory usage at 3000 yet, or are you still ignoring any part of reality that isn't convenient for your argument?
> Why do you think you can't handle most cases by using processes?
Where did I say that? Unlike you, I don't make generalizations about "most use cases" because I don't pretend to know what everyone in the world is doing.
> The goal isn't for a tool to handle every use case, it's to handle a specified set of use cases well.
What do you mean, "the goal"? You speak for every possible goal anyone using Python could possibly have now?
> I wanted to compare context switch, you decided to measure something else.
Then do it! I'd be interested to see the results, and even more interested to see how you tested it.
In any case, you can't just ignore tests you don't want to do or apparently aren't capable of. Being able to run a lot of threads can be extremely useful for networking applications, which is why I care. You don't get to decide my use case is "unreasonable" because you apparently can't compile a program that does a stripped down version of it.
That is to say, even if context switching is faster in processes than in threads, that doesn't mean there's no use to threads, because there are use cases where spinning up and tearing down is more common than context switching.
> I am pleasantly surprised anyway.
Ignorance is bliss, I suppose.
$ gcc threads.c
$ time ./a.out
real 0m10.097s
user 0m0.035s
sys 0m0.239s
$ gcc process.c
$ time ./a.out
real 0m10.168s
user 0m0.579s
sys 0m0.347s
You were pleasantly surprised by your own test showing an order of magnitude speed gain in user space by using threads?Please explain: In what sense is the overhead of starting actual OS processes, and relying on IPC "free", compared to running threads or even greenlets, and using shared process memory?
With Python's current multiprocessing utilities - you get a big discount by not having to write thread-safe code, or worry about synchronization, despite the GIL still being there. Very broadly speaking, it's "free" in the sense that the OS handles parallelism automatically at the process-level, and provides a simple communication mechanism between those processes through standard APIs. It also reduces potential attack surfaces (although this is a lesser argument).
It's also "free" in the sense that you don't need to re-write large parts of the VM, as-well as all supported libraries, and teach the entire Python community how to safely write and test multi-threaded code (something I bet upwards of 75% of the people using Python today won't manage) to support this specific form of parallelism.
If your goal is to run code (whether IO bound or CPU bound) in parallel - Python has the means to do that already today, without removing the GIL.
In that sense, never updating python again is "free" as well, because it would save the python devs the trouble of changing the interpreter. And yet I think we can all agree that Python benefits from the fact that we no longer use Python 3.5
> and teach the entire Python community how to safely write multi-threaded code
People who don't write threaded code don't need to worry about it. And people who write threaded code in python already need to worry about writing thread-safe code. The GIL doesn't prevent race conditions between individual python instructions.
> Python has the means to do that already
And as outlined above, these means are no suitable replacement for true thread based parallelism.
Your only argument is that "thread-based parallelism can't be achieved without threads", but that's not relevant to the conversation whatsoever.
The fact of the matter is that Python (already today) allows you to achieve parallelism across both IO-bound and CPU-bound workloads.
For CPU-bound workloads, the number of threads you can run in parallel is bound by the number of cores you have. For IO-bound workloads, your threads are just waiting on interrupts.
What are concrete use-cases where thread-based parallelism in Python is so desperately needed right now, that can't be achieved through process-based parallelism?
I'll give you one: real-time latency/throughput-sensitive DSP. Think real-time audio processing, or real-time algorithmic trading. Python isn't used there to begin with (for an entire flurry of reasons) - GIL or no GIL.
I also don't need to use goroutines. I could simply spin up my golang application as a couple of processes, and use pipes and other IPC to coordinate them.
"There is another way to do X" doesn't imply that other way is better.
I never said that multiprocessing is "better" than multithreading.
Both are mechanisms to achieve parallelism with different pros/cons, and there are legit cases where threads are a necessity that can't be satisfied with processes (ref my previous comment with examples). This conversation would've been much easier to have given specific constraints and examples (which nobody seems to give).
Within the context of a general purpose VM which was never built to support multithreading (CPython), and which carries the baggage of 30+ years' worth of 3rd-party packages and libraries which were never built with multithreading in-mind - I think we can both agree that the costs and risks of changing literally everything may outweigh the benefits. It's not a difficult stance to accept.
My main argument, if you distill it down to the abstract, is "use the right tool for the job" - and stop pretending that a hammer and a knife are the same thing, when they're not.
If you need hardcore, ultra efficient and parallelized workflows - Python is just the wrong tool for the job period. This isn't up for debate - it's a given fact. This isn't just about the GIL - it's about the type system, the frameworks, the bloat, etc. There's nothing wrong with Python being this way - Python fills a niche of its own incredibly well, something which others tools suck at, and Python is loved by millions for it.
Python is used today by certain demographics to perform certain jobs - and it excels really well at those jobs for those demographics. This whole discussion now (GIL vs. no GIL) is about bending and twisting Python to fit the use cases of few companies like DeepMind or Facebook - who I'm certain represent a miniscule usage compared to the millions of students, schools and universities, research institutes, web shops, hobbyists, tinkerers, etc. Those people want to get shit done quickly - and mutexes, sempahores, events, threads, synchronization primitives, atomics, etc - will do nothing but make their lives a misery, and drive them away.
Also, I keep asking you for concrete examples where the Python community absolutely needs multithreading (where multiprocessing fails) and you're not really responding which makes me feel like we're not conversing here...
Maybe not, but you seem to be implying that multiprocessing is a sufficient for the most case (or the average case ?) however ill define that average case is.
> This conversation would've been much easier to have given specific constraints and examples (which nobody seems to give).
Not quite, there is no need grand example here, processes vs threads is just a function of how chatty the interaction between the independent agent are. We should have both mechanism available and let the user choose.
> My main argument, if you distill it down to the abstract, is "use the right tool for the job" - and stop pretending that a hammer and a knife are the same thing, when they're not.
You don't get to decide in the abstract what the right tool is for other people and context we don't know anything about. The people pushing for proper threading are telling you that for them python + better threading is the right tool.
> If you need hardcore, ultra efficient and parallelized workflows - Python is just the wrong tool for the job period. This isn't up for debate - it's a given fact.
Maybe; But that's not the point here. Proper threading in python doesn't make python more efficient, but it makes the * scaling * of python performance more efficient. Those are two differents things.
> Python fills a niche of its own incredibly well, something which others tools suck at, and Python is loved by millions for it.
> This whole discussion now (GIL vs. no GIL) is about bending and twisting Python to fit the use cases of few companies like DeepMind or Facebook
> who I'm certain represent a miniscule usage compared to the millions of students, schools and universities, research institutes, web shops, hobbyists, tinkerers, etc. Those people want to get shit done quickly - and mutexes, sempahores, events, threads, synchronization primitives, atomics, etc - will do nothing but make their lives a misery, and drive them away.
This seems to me like the core of the problem, somehow you say that there is no valid need for threading, but at the same time the people needing threading are not important enough for python to care about... You have the magical ability to be plug into the python hivemind and decide which use case are valid, what are the true value of the python community and what python should and should not consider important... And after all that mental gymnastic you expect us to jump through hoops to convince you otherwise...
That's a lot of non technical assumption for a technical problem.
> Those people want to get shit done quickly - and mutexes, sempahores, events, threads, synchronization primitives, atomics, etc - will do nothing but make their lives a misery, and drive them away.
Not really. If anything, they will benefits from better libraries and more efficient C-extensions.
> absolutely needs multithreading
Nothing is absolutely needed. Even something are core as a type system is not absolutely needed. The fact that you ask for this level of qualification for a feature that pretty much every other language has is the problem.
The problem is not that deep, implementing multi-threading is not that complicated. I do agree that we need to be careful to provide a better API than raw-thread to end users (which no-gil is COMPLETLY agnostic about).
> "you don't get to decide in the abstract what the right tool is for other people and context we don't know anything about"
So which one is it? Do I get to ask for concrete examples for why this work is worth all the trouble, or should we just pretend like it doesn't matter because you say so?
"I want threads, threads fast, others have threads" is not good enough.
> "somehow you say that there is no valid need for threading"
I've never said that. All I've been asking for are concrete examples that would justify removing the GIL and the impact it would have on the ecosystem. You won't provide any, despite me asking over and over again.
> "You have the magical ability to be plug into the python hivemind and decide which use case are valid"
CPython has had a GIL for the past 30 years - and things worked out just fine. There are other Python VMs that don't have a GIL - yet CPython remains the most widely used and popular VM. So how about you stop putting words in my mouth, and start actually justifying the asks beyond hand-waving my arguments away?
> "The fact that you ask for this level of qualification for a feature that pretty much every other language has is the problem."
Again, you're twisting and putting words in my mouth. There's a clear difference between building an ecosystem with multithreading in mind from day one - and suddenly introducing it out of nowhere, 30 years into it being used by millions of workloads globally. This is why I'm asking for "qualifications" (which you won't provide), and this is why I also deem your assertion that "the problem is not that deep" to be a proof beyond any that we should end this conversation now before things get too embarrassing for you :)
Both ?
1 - There is no need for a grand example to justify the need for threading in a modern language, as we have 20/30 years of background on that topic. To repeat myself, it's all about the amount of communication between compute agent... The more communication you have , the more the process isolation/serialization cost become a problem.
2 - You yourself understand that there is are valid use case for threading, somehow those valid cases are not valid python uses. That's where the disconnect is. You don't get to say by decree what is and is not a "valid" use of python.
> "I want threads, threads fast, others have threads" is not good enough.
Neither is i don't use thread , you shouldn't use thread.
> CPython has had a GIL for the past 30 years - and things worked out just fine.
That's what the no-gil people are trying to tell you, no it's not fine and it was never fine. Effort and conversation about the replacement of the GIL are at least 10/15 years old.
> yet CPython remains the most widely used and popular VM Faulty logic , the correlation doesn't imply any causal relationship.
> and start actually justifying the asks beyond hand-waving my arguments away?
Because you aregument are not really argument, they more like strong opinion on things that are closer to esthetics and right/wrong usage of things. Happy to disagree on those one.
> There's a clear difference between building an ecosystem with multithreading in mind from day one - and suddenly introducing it out of nowhere, 30 years into it being used by millions of workloads globally.
1 - no-gil isn't out of nowhere. Conversation about this are more that 10 years old. Combined with even longer conversation in other VM/programming language.
2 - no-gil doesn't introduce threading in random workload. It allow people who want threading to use threading.
> This is why I'm asking for "qualifications"
It's your prerogative to ask for qualifications. But it's also our to decide to judge your ask and decide if they are worth our time.
Much in the same way that if feel like we don't need (in 2023) a grand example to justify why we need to add a type system, we don't need a deep conversation about threading. We all have the same information, understand the trade-off. We just have different value system and want different things out of python. And that's okay...
> the problem is not that deep" to be a proof beyond any that we should end this conversation now before things get too embarrassing for you :)
I think i have some comment somewhere explaining why no-gil was never a technically challenging problem. But the prof is simple... Sam Gross is definitely an exceptional dev, but the fact that a lone programmer come out with an acceptable solution is proof that the problem wasnt that deep.
No, they did not.
That's why the discussion about the GIL is about as old as Python3 itself.
The GIL has always been a major drawback of python, moreso because the language does in fact have threading support...only it can't use threads for parallel workloads.
This drawback was tolerated, because of the many advantages Python brings to the table, and because Python comes from an age when Moores Law was still in full effect; Powerful single cores were the norm not so long ago.
This isn't the case any more. Moores Law is done. Now we increase the number of cores, and languages that wish to remain relevant, need to reflect that.
How do you figure that? Even python2 already supported the usage of OS threads [1].
> 30+ years' worth of 3rd-party packages and libraries which were never built with multithreading in-mind
Many of these packages also don't use other forms of concurrency, but are simply encapsulated functionality that runs in a single thread. Meaning, they will not be bothered by the change.
Besides, as I have mentioned elsewhere, library maintainers always need to keep up with the development of the underlying language as well as usage patterns of the community, or their libraries become obsolete. That is true no matter what programming language we talk about.
> If you need hardcore, ultra efficient and parallelized workflows - Python is just the wrong tool for the job period
Python is already used as an orchestration language for huge numerical workloads, be it data science or machine learning. It is simple, intuitive and has by far the largest library support of any contemporary language.
There is simply no good argument, why the language that we entrust to orchestrate this scale of computing power, shouldn't itself be as efficient as possible for a dynamically typed script-language. That this is absolutely possible, is demonstrated by languages like Julia.
The fact that Python will never be as fast as Go, Rust or C++, doesn't change that.
> Those people want to get shit done quickly - and mutexes, sempahores, events, threads, synchronization primitives, atomics, etc - will do nothing but make their lives a misery, and drive them away.
Those people will for the most not even realize that the GIL is gone. If they write...
- single threaded synchronous code
- asyncio based code
- multiprocessing code
...the change doesn't matter to them. The hobbyists small webserver, or the medium companies Flask-based webapp will still run as before. And if they write threading code, and do so correctly, then it is very likely the only change they will see, is that suddenly their application runs faster under high load.The removal of the GIL neither takes away existing capabilities from Python, nor does it force everyone to write threading code.
> and you're not really responding which makes me feel like we're not conversing here...
That's because I have done so elsewhere in this thread already [2]
The key question that remains unanswered is
Why should Python even need a free threading model?
There are no good answers afaik.
I shall be more than happy to provide them:
Supporting thread based parallelism is the norm among all mainstream PLs with the one inglorious exception of JS, which doesn't because it simply can't.
And Python already does support it, it simply is limited by a legacy design decision. That was okay in a bygone age when fast single core machines were still the norm, Python was primarily a scripting language for when bash wasn't enough, and most relevant webapps were IO bound.
Today, there simply is no excuse any more. Python is the most ubiquitous language in the world, and running scaling web applications, is the lingua franca of ML, and orchestrates huge systems. Servers have hundreds of cores, and CPU bound workloads become ever more important.
It's about time Python rids itself of that needless limitation.
It can't for the same reason python can't - every part was designed without it in mind.
> Python already does support it
The language runtime might have it in a branch, but the vast majority of C code it is based on, and the scripts themselves assume otherwise.
> fast single core machines
Once again, if you are using python as a scripting language, with C libraries, and spawning processes, you are already utilizing multiple cores without any adjustment on your part.
> CPU bound workloads become ever more important.
If that's true, then you shouldn't use python. It's about 100x slower than C. How much performance can you squeeze out of dividing work into cooperating threads, that can't be easily achieved by having multiple processes.
In the "web application" example you are using, the norm is already to have many processes to handle incoming connections.
Interesting, care to explain then why Python supported threading since Python2? [1]
> Once again, if you are using python as a scripting language
> If that's true, then you shouldn't use python.
Once again, I don't. I use it as an orchestration language calling other code, and there is no good reason why the orchestration language should have an arbitrary bottleneck.
Yes, the hot code isn't written in Python. That doesn't matter to this discussion.
> In the "web application" example you are using, the norm is already to have many processes to handle incoming connections.
Outside of the python world, it absolutely isn't. I also have numerous Go based webservices, and they don't have to jump through ICP hoops to facilitate communications between workers and services.
I want to correct one thing that I see plastered all over this thread.
The GIL isn't a programming language construct. It's an implementation detail. The GIL isn't a "Python limitation" in any way, because it has nothing to do with Python.
CPython (aka. Cython), probably the most popular and widely used VM for Python, was built around a GIL.
There are other VMs out there that don't have a GIL (ie. Jython, IronPython) and can be used out of the box.
It's an implementation detail of CPython. Which is by far the most common Python interpreter. And it limits parallelisation via threads in Python, a language that otherwise natively supports threading.
So yes, it IS a needless limitation.
Yes it does. Unless we mean something different?
No, it doesn't.
def thread_function():
# thread may lose core here
value = store.get_value() # or anywhere inside get_value
# or here
update(value) # or anywhere inside update()
# or here
store.put_value(value) # or anywhere inside put_value()
The GIL only makes certain internal functionality atomic. It doesn't protect the implemented logic from causing a race condition.So unless I protect store with a lock, I can already get a race condition, GIL or no GIL.
The GIL doesn't solve or prevent all classes of race conditions (which can stem from complex interactions with databases, the filesystem, etc). Removing the GIL though, will only make things worse from that perspective - making the language more difficult to work with (for non-technical people).
This is exactly why Python should be further simplified, and not made more complex - given the population which uses it most.
The changes to the GIL don't force the average user to write threading code.
And when it's not ?
IPC is a PITA, and orchestrating processes is even worse.
> or at the isolated numeric library level
Not everything I want to parallelise in python runs in numpy. Simple example: WebService Backends. I have a 64 core server running a Werkzeug/Gunicorn application. The Service is mostly doing CPU bound tasks (data aggregation and analysis), so asyncio is pointless.
What happens is, it runs 60 worker processes. Which puts hefty limitations on any crosstalk and data sharing, because these either require IPC, or using redis/sql. Which are nowhere near as performant as actually shared memory would be.
I think, rightfully, the concern is people who will try to use this incorrectly causing major bloat to CPython.
Please explain who exactly is "drawn into this"?
If you want to use async, this change doesn't affect your code.
If you want to use multiprocessing, this change doesn't affect your code.
Even if you already use threading, and do it correctly, this change doesn't actually affect your code.
So who is "drawn into this"? And please don't say library developers. a) Having to update libraries to have them remain relevant, is normal procedure, in all languages. b) a lot of the people who want this to change are library devs.
"We also need to bring along the rest of the Python community as we gain those insights and make sure the changes we want to make, and the changes we want them to make, are palatable."
End quote.
Yes, library developers have to keep up with developments in the underlying language as well as changes in usage patterns by the community. That is true for all programming languages.
And as I have pointed out numerous times before, this change is on the wishlist of many libdevs in the python community.
Java is more backwards compatible, so is Common Lisp.
Someone here said that the thread on the Python "discussion" forum was shut down. That does not sound like everyone except for the proponents is supposed to be heard (or is even aware of the discussion).
Even for such stable languages, a library maintainer has to, at the very least, patch security problems as they are discovered.
> That does not sound like everyone except for the proponents is supposed to be heard (or is even aware of the discussion).
There is an official poll among the Python core devs, linked in the article, which shows overwhelming support for the change. Since they are the ones who have to work this out, that's the only discussion about this that is relevant.
Whatever this comment means (I honestly can't properly tell) - removing the GIL will have absolutely no impact on Python's resource utilization.
Quote: "In PyTorch, Python is commonly used to orchestrate ~8 GPUs and ~64 CPU threads, growing to 4k GPUs and 32k CPU threads for big models. While the heavy lifting is done outside of Python, the speed of GPUs makes even just the orchestration in Python not scalable. We often end up with 72 processes in place of one because of the GIL. Logging, debugging, and performance tuning are orders-of-magnitude more difficult in this regime, continuously causing lower developer productivity."
Quote: "We frequently battle issues with the Python GIL at DeepMind. In many of our applications, we would like to run on the order of 50-100 threads per process. However, we often see that even with fewer than 10 threads the GIL becomes the bottleneck. To work around this problem, we sometimes use subprocesses, but in many cases the inter-process communication becomes too big of an overhead. To deal with the GIL, we usually end up translating large parts of our Python codebase into C++. This is undesirable because it makes the code less accessible to researchers."
Now we change the world for everyone and put most of library developers through a valley of desperation for 5 years+, just so that a very few narrow use cases get the benefits they want.
Not a smart move IMHO.
I try to avoid python as much as possible, because I mainly work with Go & C++ and multi-threading with those languages is just better (imho). Bringing python a step forward and making it future proof might be a good thing... Even if this means to break some things? Not sure if dismissing the GIL is the right step, but there is a big performance gap to fix. Or maybe the AI community must move to a better suited language? Having python code in production just feels so wrong. Especially if a rewrite in another language shows the performance gap.
I'm not sure whether the SC has considered alternative approaches but it would be surprising if not
Alas, subinterpreters sound like they could be a feasible solution for many use cases as well.
In Python there's a much better option if you care about performance - use a different language!
The website you're using likely benefits immensely from thread-safe code. The browser you're using benefits immensely from thread-safe code. Your entire user experience using the modern internet is in an ecosystem of thread-safe code which you apparently are completely oblivious to.
There are a lot of threading models besides Java and C's as well, which you're apparently also unaware of.
Why would you assume you know all the possible use cases of a multi-purpose programming language?