10% doesn't look too much to me, I still don't get why today people care so much about single thread performance.
10% doesn't look too much to me, I still don't get why today people care so much about single thread performance.
If no-GIL has a 10% single thread performance hit, that means that essentially all my existing python code would be that much worse.
Why havent you implemented multithreading?
(don't get me wrong, I know the cost of implementation, but if speed matters, multithreading is a very reasonable step in python)
Because that makes programs slower in Python.
Multi-threading in Python is for when you need time-slicing for CPU intensive tasks so that they don't block other work that needs to be be done.
A lot of things are though...
With threads, I can encapsulate the use of threads in a class, whose clients never even notice that threads are in use. Sure, threads are a global resource too, but much of the time you can get away with pretending that they're not and create them on demand. Not so with multiprocess. If you use that, then the whole program has to be onboard with it.
Threads work great in Python. Well not for maximising multicore performance, of course, but for other things, for structuring programs they're great. Just shuttle work items and results back and forth using queue.Queue, and you're golden - Python threads are super reliable. And if the threads are doing mostly GIL-releasing stuff, then even multicore performance can be good.
Huh? In Python you just need a function to call, and multiprocess will run it wrapped in a process from the pool, while api-wise it would look as it would if it was a threadpool (but with no sharing in the process case, obviously).
So what would the rest of the program be onboard with?
And all this could also be hidden inside some subpackage within your package, the rest of the program doesn't need to know anything, except to collect the results.
That means you either use fork - which is a major can of worms for a reusable library to use.
Or you write something like this in your entry point module:
if __name__=='__main__':
multiprocessing.freeze_support()
once_only_application_code()
Suppose I don't realise that your library is using multiprocessing, and I carelessly call it from this two-line script: import library_that_uses_multiprocessing_internally
library_that_uses_multiprocessing_internally.do_stuff()
That's basically a fork bomb.And where do you put the multiprocessing.set_start_method call? Surely not in the library.
Huh? As far as I remember multiprocessing just sends pickled versions of the function to run and any of its dependencies (other functions, closures, etc). As long as the function doesn't use global state that's not available when pickled, it's fine. But it doesn't re-initialize your whole program for each process in the pool.
>That means you either use fork - which is a major can of worms for a reusable library to use.
How did we get into a reusable library authoring?
Yes, multiprocessing is not just turnkey to use inside a reusable library you make.
But the context were programs here, or not?
>Or you write something like this in your entry point module:
Hmmm? This is to have it support freezing the script (that is, using a tool to make it a distributable, like PyInstaller). That's not necessarily a use case most have.
That was always my premise. Maybe I didn't make it clear enough, because I tend to just take it for granted that that's how you write code, in a style that's suited for reuse.
> But the context were programs here, or not?
That's the "global" I was talking about: Code that's using multiprocessing needs to know the context that it's embedded in. Any moment I might grab that piece of code and transfer it to a library of reusable components, because that's how I work - code that starts out as part of a standalone program doesn't necessarily stay that way. Multiprocessing gets in the way of that.
That's somewhat condescending. You can write code "in a style that's suitable for reuse" without being a library author - well, without publishing public packages anyway. Re-use is not only about some totally generic package that can run under any arbitrary context willy nilly.
And of course there are tons of programs where the parts don't make sense as libraries, because they're tied to the specific functionality and overall design (whether because of the domain logic required or due to optimization or other constraints). You write them to be modular and clean, but not with "arbitrary people running my code in whatever context" in mind.
Not to mention the mountains of purpose-specific throwaway scripts, e.g. in the scientific community especially, where Python is big, there's little regard for reuse (even less so for library building), and it's not because multiprocess is stopping them :)
So, yeah, I'd say, even if not 100% suitable for generic reusable library-style code, it doesn't mean multiprocess can't be applied in a huge number of specific people's problems and codebases.
>Code that's using multiprocessing needs to know the context that it's embedded in.
If you want to speed up your Python program and there's something that can run in parallel with no shared state, you can use multiprocess to run it.
If having it as a "reusable component" that hides away the fact multiprocess is used, and that can be called in any arbitrary context, is your concern, it's a valid one, but then perhaps a specific Python program and its performance is not your main priority. Library writing is, instead :)
Else, it's enough that the user calling multiprocessing knows the function that is to be passed and its dependencies (or lack thereof). Other than that, they don't have to change their top level program's architecture.
I'm really looking forward to subinterpreters. I think they have great potential for supporting a style of multiprocessing that is both faster and better isolated.
I think you would love Trio and applying the idea to threads.
Maybe in a 100% CPU bound code, most of the code is I/O bound and no one will notice the change, just my opinion.
So, the hit on the Python interpreter wouldn't translate to a hit on those.
So? Especially since the "Faster Python" team already made Python 1.11 "10–60% Faster than 3.10", and 1.12 is even faster still, whereas their overall plan is to get it to 2-5 times faster compared to 3.9.
So at the worst case, with a 10% hit, you'd balance out the 3.11 speed, and your code would be as fast as 3.10.
There's no absolute objective "needs to" or even any static baseline. Python can have, and often has had, a performance regression that drop your code by 10% at any time. It's no big deal in itself.
Also consider a further speedup of e.g. 50% in upcoming versions (they have promised more).
If you're OK with the X speed of today's Python, you should be ok with X + 40% - even if it's not the X + 50% it could have been due to the 10% GIL's removal toll.
The impact won't be on users / Python programmers who don't develop native extensions. It will suck for people who had a painful workaround for Pythons crappy parallelism already, but now will have to have two workarounds for different kinds of brokenness. It still pays off to make these native extensions, however their authors will create a lot of problems for the clueless majority of Python users, which will like end up in some magical "workarounds" for problems introduced by this change very few people understand. This will result in more cargo cult in the community that's already on the witch hunt.
I was under the impression that the Python thread scheduler is dependent on the host OS (rather than being intelligent enough to magically schedule away race conditions, deadlocks, etc.), so you still need to manage locks, semaphores, etc. if you write multi-threaded Python code. I don't see how removing the GIL would make this any worse. (Maybe make it slightly harder to debug, but at that point it would be in-line with debugging multi-threaded Java/C/etc. code.)
Or would this affect single-threaded code somehow?
You need to do that with multithreaded Python code with the GIL. The GIL only guarantees that operations that take a single bytecode are thread-safe. But many common operations (including built-in operators, functions, and methods on built-in types) take more than one bytecode.
For about 10 minutes a few years ago, when the M1 had the best single threaded performance per buck, people cared.
Now that the M1 isnt the leader in single threaded, we are back to the 'multithread is most important'.
Which has always been true. If your program needs an improvement in speed, you can multithread it. The opposite isnt true.
What do you mean by "the opposite"? "If your program doesn't need an improvement in speed, you can't multithread it"? "If you can multithread your program, then it doesn't need an improvement in speed"? Well, yeah, obviously both of those statements are false but they're also quite useless, so who cares?