Scaling NumPy on Free-Threaded Python
labs.quansight.org
labs.quansight.org
I was taken aback for a moment that this work originated from a report on StackOverflow. I had thought SO was effectively dead and abandoned by its community. But maybe I shouldn't project my own experience onto everyone else.
It's a less marketable product in the modern attention economy, but that isnt the same as a dead community. It would be interesting to see a plot of posts/discussion liveliness per unit time and seeing is quality and depth of both questions and answers has changed and how they have changed.
sum((np.sin(np.cos(np.sin(np.cos(x + i)))).sum()
for i in range(n_loop)))
This is not like: release GIL
# i = 0
compute y = x + 0 (numpy broadcasting sum)
compute z = np.cos(y) (elementwise)
...
add to running total
# i = 1
compute y = x + 1
...
reacquire GIL
Instead it is: # i = 0
look up "+" operation
release GIL
compute y = x + 0
reacquire GIL
look up "np.cos" operation
release GIL
compute z = np.cos(y)
reacquire GIL
...
# i = 1
look up "+" operation
release GIL
compute y = x + 1
reacquire GIL
...
So there was work being protected by the GIL, that suddenly is exposed to lock contention with free threading.Of course, without free threading, the lock contention would be way worse, but this time the GIL is the lock being contended. Numpy has to reacquire the GIL whenever it returns from a function call, and this expression is made up of multiple calls. To multithread effectively with numpy (in non-freethreading) you'd normally aim to vectorise into a small number of calls in big arrays.
The composition of +, then np.cos, etc. is not too bad if these are big arrays, but the problem is the pure Python iteration over the range which is, presumably, quite large. You could vectorise over the range:
x[..., None] + np.arange(n_loop, dtype=np.float64)
but this is the start of a new conversation.But this comment thread is more of a meta discussion: why would lock contention issues only show up now, in free threaded mode, if numpy already released the GIL anyway? That's what I answered above: yes the GIL was sometimes released by numpy, but these operations really were happening with the GIL locked (in non free threaded build).
It's a new issue because the reference count needs locking, whereas previously it didn't because it was implicitly synchronised due to the GIL being locked.
But the original commenter asked: hang on, I thought the GIL wasn't locked for numpy?
Now you have reached the start of the conversation.
Unclear why you still need the lock here in that case. The idea that this flag may get updated during runtime and impacts how the software works when set seems to clash with the idea we need take no action having performed a relaxed (ie non-synchronising) load and seen it wasn't set at some previous time.
Maybe there's something I don't understand about these internals, which may be as simple as "It's just advisory so if we don't trace when we should no big deal".
https://en.wikipedia.org/wiki/Amdahl%27s_law
Even a tiny bit of serial instruction will limit the speed up
Every doubling of the number of workers halves the execution time cleanly in the plot, from 40 seconds to 20 seconds to 10 seconds. Eyeballing this for 32 over 16 workers is difficult, but it still seems close to halving the total time once again. So there's not a lot of Amdahl flattening, it's just the plain physics of looking at a inverse-proportional curve.