Speeding up Cython with SIMD
pythonspeed.com
pythonspeed.com
If you're writing a compiled extension, there's pybind11 or, better yet, pyo3, where you can also get access to internals of polars and other libraries.
In numeric Python, and especially with computers becoming progressively faster, it's rarely the case than the layer of pure Python is where the cpu time is spent. And in the rare case when you need to do something funky for-loop-style with your numpy data, there's jit-compiled numba...
I have a feeling this is why ultimately why PEP 703 is being accepted despite the setbacks to the Faster CPython effort-- while faster CPython is a great goal, it is rarely the bottleneck nowadays.
How much of a setback that should be? The initial target was a 5x improvement over 4 releases, IIRC. What are current estimates like?
-mavx2 -mfma -mavx512pf -msse4.2 etc
Eventually I'll write that other article; I've been wondering if it's possible to have infrastructure to support both modern and old CPUs in Python libraries without doing runtime dispatch on the C level, so this may involve some coding if I have time.
I know, thats not the point and average was only picked as a simple example, but still...
Just tested how would it be without compile nonsense.
```
a = np.random.random(int(1e6))
%%timtit
np.average(a)
%timeit
np.average(a[::16])
```
And my result is that no matter how uncontiguous in memory (here I take every 16 elements like what they did, and I tested for 2,4,8,16), we are doing less operations so it always end up faster. Contrastingly their SIMD compiled code is 10-20X slower in uncontiguous case.
And for a larger array that is 16X of the contiguous one, but we only take 1/16 of its element, the result is like 10X slower as shown by the article. But I suspect that purely now you have a 16X larger array to load from memory, which itself is slow in nature.
```
b = np.random.random(int(16e6))
np.average(b[::16])
```
Which conclude that people should use Numpy in the right way. It is really hard to beat pure numpy speed.
I can only imagine that this is already backed into Numpy now.
But that's not the function in the article. The article implements `(a + b) / 2`.
And, on my system, simple `return (arr1 + arr2) / 2` takes 1.2ms, while the `average_arrays_4` takes 0.74ms.
Could this be the reason?