Here's a typical example of the kinds of optimizations this guide teaches you, in this case by avoiding the creation of temporary copies of Numpy arrays in memory:
# Create two int arrays, each filled with with one billion 1's.
X = np.ones(1000000000, dtype=np.int)
Y = np.ones(1000000000, dtype=np.int)
# Add 2 * Y to X, element by element:
# Slowest
%time X = X + 2.0 * Y
100 loops, best of 3: 3.61 ms per loop
# A bit faster
%time X = X + 2 * Y
100 loops, best of 3: 3.47 ms per loop
# Much faster
%time X += 2 * Y
100 loops, best of 3: 2.79 ms per loop
# Fastest
%time np.add(X, Y, out=X); np.add(X, Y, out=X)
100 loops, best of 3: 1.57 ms per loop
That's a 2.3x speed improvement (from 3.61 ms to 1.57 ms) on a simple vector operation (your mileage will vary!).[1] This only scratches the surface. The guide goes into quite a bit of explicit detail about how Numpy arrays are constructed and stored in memory and always explains the underlying reasons why some operations are faster than others. In addition, the guide has a section titled "Beyond Numpy" that points to even more ways of improving performance, e.g., by using Cython, Numba, PyCUDA, and a range of other tools.I highly recommend reading the whole thing!
--
[1] Example copied from here: https://www.labri.fr/perso/nrougier/from-python-to-numpy/#an...