Accelerating Python Libraries with Numba (Part 2)
continuum.io
continuum.io
Here is the direct link to the notebook: https://www.wakari.io/sharing/bundle/aron/Accelerating_Pytho...
from the blog:
https://gist.github.com/ahmadia/5638980
Update:
At the request of several commenters, here is a test script and benchmarks that we ran on PyPy and Anaconda Python (with Numba). The results are not tuned (I am not a PyPy expert!) so we did not post them in the blog. We’d be happy to look deeper into this with the PyPy developers. While PyPy is not currently installed on Wakari, we are looking at a number of ways we can install and support the PyPy community.
Update:
Well this is interesting ;D
time python test.py
real 0m11.169s
user 0m11.141s
sys 0m0.022s
time pypy test.py
real 0m0.259s
user 0m0.239s
sys 0m0.018s
cat test.py def python_sum(y): N = len(y) x = y[0] for i in xrange(1,N): x += y[i] return x
python_sum(xrange(100000000))
It'd be interesting to see a comparison to Matlab's JIT. Is Numba competitive?
The C was compiled with the same flags used to compile Python. In this case: -O2 -g.
I don't have access to a MATLAB license to compare, but I would love to see this comparison done. Let me know if you need any help putting it together.
>>> x = list(range(100000000)) + ["break"]
After some time, you get a list with a hundred million integers with a string at the end. >>> len(x)
100000001
Now take the sum: >>> sum(x)
On my machine, I get a traceback, but after about half a second of hesitation. Python's been iterating through each of those elements, testing to see if it's an integer and building a running sum, only to throw an exception because it didn't expect the string at the end. That means that in order to avoid memory corruption (or other terrible fates worse than an exception), Python must double-check each element of the list as it goes; it cannot take shortcuts.