Automatic SIMD vectorization support in PyPy
morepypy.blogspot.com
morepypy.blogspot.com
Their framework is (theoretically) able to trace any Python operation, no matter how dynamic, and speed it up. This means that if you want to speed up pure Python code, PyPy is really the only game in town.
The downside is that when scientists use Python, Python is used as a beautiful API on top of very optimized code written in Fortran or C. If you want to do numerically complex code in Python, you are much, much better off using numba. numba is much less ambitious than PyPy - it handles a small subset of Python (basically, NumPy) but it is very, very fast and very efficient at speeding up pure NumPy code. In my experience, 100x speed ups (over pure NumPy code) are not that uncommon.
The founder of continuum (Travis Oliphant) wrote a blog about his technical vision for a Python jit: http://technicaldiscovery.blogspot.it/2012/08/numba-and-llvm... and http://technicaldiscovery.blogspot.it/2012/07/more-pypy-disc...). Basically, the continuum team made a big bet that a very efficient JIT that targets only numerical python would be more useful than a generic JIT that can theoretically handle all of Python. For my use case (scientific coding) - numba is far superior.
In the short term Numba is much more practical for numerics. In the longer term Pyston looks promising - it's actually similar to Numba in that it also uses LLVM, I imagine there could be synergy between the two...
Pandas is one, others include scikit-learn, scikit-image, Astropy, Bioinformatics libraries, stats libraries, etc... which all have heavy C/Cython use and depend to varying degrees on the Python C-api. Porting NumPy barely scratches the surface of scientific python.
Going on a tangent, and though I realize it's a lost battle, I wish people would stop saying that NumPy is the base of scientific programming in Python. As Biopython shows, it isn't required for at least some of bioinformatics.
My own research[1] deals with chemical graphs, and NumPy/SciPy/etc. are nearly irrelevant to that research.
[1] For example, given a set of 100 structures, what is the largest substructure (based on the number of bonds) which is in at least 90 of the structures?
It is indeed fairly easy to create a highly performant JIT for numerical operations on N-dimensional arrays (vector,matrix..etc..); and depending on your field, it might represent 99% of the execution time.
For example, I wrote a really simple JIT compiler for Matlab which performs sometimes better than raw C (and it's backed by LLVM vectorizer to generate the correct SIMD assembly). Link to the Master thesis if you are interested: https://www.dropbox.com/s/caz7d4d08xhbwcu/thesis.pdf?dl=0
So what the graphs end up highlighting is that the CPython (NumPy?) has decent overhead on tiny calls that don't really matter for gross performance. As you say, 128 elements is tiny: that's where I'd consider starting a graph, not ending it. At 4 elements, you're looking at about 40 clock cycles to do a simple vector-add in the scalar case (I think x86 L1 cache hits are ~3 clock cycles and there are two load/store units; if these numbers are wrong, my estimate is off). This also means that there's absolutely no measurement of the overhead of the JIT autovectorizing operation (!), which is quite significant because the usual resistance to adding autovectorizers in JITs [1] is that they're way too slow for the speedup they give.
[1] As far as I'm aware, the only other JIT of note that includes an autovectorizer is the Hotspot JVM. The approach taken by other people (e.g., .NET CLR) is to effectively expose a SIMD primitive in the bytecode and get compilers to target those via static autovectorization instead of autovectorizing at runtime.
128 elements is still useful for when you have an inner loop that has a small vector operation. Most of the time this is not the case though.
I definitely agree that overhead should be included in the article, you wouldn't expect to see it in those graphs, because it's a once off, right?
Vectors are only fixed-width, so even though it is an "embarrassingly parallel" problem, you still only expect it to asymptote towards a fixed speedup. Moreover, the bigger the matrix, the more likely you are to see cache size and memory bandwidth effects in your performance numbers, meaning the SIMD could be less of a win in the limit.
There are a lot more details sprinkled across several pages: http://morepypy.blogspot.mx/2013/11/numpy-status-update.html http://pypy.org/numpydonate.html
Not sure about 4.0, but I think it was the working number for the next release (to avoid confusion with Python 3, I guess they call that pypy3 and have a version number also).
The old scheme (PyPy 2.x) resembles Python 2 versioning too much (PyPy 2.6 vs. Python 2.7) and may cause confusion that PyPy version correspond to Python version it implements (while in fact it is not). Say, if the next release is PyPy 2.7, some may assume it is a PyPy implementation of Python 2.7 (even though PyPy 2.6 is already Python 2.7). The situation is also confusing with Python 3 as well (PyPy3 2.6 for Python 3.2.)
I think PyPy's initial plan was to use YY.MM for versioning, so the tentative version was 15.11 but now looks like they decided to follow the old scheme but with major version > 3 instead (so the next release is PyPy 4.0.0 and PyPy3 4.0.0).