How fast can we make interpreted Python?
arxiv.org
arxiv.org
I guess what I'm trying to say is, while this is a laudable effort, it seems what Python (and Ruby) really needs is a way to free itself from the chains of extension API compatibility.
My experience with numerical algorithms in Python has been a 10000x speedup and 10x memory reduction by using Numpy arrays and C/fortran functions. And when Numpy doesn't have a necessary function, a few extra lines of static C types and on-import compilation with Cython has still resulted in a 1000x speedup.
"Optimizing Python" needs to look at the whole existing ecosystem of C libraries - to be useful the end result has to be faster than what's possible now, not just faster than pure Python.
The speed improvements are worth it given how easy it is to use, I have seen literally a 1000x speedup on a time-series data analysis script from adding about 10 lines of static type annotations.
The functionality I like best about it are:
* It can be used transparently without any makefiles or compilation steps. Add "import pyximport; pyximport.install()" at the top of your main script, and now every imported Cython-capable module is compiled on the fly at runtime.
* You can start out with no changes to your python modules. All libraries and features still work within compiled modules. Then you can slowly start adding static types to a few variables at a time. The annotated variables become very fast native C integers/doubles/functions, instead of Python objects.
* Install it with pip, the packages in your OS distribution may be out of date,
* Rename your numeric code module's .py file to .pyx,
* Use pyximport from the main script ( http://docs.cython.org/src/userguide/source_files_and_compil...) to have it compiled at runtime without any extra build steps, and then
* Start experimenting with adding a few "cdef"s (http://docs.cython.org/src/quickstart/cythonize.html and http://docs.cython.org/src/tutorial/numpy.html)
> Not yet! Lots of constructs (like catching exceptions, constructing objects, etc...) aren't implemented in the Falcon virtual machine.
"... However, this doesn't mean that programs which use these constructs won't run. Any missing functionality is routed through the Python C API, foregoing any potential performance benefit you might have gotten from Falcon. So, though Falcon isn't a complete Python implementation, it should still run all of your code."
https://github.com/numba/numba
Intro to Numba, parts 1 and 2:
The big performance win is nice (this benchmark:
https://github.com/rjpower/falcon/blob/master/benchmarks/mat...
), but it is still going to be a lot slower than doing it the horrible way (calling out to Numpy).
The benchmark is more aimed at demonstrating the reduction in loop overhead than it is at numerical speed, but numerical stuff is usually where you end up with a lot of loop overhead...
There was just something magical about Psyco, just adding two lines to your script and seeing it run 100x faster. It was really great at speeding up floating point math.