NumbaPro CUDA Python speeds up a Monte Carlo pricer 14x
continuum.io
continuum.io
http://www.wallstreetandtech.com/it-infrastructure/bloomberg...
http://developer.download.nvidia.com/compute/cuda/2_2/sdk/we...
In my experience, it's usually more like 2x-5x for a straight translation of the algorithm from numpy to C. (Though I have to confess that I often write rather naive C...)
Numpy performs quite well if you avoid a few common pitfalls. It certainly can be beat, but it's often faster than people think it is.
Things like optimizing memory accesses to prevent bank conflicts, using per thread-group shared memory effectively to reduce global memory accesses are a bit tricky to get right, and I'm not sure how well auto-generated code would be able to do this.
http://www.hpcwire.com/hpcwire/2011-12-13/ten_ways_to_fool_t...
Do you have any good links that elaborate on this kind of technique? I had an idea that sounds a lot like this once, but being fairly ignorant about the topic I had no idea what to call it or how to implement it.
Thanks.