This particular problem is so embarrassingly parallel, I am surprised the speedup is only 14x.
http://www.hpcwire.com/hpcwire/2011-12-13/ten_ways_to_fool_t...
In my experience, it's usually more like 2x-5x for a straight translation of the algorithm from numpy to C. (Though I have to confess that I often write rather naive C...)
Numpy performs quite well if you avoid a few common pitfalls. It certainly can be beat, but it's often faster than people think it is.
Things like optimizing memory accesses to prevent bank conflicts, using per thread-group shared memory effectively to reduce global memory accesses are a bit tricky to get right, and I'm not sure how well auto-generated code would be able to do this.