Man this community and it's diverse group of amazing people. I don't there is a single place on the internet as valuable as this community.
Sometimes I just have to pinch myself...
I weep to think of of all the cycles wasted on expensive tax payer funded hardware due to using Python over much faster Intel Fortran. We are talking like orders of magnitude.
I once ported a fellow undergrads astronomy program from Python to C+CUDA and it ran on their personal workstation in less than a day, when it had been using the department cluster for a week before :/
People often hate on Fortran because its old and creeky when it often as fast or faster than C...
Even then, in a lot of cases programmer time is infinitely more expensive than CPU time, so it often still makes sense if Python is the more accessible choice.
I think this is generally true outside of scientific computing, but often not the case here. Esp. if you consider the low salaries of scientists in the public service compared to private industry.
Even in the private sector I have routinely made things an order of mag. faster saving the need to massively scale out. People also forget about total cost of ownership, 10 servers costs much less than 10x what 100 servers require operate when you consider cooling, networking, power, part replacement etc.
But I agree, if you're writing a massively parallel simulation in Python and running it on a supercomputer it's a net waste of money. There's a big trend in the government projects to build libraries like Trilinos or MOOSE, which work kind of along the same philosophy as NumPy - let specialists implement the parts that really matter (in terms of computational expense) in a compiled language and let the scientists glue together those parts in a higher level language. It's a good strategy IMO, especially for technologies like CUDA where there's a huge benefit but also a steep learning curve.