1) That particular example is 10 or 20 seconds on CPython, depending on CPython version, machine, etc.
2) With the exception of Luajit and JS, any dynamic interpreted language will take in the range of 5-20 secs to run equivalent code. I'm pretty sure Ruby and PHP are in the same range, no time to benchmark it now. Besides, if you use the Pythonic version of this loop (using list comprehensions), it will take more like 5 or 10 seconds and not 20. Run the same code on Pypy, and it takes under a second.
3) In any case, if performance is really an issue, the correct answer (as the parent discovered) is to use Numpy or Cython or C extensions for that part of the code. This is not uncommon knowledge.
4) I've been using Python in production for a decade and have never thought of this (~0.2 vs ~10 secs) as a problem because I never used pure Python for critical numeric code (or Ruby, or even JS). I did that in C or Cython, or Numpy (or OCaml, on one occasion). If I had to do something like that now, I'd probably use Rust and call it from Python (assuming I was writing the app in Python).
5) Any language has edge cases like this that can surprise you if you're not familiar with the ecosystem. The higher the level of abstraction, the more edge cases. Look at even the Julia thread directly below this one. It's a massive difference between the naive, Julia beginner version (no offense, I would have written it the same way, probably) and the experienced Julia programmer version. Even a language like Go has performance gotchas.
Python certainly has it's problems (like building large, robust programs, although that's being actively addressed with MyPy types), but I honestly don't agree that this is one of them.