On GPU/TPU, it is not going to reach perfect 100% hardware usage, but it is going to get close enough (far above vanilla Python performance) and be significantly more productive than alternatives.
That makes it a sweet spot for research (where you will want to tweak things as you go) and extremely complex codes (where you already need to put your full focus on the correctness of the code). I highly recommend it to domain experts who need performance for their research project.
I write in a mix of C++ and Python, and have also dabbled with tokamaks, and I think this is a (common) misunderstanding.
Fundamentally, you are optimizing the speed of project progress. Now, if you don't need your results in real time, e.g. because you are building a simulator which is not in a control loop for instance, you are often better at taking the easy language and ignoring the compute performance of the language. The compute performance of languages is a fixed multiplier, and with numerical code you might see things up to 10x.
But having readable code, which is easy to manipulate and change to test ideas, is speeding up the project progress by such a large factor, that it is hard to keep up with other languages. The reason python is omnipresent in machine learning, is not for a lack of trying of other languages. Python is just very good at allowing you to keep up with a fast-moving field.
The metric being optimized is not just performance, but also the ability to build reasonably performant workflows with arbitrary differentiable (i.e., ML) inputs and outputs.