"Go fast" has multiple dimensions.
For one example, I had a command-line tool which needed to compute something related to hypergeometric distribution. (It's been a few years; I forget the details.)
This available in scipy, which need numpy. Most of my program's run-time was spent importing numpy. (The following timings are best-of-3.)
% time python -c pass
0.027u 0.011s 0:00.04 75.0% 0+0k 0+0io 0pf+0w
% time python -c "import numpy"
0.212u 0.059s 0:00.17 152.9% 0+0k 0+0io 14pf+0w
% time python -c "import scipy"
0.252u 0.077s 0:00.32 100.0% 0+0k 0+0io 14pf+0w
While 0.2 seconds doesn't seem like much to people used to spending hours developing a notebook, or running some large matrix computation, I could make my program 8x faster by writing the dozen or so lines I needed to evaluate that function myself.
Numpy is not about making short-lived programs fast.
In any case, this specific choice of importing all submodules is not to cache things up "for obvious [performance] reasons", because there is no machine performance improvements.
Instead, it's to make the API easier/faster to use. When you're in a notebook and you need numpy.foo.bar() you can just use it, and not have to go up and use an import statement first.