>given that more code ≈ more cache misses
More code isn't randomly jumped around in, so there are not necessarily (or in practice) more cache misses. Hot paths in Julia compile down better than in many other languages (including C/C++/FORTRAN, where aliasing is much more of a problem), specifically since multiple dispatch gets resolved as much as possible (and often completely) at compile time.
If anything, by making per type versions of things, there are less cache misses since things that are similar stay together, needing less instructions on hot paths to look up types and do other stuff that other systems do hit.
The proof is in the pudding - write (or find) some good code doing similarly complex things, and test them. Julia has C/FORTRAN speeds (and better) for a lot of important tasks, with the development flexibility of Python.