How about Dylan? :-)
I think one of the nice thing about Julia's "just ahead of time" monomorphization/devirtualization is that it allows a level of dynamism that also works on GPUs/TPUs. This post and linked paper help me understand a little bit of it: https://discourse.julialang.org/t/julia-inference-lattice-vs...
Is this level of dynamism required for conventional ML? Probably not. For physics-informed ML and probabilistic languages? Probably more likely.