When you ("you" in this comment is probably most usefully read as "a compiler" :) ) monomorphize a generic function or data structure with the various type parameters that it's been called with, you make a copy of the code for every set of type parameters it's been called with.
Each of these copies requires space for the instructions...this means a larger binary on disk, more memory required when it runs, and more pressure on the CPU's caches, especially precious L1I.
The cache pressure issue is typically the big one worth thinking about, although binary size itself can be an issue for some cases - I've primarily got embedded use cases in mind, but I'm sure there are others I'm not thinking of.
Anyway, back to cache pressure. If there are a whole bunch of different monomorphizations of a given function, and they're all called frequently, that could mean that the CPU will frequently need to refer to slower L2$, or much slower L3$ (or much much slower RAM) to load instructions. That's no bueno from a performance standpoint.
Because of this, there are cases when dynamic dispatch can outperform monomorphization - it's definitely not as simple as "vtable slower, monomorphization faster" across the board, even if that is an OK rule of thumb.