If I understand this correctly, saying that vowels tend to be close to vowels is just a special case of using n-grams. If you have a fingerprint with the common n-gram distribution for the target language (or even subject), you get an optimization problem where you try to guess the substitutions such that the angle between the fingerprint vectors are minimized.
If it cannot be solved analytically, it seems something like a GA should solve it well.
Is there a standard method for solving substitution cryptos?