Uncommon words have more information content than common words. So, common words having larger embedding scale is an issue here.
If you want to measure similarity you need a scale free measure. Cosine similarity (angle distance) does it without normalizing.
If you normalize your vectors, cosine similarity is the same as Euclidean distance. Normalizing your vectors also leads to information destruction, which we'd rather avoid.
There's no real hard theory why the angle between embeddings is meaningful beyond this practical knowledge to my understanding.