Although, I suppose if we treat "apple" and "Apple" as different words, that would help.
Fun fact: One of the current NLP problems is detecting which words are names. Apparently it's really tough, especially with Twitter data!
Unfortunately I see a problem in having to specify an exact position per word. If you think of the position of english "Apple" in the Spanish word space as a distribution instead of a specific location, then it ideally should be a two-mode distribution, with one peak next to Apple and one peak next to manzana. If you must use a normal distribution, the variance must be wide enough to cover both words -- a huge problem, since (a) that assigns a lot of probable values to one word and (b) the mean value (expected value) lies between them, not at the semantic location of "apple" at all.
An interesting paper looked at how these associations changed over time [1]. It was also featured recently on The Morning Paper [2], in case you prefer a summary with added context.
Although those ambiguities make things a bit more difficult, you can usually leave the job of disentangling them to a later stage in the language-modeling process, which will have more context it can use to disambiguate which word sense was used.
[1] https://arxiv.org/abs/1703.00607
[2] https://blog.acolyer.org/2018/02/22/dynamic-word-embeddings-...