Excellent article, thank you. My snag in thinking about word2vec is how the vector model stores information about words with multiple, significantly different definitions, such as 'polish', Eastern Europe or glistening clean.
[1] https://nlpprogress.com/english/word_sense_disambiguation.ht...
Edit: formatting