Wikipedia2Vec: Optimized Implementation for Learning Embeddings from Wikipedia
arxiv.org
arxiv.org
I recommend to use T-SNE instead of PCA, which can be selected by the button at the bottom left.
https://en.m.wikipedia.org/wiki/Word_embedding
What is an "entity"?
And yes that page is the embedding definition they are referring to.
Given an input text, what would be a good way to extract a list of entities? The word sequence should be usable to determine which is an entity or just a word.
Is there the possibility to do fine tuning a la Bert or Elmo?
No reason you can’t fine tune these on your task, that’s true of any word embeddings not just Bert or Elmo.
You need to extract entity names using an NER software (e.g., SpaCy, Stanford NER), and resolve the names to knowledge base entities using the entity linking method.