I'm currently busy with my MSc thesis on learning to link plain text documents to semantically relevant Wikipedia articles (and some of the cool machine learning-y things you can do from there).
I have 2 questions about your work:
1. I'm not sure if you're familiar with Doc2Vec but it allows you to train a Word2Vec model while also learning a vector for each document in the training corpus. Wikipedia is commonly used as a training corpus so you can get a "DocVec" for each Wikipedia article in the same vector space as your Word2Vec model (i.e. the DocVec for the Wikipedia page "Machine Learning" is nearby the WordVec for "mathematics"). Did you consider/compare with using Doc2Vec to learn a vector for each Wikipedia page and then use those as your entity vectors?
2. Your "Features" page says you convert an entity name to a link pointing to an entity if the entity name is unambiguou". In the case of ambiguous entities from a link (which happens often - my research is only learning links to Wikipedia articles from plain text documents), did you consider using the entity vectors (or some simple model built on top of the word vectors of the target page) to disambiguate?