Primer on Neural Network Models for Natural Language Processing[pdf]
u.cs.biu.ac.il
u.cs.biu.ac.il
http://arxiv.org/pdf/1402.3722v1.pdf https://levyomer.wordpress.com/2014/04/25/word2vec-explained...
However as I've understood it, the negative-sampling is a big part in why those models are so calculation-efficient, combined with Hierarchical Softmax to reduce the complexity further.
Note that negative-sampling and hierarchical-softmax are actually alternative choices to interpret the hidden-layer and to arrive at error-values to back-propagate. Each can be used completely independently.
If you enable them both, you're training two independent hidden layers, which then in an interleaved fashion update the same shared input-vectors. (Essentially, it's joint training of each example via the hierarchical-softmax codepath to nudge the vectors, then via the separate negative-sampling codepath to nudge the vectors.) So the actual combination doesn't reduce the complexity – it's additive to model state size and training time – and I think most projects with large amounts of data just use one or the other (usually just negative-sampling).
However, I would still not agree that the comment-linked article explaining negative sampling really explains how word2vec works, well enough, or maybe I just didn't understand.
Either way I recommend looking at this article as well if anyone wants to understand word2vec. http://www-personal.umich.edu/~ronxin/pdf/w2vexp.pdf
Nice one.
I would gladly welcome if you or someone could write a guide that has the comprehensiveness of the PDF above but with more NLP domain-specific discussion and concrete examples.