Parallelizing Word2Vec in Multi-Core and Many-Core Architectures
arxiv.org
arxiv.org
Performance takes a big hit however: gensim is ~3X faster, mainly because it uses BLAS primitives in C.
If you train on the big dataset that is produced by demo-train-big-model-v1.sh (which includes news corpora from 2012 & 2013, the 1BN word dataset from statmt, the UMBC web corpus and the whole wikipedia) using only one thread, accuracy on the google analogy dataset drops to 68% (down from ~71.5% when using 20 threads.)
This is due to the learning rate algorithm used: Learning rate is linearly reduced with the number of processed words. When K threads are used, the input dataset is split into K parts, processed in parallel, which means that more parts of the dataset have a chance to influence the resulting vectors in the beginning of the computation (where learning rate is relatively high.)
In the limit, performance would indeed suffer as all updates would happen in parallel.
[1] Recht, Benjamin, et al. "Hogwild: A lock-free approach to parallelizing stochastic gradient descent." Advances in Neural Information Processing Systems. 2011.
The summary would be that accuracy per pass suffers slightly, but since the speedup is close to linear for the first dozen or so cores, each pass is much faster to run. The result is that the wall time to achieve a given level of accuracy is much shorter despite the slightly lower accuracy per pass.
If the authors are reading this, it'd be nice to see the actual accuracy on the google analogy dataset (the paper states it is within 1% of the reference implementation) and the performance on the large dataset produced by demo-train-big-model-v1.sh.
Incidentally, at Yahoo we can learn from a dataset of 1066 Billion Words, on a dictionary of 1.42 Billion unique terms, in 7344" (~145M words/sec.)
Tokens are randomly distributed over a set of workers. Each worker iterates over its edges in parallel with all other workers and performs the appropriate computation.
Drawing negative samples is done in two steps. We first draw a worker W from a suitable distribution over the workers and then draw a word from W. The overall word sampling is the same as for the reference implementation (ie, unigram distribution raised to 3/4.)
This work will soon be made public [1].
[1] Stergios Stergiou, Zygimantas Straznickas, Rolina Wu and Kostas Tsioutsiouliklis, ``Distributed Negative Sampling for Word Embeddings''. AAAI 2017.