Word vectors are awesome but you don’t need a neural network to find them
multithreaded.stitchfix.com
multithreaded.stitchfix.com
This is all well and good, but from an industry practitioner standpoint this doesn't explain why one would avoid using, or actually stop using word2vec.
1. Several known good word2vec implementations exist, the complexity of the technique doesn't really matter as you can just pick one of these and use it.
2. Pretrained word vectors produced from word2vec and newer algorithms exist for many languages.
Why should someone stop using these and instead spend time implementing a simpler method that produces maybe good enough vectors? Being a simpler method isn't a reason in of itself.
Also, Glove (which is one of the newer word2vec alternatives you mentioned) works pretty similar to the approach in the blog post, but has some tweaks that make it perform better than this.
That depends on your corpus and vocabulary size. word2vec is O(n) where n is the length of the training corpus. Vanilla SVD is O(mn^2) for a co-occurence matrix size of m x n. I have been training word2vec and GloVe embeddings on 'web-scale' corpora and training time is usually not a problem.
the SVD approach is often much faster and almost as effective
SVD-derived embeddings does bad on some tasks, such as analogy tasks (see Levy & Goldberg, 2014).
I agree with your parent poster that the author does not really provide good arguments against the use of word2vec et al.
Moreover,
But because of advances in our understanding of word2vec, computing word vectors now takes fifteen minutes on a single run-of-the-mill computer
Which advances in our understanding? SVD on word-word co-occurrence matrices was proposed by Schütze in 1992 ;). There have been many works since then exploring various co-occurrence measures (including PMI) in combination with SVD.
But on twitter, Chris promised a sequel post, with extra tricks and tips. So I see this as a warm up :) Looking forward to part 2.
And your blog post is much better :)
SVD must be done at-once, and you need to use sparse matrix abstractions for raw word vectors. The implementation and abstractions you use actually make it more complex than word2vec imho.
Word2vec can train off pretty much any type of sequence. You can adjust the learning rate on the fly (to emphasize earlier/later events), stop or start incremental training, and with Doc2Vec you can train embeddings for more abstract tokens in a much more straightforward manner (doc ids, user ids, etc.)
While word2vec embeddings are not always reproducible, it is much more stable with the addition of new training data. This is key if you want some stability in a production system over time.
Also, somebody edited the title of the article, thanks! The original title of "Stop using word2vec" is click-bait FUD rubbish. I think in this case we're trying too hard to wring a good discussion out of a bad article.
There is a temptation to use just the word pair counts, skipping SVD, but it won't yield in the best results. Creating vectors not only compresses data, but also finds general patterns. This compression is super important for less frequent words (otherwise we get a lot of overfitting). See "Why do low dimensional embeddings work better than high-dimensional ones?" from http://www.offconvex.org/2016/02/14/word-embeddings-2/.
Are you familiar with http://www.offconvex.org/2016/07/10/embeddingspolysemy/?
I read a blog post maybe a year ago that explained word embeddings under an 'atom interpretation', in which word embeddings are really sparse combinations of 'atoms' (I don't really remember more than that). It was very interesting, but then I forgot about it. Trying to find the post again I only came up with the above paper, which is probably the same idea, but it wasn't what I originally read. Wish I could find it.
It's a interesting article but the author didn't really provide good arguments why I should stop using w2v.
[0] http://www.kdd.org/kdd2017/papers/view/struc2vec-learning-no... [1] https://arxiv.org/abs/1603.07185
Word2vec is not deep learning (the skip-gram algorithm is basically a one matrix multiplication followed by softmax, there isn't even place for activation function, why is this deep learning?), and it is simple and efficient. And most of all, there is no overhead using word2vec, just a difference between pre-trained vectors or trained ones.
I don't understand what this article tries to say.
Once you have that, we can talk about actually replacing word2vec and similar solutions.
You can replace pre-trained word2vec in 12 languages (with aligned vectors!) with ConceptNet Numberbatch [1]. You can be sure it's better because of the SemEval 2017 results where it came out on top in 4 out of 5 languages and 15 out of 15 aligned language pairs [2]. (You will not find word2vec in this evaluation because it would have done poorly.)
If you want to bring your own corpus, at least update your training method to something like fastText [3], though I recommend looking at how to improve it using ConceptNet anyway, because distributional semantics alone will not get you to the state of the art.
Also: what pre-built word2vec are you using that actually contains valid word associations in many languages? Something trained on just Wikipedia? Al-Rfou's Polyglot? Have you ever actually tested it?
[1] https://github.com/commonsense/conceptnet-numberbatch
Training on just Wikipedia is not a representative model of any language, because not all text is written like an encyclopedia.
My general criticism is that it’s easy to invent a new method, or popularize an existing method that’s superior. But that’s not the hard task. Building the algorithm is pretty easy.
Getting a corpus of colloquial, professional, and slang texts in dozens of languages, and enough computational power to actually train them all, is usually the hard part for us doing independent research, or working on open source projects without major corporate sponsors. The reason we use word2vec and similar solutions is not because of the approach it uses, but just because it provides a well-working complete package.
There is also a Python function included that does the normalization. But it's pretty simple and if you're already using word2vec you're already doing it.
If you are looking for word embedding for production use, checkout fasttext, lexvec, glove, or word2vec. Don't use the approach described in this article.
This is corroborated by the link elsewhere in this thread that shows SVD enjoying lower wall clock time on the 1.9B word Wikipedia dataset [2].
[1] https://en.wikipedia.org/wiki/Lanczos_algorithm#Application_...
Re. [2], these measurements by Radim have several issues; First, word2vec is a poor implementation, CPU-usage wise, as can be seen by profiling word2vec (fastText is much better at using your CPUs). Second, even Radim states there that his SVD-based results are significantly poorer than the w2v embeddings ("the quality of both SPPMI and SPPMI-SVD models is atrocious"). Third, Radim's conclusion there is: "TL;DR: the word2vec implementation is still fine and state-of-the-art, you can continue using it :-)".
So I don't really get your points. Instead of referencing websites and blogs, lets take a deeper look at a "proponent" for count-based methods, in a peer-reviewed setting. In Goldberg et al.'s SPPMI model [1,2] they use truncated SVD. (FYI, that proposed model, SPPMI, is what got used in Radim's blog above.) So even if you wanted use SPPMI instead of the sub-optimal SVD (alone), you would first have to find a really good implementation of that, i.e., something that is competitive to fastTest. Also note that Goldberg only used 2 word-windows for SGNS in most comparisons, which makes the results for neural embeddings a bit dubious. You would typically use 5-10, and as shown in Table 5 of [2], SGNS is pretty much the winner on all cases as it "approaches" 10-word window. Next, I would only trust Hill's SimLex as proper evaluation targets for word similarity - simply look at the raw data of the various evaluation datasets yourself and read Hill's explanations why he created SimLex, and I am sure you will agree. "Coincidentally", it also is - by a huge margin - the most difficult dataset to get right (i.e., all approaches perform worst on SimLex). However, SGNS nearly always outperforms SVD/SSPMI on precisely that set. Finally, even Omar et al. had to conclude: "Applying the traditional count-based methods to this setting [=large-scale corpora] proved technically challenging, as they consumed too much memory to be efficiently manipulated." So even if they "wanted" to conclude that SVD is just as good as neural embeddings, their own results (Table 5) and this statment lead us to a clearly different conclusion: If you use enough window size, you are better off with neural embeddings, particularly for large corpora. And this work only compares W2V & GloVe to SVD & SPPMI, while fastText in turn works a lot better than "vanilla" SGNS and GloVe. What I do agree with is that properly tuning neural embeddings is a bit of a black art, much like anything with the "neural" tag on it...
QED; This article is horseradish. Neural embeddings work significantly better than SVD, and SVD is significantly harder to scale to large corpora. Even if you use SPPMI or other tricks.
[1] https://papers.nips.cc/paper/5477-neural-word-embedding-as-i... [2] https://www.transacl.org/ojs/index.php/tacl/article/view/570
ADDENDUM: To which I should add, to avoid more discussions, that parallel methods on dense matrices exist that essentially use a prefix sum approach and double the work, but thereby decrease the absolute running time [2]. However, as that exploits parallelism and requires dense matrices, that does not apply to this discussion.
[1] https://link.springer.com/chapter/10.1007%2F978-1-4615-1733-...
In the best case, as determined by Halko et al., you low-rank k approximation of a n times m term-document matrix is O(nmk), and randomized approximations get that down to O(nm log(k)) [1]. And, according to Rehurek's own investigations [2], those approximated eigenvectors are typically good enough. I.e., in both cases, the decomposition scales with the product of documents and words, not their sum. Therefore, this is clearly not a linear problem.
On top of that, when these inverted indices grow too large to be computed on a single machine, earlier methods required k passes over the data. These newer approaches [1,2] can make do with a single pass, meaning that the thing that indeed scales linearly here is the performance gains of scaling your SVD among a cluster with these newer approaches. Maybe this is the source of confusion for some commenters here.
[1] https://authors.library.caltech.edu/27187/
[2] https://link.springer.com/chapter/10.1007%2F978-3-642-20161-...
[1] E.g., for each word in a document or a search string, it would generate not just its base form, but also a list of top 3 base forms that are different, but similar in meaning to this word's base form (where the meaning is inferred based on context).
In general, if you want to search over millions of documents, use Annoy from Spotify. It can index millions of vectors (document vectors for this application) and find similar documents in logarithmic time, so you can search in large tables by fuzzy meaning.
What's nice with counting methods is that you can simply add matrices from different collections of documents.
https://rare-technologies.com/making-sense-of-word2vec/
(also contains benchmark experiments with concrete numbers and Github code -- author here)
log( P(x|y) / ( P(x)P(y) ) )
Because the skipgram probabilities are sparse, P(x|y) is often going to be zero, so taking the log yields negative infinity. The result is a dense PMI matrix filled (mostly) with -Inf.
Should we be adding 1 before taking the log?
Do I need to reread some papers?
except this tweet which seems to be from a conference https://twitter.com/ic/status/756918600846356480?lang=en
does anyone have any info about it?
Word2Vec is based on an approach from Lawrence Berkeley National Lab posted in Bag of Words Meets Bags of Popcorn 3 years ago 2 "Google silently did something revolutionary on Thursday. It open sourced a tool called word2vec, prepackaged deep-learning software designed to understand the relationships between words with no human guidance. Just input a textual data set and let underlying predictive models get to work learning."
“This is a really, really, really big deal,” said Jeremy Howard, president and chief scientist of data-science competition platform Kaggle. “… It’s going to enable whole new classes of products that have never existed before.” https://gigaom.com/2013/08/16/were-on-the-cusp-of-deep-learn...
Spotify seems to be using it now: http://www.slideshare.net/AndySloane/machine-learning-spotif... pg 34
But here's the interesting part:
Lawrence Berkeley National Lab was working on an approach more detailed than word2vec (in terms of how the vectors are structured) since 2005 after reading the bottom of their patent: http://www.google.com/patents/US7987191 The Berkeley Lab method also seems much more exhaustive by using a fibonacci based distance decay for proximity between words such that vectors contain up to thousands of scored and ranked feature attributes beyond the bag-of-words approach. They also use filters to control context of the output. It was also made part of search/knowledge discovery tech that won the 2008 R&D100 award http://newscenter.lbl.gov/news-releases/2008/07/09/berkeley-... & http://www2.lbl.gov/Science-Articles/Archive/sabl/2005/March...
A search company that competed with Google called "seeqpod" was spun out of Berkeley Lab using the tech but was then sued for billions by Steve Jobs https://medium.com/startup-study-group/steve-jobs-made-warne... and a few media companies http://goo.gl/dzwpFq
We might combine these approaches as there seems to be something fairly important happening here in this area. Recommendations and sentiment analysis seem to be driving the bottom lines of companies today including Amazon, Google, Nefflix, Apple et al.
https://www.kaggle.com/c/word2vec-nlp-tutorial/discussion/12...