Understanding word vectors
gist.github.com
gist.github.com
For my problem (1000+ total classes, 1 class per input), I experimented with Naive Bayes + TFIDF (~50% accuracy, < 1 sec training), then Word2vec + CNN model on GPU (~70% accuracy, 6 hrs training), and finally FastText (99% accuracy, 10 minutes training).
FastText [0] in particular is quite impressive. It is essentially a variant of Word2Vec that also supports n-grams, and there is a default implementation in C++ with a built-in classifier that runs on the command-line (no need to setup Tensorflow or anything like that).
Despite it only running on plain CPUs and only supporting a linear classifier, it seems to beat GPU-trained Word2Vec CNN models in both accuracy and speed in my use cases. I later discovered this paper from the authors comparing CNNs (and other algorithms) to FastText, and their results track my experiences [1].
This goes to show that while GPU-accelerated models are cool, sometimes using a simpler, more suitable model can have a significantly better pay-off.
Other approaches like doc2vec are also used.
"Often the w2v embedding layer will be the first layer of a network, through which a document of word representations will be passed. The output of the embedding layer will be a 2 dimensional tensor with the embedding dimension in one direction and the number of input words in the other. The next step is often to apply zero or more convolutional units over the word direction before some kind of pooling to get a 1 dimensional output. The output of the pooling layer will be a document representation based upon the word vectors. Other approaches like doc2vec are also used."
I thought the approach you described was doc2vec. If not, then does it have a name/citation?
Either way... Very nice, thanks!
[0] http://www.aclweb.org/anthology/P/P08/P08-1028
For the FastText model, it would be whatever FastText is doing under the hood. I haven't peeked at the code to see what it is doing, but Section 2 of the paper I cited above seems to imply that an average is taken.
[0] https://github.com/keras-team/keras/blob/master/examples/imd...
As with anything, your mileage may vary.
One aspect of FastText that definitely helped in my case was n-gram support (both word and character, tunable via command-line arguments). In my corpus, I have short fragments of sentences containing misspelled words, incorrect grammar etc. plus my test set has out-of-vocabulary words.
n-grams are more robust to these than Word2Vec which uses a static vocabulary.
There is no silver bullet. The best solution is to collect more data to bring the classes to a balance. The second best approach is to try algorithms like SMOTE.
To clarify: fasttext does not support n-grams of words, but instead considers n-grams of characters within words.
See options -minn, -maxn and -wordNgrams.
In the "Doing bad digital humanities with color vectors", if you consider colors as 3D vectors, which they do, you'll see that summing enough uniformly sampled vectors always gives medium browns, because that's the color in the middle of the colorspace. Instead, you should model colors in a polar space and sum vectors in that space. This will prevent going inside the sphere and losing color saturation.
It's explained quite well here in the Interpolation section: http://www.inference.vc/high-dimensional-gaussian-distributi...
If you want to understand contemporary use of words embeddings in ML, a nice simple model is explained, with full code, here: https://blog.keras.io/using-pre-trained-word-embeddings-in-a...
The original model comes from Kim 2014, https://arxiv.org/abs/1408.5882 It's a very neat use of CNNs for language processing, instead of the more popular RNNs/LSTMs. CNNs have the advantage of training much faster.
I mostly see word2vec and fasttext, neither of which are CNN nor rnn
CNNs and RNNs would be used in the next stage of the pipeline for whatever your task is (machine translation etc), probably as some way of combining the word vectors. RNNs are especially nice as they can deal with variable length sentences. Note also that it's possible for systems to learn their own systems as part of training, rather than using pre-trained ones from word2vec etc.
Word2vec and fasttext use nn that are so shallow that the can be described as types of linear regression.
Another problem of word vectors is that any word might actually have multiple senses, while vectors are just point estimates. If we wanted to be correct, we needed first to find the right sense for each word in a phrase and only then assign the vector. There is research in "on-the-fly" word vectors that adapt to context, but it's much harder to use.
A third problem of word vectors is out of vocabulary words and words with low frequency. For OOV, the usual solution is to create character or character-ngram embeddings that can be used to compute embeddings for new words. For low frequency words we usually ignore them (apply a cutoff).
Then there is the problem of phrases and collocations - some words go together, such as "New York" and "give up". The meaning of the phrase is different from the sum of the meanings of the component words. In these cases we need to have lists of phrases and replace them in the original text before training the vectors, so we have proper vectors for phrases.
By the way, one amazing tool that goes with word vectors is the library 'annoy' which can do similarity search in log time. So you can do approx 1000 lookups per second per CPU even if the database contains millions of vectors, pretty good. Annoy can be used to find similar articles, or music recommendations. Another remark - my preferred word vectors are computed with Doc2VecC (a variant of doc2vec with corruption). Doc2VecC seems more apt to discriminate between topics, but the secret is to feed it gigabytes of text.
Playing with word vectors has taught me intuitively how it is to navigate a space of high dimensionality. It feels different than 3d-space because each point has a shortcut to other points, each point leads to hundreds of other places which might be far apart. It's like a kaleidoscope where a small change can create a very different perspective.
doc2vecC: https://github.com/mchen24/iclr2017
Interesting! I wonder if you could you e.g. arbitrarily split a word into some number of symbols, e.g. two, and each time you're going to apply a training update, only apply the update to one of the symbols -- perhaps initially choosing the symbol facing the greatest loss (forcing the symbols apart in vector space), and then eventually switching over to picking the symbol with the smallest loss (letting each settle onto its own precise meaning)?
This is interesting, and there seems to be a bit of debate about it (at least with compositional distributional semantic models). [0][1] seem to show that sense disambiguation helps in some contexts, and [2] show that they don't in others. It doesn't seem immediately clear who is right here. I agree with you though that it seems pretty likely that disambiguating would be helpful.
[0] https://www.aclweb.org/anthology/W13-3513
I'm actually intrigued that nobody has made a text processor named after https://en.wikipedia.org/wiki/Amelia_Bedelia yet. :)
This seems like it should obviously make sense, but in practice it doesn’t necessarily. Usually the sense in defined by the context, and so the vector can carry the meaning of multiple contexts at once.
There are some specific tasks where this breaks for some words. I can’t remember one I’ve seen right now, but thinking about the “bark” example, it can cause a problem when both the “woodyness” and “noisyness” of a word cause completely different results. In practice that’s pretty rare.
Generally though, the non-relevant meanings don’t have any negative effects as they can be ignored.
There has been work on representing words not as vectors but as multimodal Gaussian distributions in order to try to deal with polysemy, such as [0], of which an implementation (which I have not tried to use) is available on GitHub at [1].
>It feels different than 3d-space because each point has a shortcut to other points, each point leads to hundreds of other places which might be far apart. It's like a kaleidoscope where a small change can create a very different perspective.
I appreciate that the author of the parent comment has found some sort of intuition, but I would caution others from trying to use the above quote in order to develop their own intuition as it is meaningless in any rigorous sense.
And just for getting idea why does it work, and play with examples in your browser: http://p.migdal.pl/2017/01/06/king-man-woman-queen-why.html
https://blog.keras.io/using-pre-trained-word-embeddings-in-a...
ah shit this makes me want to stand on a desk.
Allison we know this is a lot of work for a very small group but, if you see this, a couple of us here would be super stoked if you could mirror your articles somewhere else as well!
Regular Github is free and clear, so we're good there.
What's it like being a software dev in North Korea?
Almost all of the "unique" features listed there are in fact a standard part of Gensim. Fast approximative queries (using Annoy), memory mapping with lazy loading, ngrams features, format convertors, Python interface, parallelization, pre-trained models for download…
There is a way to promote cool new libraries, but this ain't it.
Shoot me an e-mail (link in HN profile), we just created this a few days ago, so it hasn't been up long, and happy to fix any disagreements in the benchmarks. You're absolutely right about Annoy Indexing, but to be fair, I don't think it was part of Gensim when I started using it :). Gensim's a great library and Magnitude's not meant to be an attack on it (in fact we use it for our own converter) and we provide zero-training which Gensim does handle.
Gensim has many warts for sure, but a benchmark that gives a green tick to itself in "Simple Pythonic interface" and a red check mark to Gensim, while copying the Gensim interface almost verbatim, was not created in good faith.
For the record, the claim was "Pythonic interface" not "Python" interface because we support some Pythonic syntactic sugar like "cat in vectors" with the "__contains__" method and "for key, vector in vectors" with the "__iter__" method. It wasn't meant to be in bad faith, but I could see how that claim could be misinterpreted, so I will remove it.
The interface is very similar to Gensim's, but Gensim is after all open source, and we made it very similar on purpose so it could be easily swapped out in our own internal codebase :).
Like I said, I think Gensim's a great library! Thanks making us aware of your concerns, I also sent you an e-mail. I'll update the repository later today to remove the comparison.
For those who like Tsne. Check out the relatively new UMAP, which seems to be faster and better.