Show HN: A search engine that lets you 'add' words as vectors
insightdatascience.com
insightdatascience.com
daughter + male - female -> { The Eldest (book), Songwriter, Granddaughter }
(Hopeless; should have had "son" in there)
pc - microsoft + apple -> { Olynssis The Silver Color (Japanese book), Burger Time (arcade game), Phantasy Star (series of games) }
(Hopeless; should have had "Mac" in there)
violin - string + woodwind -> { clarinet, oboe, flute }
(OK)
mccartney - beatles + stones -> { Rolling Stone (magazine), carvedilol (pharmaceutical), stone (geological term) }
(Poor; should have had Jagger or Richards or something of the kind in the top few results)
sofa - large + small -> { relaxing, asleep, cupboard }
(Poor; I'd have hoped for "armchair" or something of the kind)
So some of your examples, when clarified, are a bit clearer: Paul McCartney - Beatles + Rolling Stones = Mick Jagger is in the 3rd spot. (http://www.thisplusthat.me/search/Paul%20McCartney%20-%20Bea...) Change stone -> Rolling Stones.
Thank you for the comment!
The idea being that the words used to search on your website are likely to be conceptually similar (one doesn't add apples to oranges), so you can assume that "queen (sovereign) + man (noun) - woman" makes more sense that "queen (band) + man (unix command) - woman" because "man (noun)" and "woman" are likely to belong to similar clusters and not "man (unix command)".
I'm sold. This is really cool! (Though it's worth noting that a Google search with the same terms returns the exact same result...)
fluid dynamics + electromagnetism : expected magnetrohydrodynamics, got Maxwell’s equations and classical mechanics (not useful);
verse + 5 - rhyme : expected blank iambic pentameter, Shakespeare, etc.: got nonsense;
writer + American + Russian + Great - Nobel Prize : expected Nabokov, got Meirkhaim Gavrielov + 1 nonsense result;
plant + illegal - addictive : expected cannabis, chronic, etc; got “Plants” (thanks) and “Nuclear Weapon” (?!? ) and some Hungarian village.
EDIT: I thought maybe I wasn't being sufficiently imaginative, so I tried "Nixon + Clinton - JFK" and got nothing that looked interesting. Then I noticed that the "Nixon" part of my query was "disambiguated" to something like "non_film", and the word "Nixon" was just stripped out. I think this thing is just broken.
Some queries I tried:
medicine - science + superstition
aquarium - wet + dry
mario - nintendo + sega
luigi - green + red
pocahontas - native americans + aliens
What is the dimensionality of each word vector and what does a words position in this space "mean"? What is this dimensionality determined by? Have your tried any dimensionality reduction algorithms like PCA or Isomap? It would be interesting to find the word vectors that contain the most variation across all of wikipedia. Have you tried any other nearest neighbor search methods other than a simple dot product, such as locality sensitive hashing?
I guess most of those questions are about the word2vec algorithm, but you are probably in a good place to answer them after working with it. Anyways, cool work, and I am glad you did it in python so I can really dig in and understand it.
Word2vec can be viewed as dimensionality reduction. In some sense, the dimensionality of words is the same as the vocabulary size, if a one-hot encoding is used (which is most common). For word2vec specifically, the output vector dimension is a parameter specified during training, so you can choose it however you want. I suspect the OP chose something fairly large (>256) because of the performance boost.
Nearest neighbor methods work well with word2vec words: https://code.google.com/p/word2vec/#Word_clustering
The paper associated with word2vec is fairly easy to read and understand, even if you don't have a background in NLP or neural networks: http://arxiv.org/pdf/1301.3781.pdf
>> What is the dimensionality of each word vector and what does a words position in this space "mean"? What is this dimensionality determined by?
Each dimension roughly is a new way that words can be similar or dissimilar. So I've got 1000-dimensional vectors, so words can be similar or dissimilar in only one thousand 'ways'. So associations like 'luxury', 'thoughtful', 'person', 'place' or 'object' are learned (roughly speaking). Of course, real words are far more diverse, so this is an approximation. The 1000 dimensions is configurable, and in theory more dimensions means more contrast is captured, but you need more training data. In practice, the number 1000 is chosen because that maxes out the size of my large memory machine. That said, the word2vec paper shows good results with 1000D, so it doesn't seem to be a bad choice.
>> Have your tried any dimensionality reduction algorithms like PCA or Isomap?
Yes! I've tried out PCA, and some spectral biclustering using the off-the-shelf algorithms in SciKits Learn. I only played around with this for an hour or so but got discouraging results. Nevertheless, the word2vec papers actually show that this works really well for projecting France, USA, Paris, DC, London, etc. on a two-dimensional plane where the axis roughly correspond to countried & capitals -- exactly what you'd hope for! I wasn't able to replicate that, but Tomas Mikolov was!
>> It would be interesting to find the word vectors that contain the most variation across all of wikipedia.
Hmm, interesting indeed! I'm not sure how I'd got about measuring 'variation' -- would this amount to isolating word clusters and finding the most dense ones? Something like finding a cluster with a hundred variations of the word 'snow' (if you're Inuit)? I'd be willing to part with the raw vector database if there's interest.
>> Have you tried any other nearest neighbor search methods other than a simple dot product, such as locality sensitive hashing?
Only a little bit, although I'm very interested in finding a faster approach than finding the whole damn dot product (see: https://news.ycombinator.com/item?id=6720359). I worry that traditional location sensitive hashes, kd-trees, and the like work well for 3D locations, but miserably for 1000D data like I have here.
I should reiterate out that most of the hard work revolves around the word2vec algorithm which I used but didn't write. It's awesome, check it and the papers out here: https://code.google.com/p/word2vec/
Whoa, that was alot. Thanks!
https://raw.github.com/dhammack/Word2VecExample/master/visua...
https://raw.github.com/dhammack/Word2VecExample/master/visua...
https://raw.github.com/dhammack/Word2VecExample/master/visua...
And more in https://github.com/dhammack/Word2VecExample/tree/master/visu...
Maybe something like this could help: http://radimrehurek.com/gensim/similarities/docsim.html
Cool stuff, thanks.
Kinda makes sense.
It's really fast for this kind of stuff. Happy to give details about how to use it if you need.
Edit. In case you're interested in the source: https://github.com/cemoody/wizlang
Since MapReduce is used, perhaps the model is already being trained on small batches making incremental updates possible.
So there's some exciting work to be done in parallelizing and streaming the word2vec algorithm!
nice work!
http://www.technologyreview.com/view/519581
For example, the operation ‘king’ – ‘man’ + ‘woman’ results in a vector that is similar to ‘queen’.
http://www.thisplusthat.me/search/Saturn%20-%20Rings%20%2B%2...
http://www.thisplusthat.me/search/Chrome%20%2B%20open%20sour...
http://www.thisplusthat.me/search/Unix%20%2B%20Open%20Source
And it sounds really close to what I was trying at Elucidate.
I was hoping for the DS or Gameboy but expecting at least something handheld.
Result:
1. Gameboy Advance
2. Nintendo DS
3. Nintendo Gamecube
Slavoj Žižek - Jacques Lacan - Hegel
which yielded an internal server error, probably due to the diacritics not being encoded properly.Ouch, that's cold
Seems like this is how Numenta's AI works: http://www.youtube.com/watch?v=iNMbsvK8Q8Y
I think it should be Waterloo.
The results were... Interesting.
http://www.thisplusthat.me/search/Dick%20Cheney%20-evil%20%2...
On a lighter note I tried "sarah palin + sexy" and I got John Mccain, Hillary Clinton and Mitt Romney.
sleep - sleep
:)
superman - male + female:
- Lex Luthor (hmm..)
- Superman's pal Jimmy Olsen (haha, what?)
- Wonder Woman (That'll do it!)Just kidding! :)
You could also say...
ThisPlusThat.me - another rant + something cool
Thanks for posting this, very interesting work!
iPad - cool -> Windows Phone
Expected: HN, Got: Digg