Show HN: 3d Word Vectors with Neural Networks, T-SNE, and WebGL
wordcloud.ersatz1.com
wordcloud.ersatz1.com
I ran some visualizations of my own a while back, and they're a lot less aesthetically pleasing, but you can see more structure and clustering.
https://raw.github.com/dhammack/Word2VecExample/master/visua... https://raw.github.com/dhammack/Word2VecExample/master/visua...
(It could also just be a WebGL thing, because the third dimension seems to have a lot less variance than the other two)
I trained word2vec on wikipedia and trained a bunch of different models w/ TSNE for between 800 and 1500 iterations. In 2d, the results actually appear fine (similar to your graphics + others that have been demonstrated). Unfortunately, at least for me, this doesn't seem to be translating to 3d.
It's either a mix of the quality of the vectors (maybe I need to try some other word vectors) and/or T-SNE needs to be optimized a bit better (and/or there's a bug). It's not quite where I want it to be yet, so consider this a "cool tech demo we're still working on" kind of thing.
However, I'm wondering if there's something I'm missing that makes it not suitable for 3d--IE, maybe the assumptions being made to speed things up break down after the 2nd dimension. Also, there's interesting discussion in the literature about whether or not T-SNE is a good dimensionality reduction technique in general (as opposed to only a very powerful visualization technique), so my next step is probably going to be running the vectors through an autoencoder to generate 3d coordinates and then plotting those and comparing the visualizations.
Re: another example of TSNE w/ text--yeah, I've only seen this http://homepage.tudelft.nl/19j49/t-SNE_files/semantic_tsne.j... which seems to work but isn't interactive. Frankly, I'm surprised we got it to work with three.js--we're able to render as many as 250,000 unique words and it runs smooth (it just takes longer to download--this demo has 25,000).
Edit: it also seems to put the entire article on a single line with no punctuation. But Google's word2vec examples run with each sentence on one line. Wouldn't this make a difference in training?
Btw, nice LDA visualization!
Used t-SNE on a vector clustering which came from LDA.
These are a bunch of "word vectors" (generated w/ word2vec https://code.google.com/p/word2vec/) put through a dimensionality reduction/visualization technique called T-SNE and then plotted in 3d. As far as what the clustering represents, check out T-SNE (http://homepage.tudelft.nl/19j49/t-SNE.html), but the short answer is--it's hard to say...
Here's a longer explanation from a webinar I did yesterday where I demoed this: http://www.youtube.com/watch?v=wmlj5uTUTFY (skip to 12:11 to get to the applicable section)
"boat" is by "scarecrow", "adopted", and "feelings"
"window" is by "insulted", "prize" and "arson"
Rotating around the word-in-question didn't pop up any more similar words.
t-SNE is known for producing immediately interpretable clusterings, but this seems a bit obscure.