Making Sense of Everything with words2map
blog.yhat.com
blog.yhat.com
The NNSE paper has associated code already, but I found setting the sparseness preference parameter was very hit-and-miss, which is why I preferred the explicit sparse-by-percentage measure in my work.
One infers from a quick read ~"Algorithms are now like people, and can learn about anything." But careful parsing of the commas shows that the sentence is true, but in the precise sense that "People can learn about anything. Now, algorithms can also learn about anything." - and the extent of learning/understanding is not being compared.
Perhaps I'm nit-picking, but this statement appears to have been constructed to support an AI pitch, and is literally true, but no 'actual AI' is involved (and no-one is actually claiming it is... unless you /want to believe/).
These techniques are impressive, and yhat is demonstrating that they are very capable. It's just that I feel a little sad that the 'AI pitch' is being turned on, when the 'really good tech' is a much more valid way to understand what they're doing.
i.e. it's nice for visualisation, and it removes noise. It would be interesting to see discussion from y-hat about where the sweet spot between lack of noise and still keeping relevant information is. I think because the subject matter is pretty simply to cluster, 2D works well enough and keeps everything simple.
word2vec -> t-SNE(2D) -> HDBSCAN
by
word2vec -> t-SNE(10D) -> HDBSCAN -> t-SNE(2D) (or just 2D projection that separates well the centroids)
?
----
About clustering in 300D, specifically for word2vec vectors, to which cosine similarity is important: Random-Projections LSH is supposed to "be compatible" with the cosine similarty. I have been toying with it and getting sort of good recalls (70%) with 100D.
It might help a lot with regards to finding the nearest neighbours. Or were you meaning that the clustering itself will be meaningless, regardless of our ability to find neighbours?
To my mind this isn't so bad -- the notion of overlapping clusters is something I think most people actually accept (think of clusters as tags rather than a partitioning and you can get the idea). However I can see why you may prefer to have your clusters present more clearly in 2D.
There's often no need to cluster in a higher space because the outcome is the same, and clustering is more expensive when there's more dimensions.
Some things that come to mind:
I'd be interested to see other vector operations such as projection of one word into another in the examples. Also, only nouns yet.
How is ≈ defined, if the distance to the closest word vector is not necessarily unique?
Finally, what is the proportion of words that maintain human meaning when averaged to those that are nonsense? What are the most "meaningful" words, in that sense?
https://lvdmaaten.github.io/tsne/
anyone looking for an explanation of word2vec may find this helpful:
Is it ok to reference your document for our papers? MIT licence is awesome and let us reuse your tech. Our site is at www.shoten.xyz if you are interested to know what we are doing
Then, second step, I augment the rank (increase or decrease) subset of top results with predetermined queries that match these results.
For example among two equal documents & search query "Who is Tommen" if the second document has more people clicking then I increase the pagerank for second document (by a function of how many more people prefer the second document)
electricity + silicon ≈ solar cells
virtual reality + reality ≈ augmented reality
--
These always seem impressive in word vector models, but in reality, I imagine that "robot" and "cyborg" were already pretty close. The fact that adding "human" nudged the vector closer is likely not as meaningful as it would be nice to believe. The same for "electricity/solar cells" and "virtual reality/augmented reality"
Still a really nice application for word2vec, and I'm looking forward to seeing other similarly practical implementations in future.
K-means is clustering and similar to this,correct.