Word2Vec Explained. Explaining the Intuition of Word2Vec
towardsdatascience.com
towardsdatascience.com
While their articles always look nice, their content is all written quickly by data scientists wanting to polish their resume with the ultimate aim of rapidly generating content for TDS that will match every conceivable data science related search. This post clearly exists solely so that TDS can get the top spot for "Word2vec explained" (which they have). As evidence of this tactic you can see that there already is a TDS post "Word2vec made easy" [0], offering nothing substantially different than this one.
The problem is that content is almost never useful, it just looks nice at first skim through. The authors, at no real fault of their own, are just eager novices that rarely have new perspective to add to a topic. It's not uncommon to find huge conceptual errors (or at least gaps) in the content there.
I personally encourage everyone at every level to write about what they can, but the issue is that TDS has manipulated this population of eager data scientists in order to dominant search results on nearly every single topic they can cover related to DS, which has made searching for anything tedious.
Compare this post to the fantastic work of Jay Alammar [1]. Jay's post is truly excellent, covering a lot of interesting details about word2vec and providing excellent visuals as well.
I'm assuming TDS will fold as soon as DS stops being a "hot" topic (which I think we'll be in the relatively near future), and will personally be glad to see the web rid of their low signal blog spam.
0. https://towardsdatascience.com/word2vec-made-easy-139a31a4b8... 1. https://jalammar.github.io/illustrated-word2vec/
You should check out Amazon’s MLU’s interactives - they’re like mini nyt articles on different algorithms:
It's unusual that this article got vouched.
Yes, perhaps it's a bit light-weight, but as an introduction I thought it did a good job.
https://transorthogonal-linguistics.herokuapp.com/TOL/boy/ma...
(Which can reproduce an old XKCD about the 'purity' of other scientific fields compared to math: https://twitter.com/RadimRehurek/status/638531775333949440)
Has anyone approached similar techniques on non-text corpuses?
In your proposed use case I would bet that you will “see” the kind of similarity you’re looking for based on vector similarity, but I also expect it to largely be an illusion due to confirmation bias. It will be much harder to make that similarity actionable to solve the actual business use case. (Like 30% of the time it’ll work like magic; 60% of the time it’ll be “meh”; 10% of the time it’ll be hilariously wrong.)
I tried this approach and it did improve the overall performance. The next step would be fine tuning the transformer model. I want to see if I could do it without disturbing the existing weights too much. Here's the library I used to get get the embeddings
Look into node2vec libraries for instance
https://www.kdd.org/kdd2018/accepted-papers/view/real-time-p...
The intuition behind doc2vec it's a bit harder to grasp. I understand the role of the "paragraph word": it provides context to the prediction so in "the ball hit the ---" in a basketball text the classifier would predict "rim" and in a football one "goalpost"(simplifying). But I still don't get why similar texts get similar latent representations.
A weakness of TDS monopolizing data science SEO is that it's hard for better techniques to surface.