Text Matching Using Cosine Similarity
kanoki.org
kanoki.org
Right now on the front page there is another article about "deep learning" classifiers which are competitive with this technology and it would be nice to see an objective comparison. (e.g. a co-worker made a 0.93 accurate text classifier in 2004 w/ bag o' words and the SVM)
Nice.
That annoying quip aside, as with many things in data processing, it's case to case. In TF-IDF you lose ordering information by definition. This is probably fine for this use case but it does mean that if ordering does matter since a set of stores share the same words.in different ordering, this will fail to resolve the difference. The author says he did due diligence on the data but there are other ways this can fall short. For example ["Walmart", "5280"] compared to ["Store", "5280"] is going to not be so similar as one would want due to the down-weighting of the identifying number in TF-IDF. So imo the disadvantage mentioned for using BoW over TF-IDF is actually not a disadvantage sometimes. As with everything it depends on your problem and data.
To the author I would hope in the future you remove statements like "to a novice, it seems easy to use X". There is nothing novice about going into a problem with an idea and trying it if it seems to fit the use case.
https://www.johndcook.com/blog/2010/06/17/covariance-and-law...
Basically, everything boils down to high-school trigonometry, because it's the most natural way to define distances in our world. I still marvel at it.
Thanks for this insight! Can you also suggest a book or other materials that I can read to understand this in greater depth?
perhaps the most important continuous distribution family because of the Central Limit Theorem
instead of
perhaps the most continuous distribution family because of the Central Limit Theorem
As for a single book, not really. See my response to the sibling comment for some more insight.
You are right that the normal distribution is not continuous because of the CLT. That was a typo. I meant "most important continuous" distribution.
Yes, other exponential distributions are convenient to work with, but the Gaussian maximum-likelihood estimate has a closed-form expression defined with just linear algebra. That's really damn nice in my opinion.
The point is that distance (in most real-world applications) is equivalent to the length of the hypotenuse of a right triangle in Euclidean space. So in my opinion, yes, high-school trigonometry governs a lot of advanced mathematics, including in statistics and machine learning.
Call it numerology if you want, I'm just trying to get people excited about math here.
- Lacks a snippet of the data in question
- Lacks proper notation ("Vector(A) = [5, 0, 2]"? never seen anything like this)
- Has grammar and formatting mistakes everywhere, even failing to properly copy-paste Wikipedia's end quotes
- Uses a batteries-included sklearn implementation which doesn't go into any interesting details
This is just stupid, people writing about stuff that has been there for some time as if it is new. I Can't believe this link is there on the front page of HN. Sorry if it seemed harsh, but somebody had to say it.
Whether or not this specific article should be on the front page of HN is a different question.
I can find plenty of blog posts which do a better job at all these criteria with a simple google search.
I don't see any value added with this post.
Maybe the author wrote it to teach whatever he has learnt, but it's not worthy of HN front page.
You could have added value just as easily by saying something like "this article from 2013 does a good job of explaining the same topic: ..."
But you chose to be rude.