How does cosine similarity work?
tomhazledine.com
tomhazledine.com
Instead just try to think about what it is: the sum of term-by-term products of normalized vectors. A product is the soft version of a logic AND, and it makes intuitive sense that vectors A and B are similar if there are a lot of traits that are present in both A AND B (represented by the sum) relative to the total number of traits that A and B have (that's the normalization process).
Forget about angles and geometry unless you are comfortable with N-dimensional space with N>>3. Most people aren't.
we absolutely are doing geometry here, given we're talking about metrics in a vector space – and this is trigonometry you learned by the first year of high school.
In the end, vectors (long lists of numbers a1, a2, a3, … an) start looking more like discrete functions f(i) = ai. And you can extend the same concept all the way to continuous functions - they’re like infinite dimensional vectors. For continuous functions over a finite interval the dot product (usually called the inner product in this domain) is just the integral of the product of two functions, and the ‘magnitude’ of a function is its RMS, and that means functions have a ‘cosine similarity’ which is not remotely geometric. There isn’t any geometric sense in which there is an ‘angle between’ cos(x) and sin(x) except it turns out that they have a cosine similarity of 0 so it implies the ‘angle between’ them is 90°, which actually makes a lot of sense. But in this same sense there’s an ‘angle between’ any two functions (over an interval).
But we are not doing geometry here.
Implausible how? “geometric” doesn’t mean “embeds nicely in 3D space”.
What’s wrong with talking about the angle between two L^2 functions defined on an interval? Geometric reasoning still works? If you take a span of two functions, you have a plane. What’s the issue?
When people say “hyperplane” they are generally talking about something with more than two dimensions.
(So, when the ambient vector space is finite-dimensional, the dimension of a hyperplane is one less than the dimension of the ambient vector space.)
No. They can be expressed as lists of numbers in a basis if the vector space is equipped with a scalar product but the vector itself is an object that transcends the specific numbers it is expressed in.
What you’re saying here is totally wrong and I recommend you check out the Wikipedia page on vector spaces. The geometrical object “a vector” is the more fundamental thing than the list of numbers
Vectors are not purely geometric objects. Geometry is a lens through which we can interpret vectors. So is linear algebra. The objects behave the same and both perspectives give us insights about them.
Insisting vectors are only geometric is like saying complex numbers are geometric because they can be thought of as points on the complex plane.
It obeys the normal rules you learned in geometry. For example, pick three functions a,b,c. The functions form a triangle. The triangle obeys the triangle inequality—the distances satisfy d(a,b) ≤ d(a,c) + d(c,b). The angles of the triangle sum to 180°.
This sounds an awful lot like geometry to me.
I suspect you’re using American terminology. When talking about school years it’s often useful to talk about the year or grade of school, like “9th grade” or “year 9” as it’s more universal.
I would expect most people to know about trigonometric functions by age 12, yes. (I entered high school at 11 and the first topic tackled in maths classes was elementary trigonometry.)
So I tend to sympathize with the by ear folks.
I can intuit a lot of things about music and even visualize some of it, but eventually I hit limitations. What I learn through these intuitions still applies as my ability to mentally visualize or model the music begins to fail, though.
It's similar with vectors. Once you have the orchestral equivalent of vectors, there's no way I'm visualizing it and doing mental geometry. However, what I learned and the modelling I developed from the "casio keyboard playing jingles" equivalent of vectors is still useful and applicable.
I guess this is the point where playing by ear or mentally modelling things fails, and notation is far more helpful. Yet if a lot of us approach these complex works from the notation angle first, we might feel pretty lost and uncertain about what we're doing with it and why.
I can tell I'm not articulating this well, but I like the musical analogy and wanted to get that out.
Richard Hamming has a whole section lecture to make everyone realize precisely this [1]. This was an eye opener to me.
I took many classes in school where we worked with higher dimensional spaces. You wouldn’t send a physics major a lecture on physics, say it was “eye opening”, and expect them to feel the same way about it. It is stuff they have already seen before. Maybe their eyes are already open.
To be honest it’s kind of rude.
The whole point of measuring similarity this way is that any two vectors exist in a two-dimensional space, which is where you measure the angle between them. Why would you need to be comfortable with high-dimensional spaces?
(A) Look at this space. Every point within it can be reached by combining these two vectors.
(B) Look at this space. No point outside it can be reached by combining these two vectors.
Saying that two vectors span a space is claim (A). Saying that the space they span contains them is... much weaker than claim (B), but it's related to claim (B) and not to claim (A).
I.e. simply a quantitive difference.
This article isn't talking geometry, it's trigonometry. And half the article is visual anyway.
If those points are really close together, then angle between the two vector lines is very small. Loosely speaking cosine is a way to quantize how close two lines with a shared origin is. If both lines are the same, the angle between them is 0, and the cosine of 0 is 1. If two lines are 90 degrees apart, their cosine is 0. If two lines are 180 degrees apart, their cosine is -1. So cosine is a way to quantify the closeness of two lines which share to same origin
To go back with 2 points in space that we started with, we can measure how close those 2 points are by taking the cosine of the lines going from origin to the two points. If they are close, the angle between them is small. If they are the exact same point, the angle between the lines is 0. That line is called a vector
Cosine similarity measures how closes two vectors are in Euclidean space. That’s we end up using it a lot. It’s no the only way to measure closeness. There are many others
I have no idea if this a common optimization or if it was something very niche. It was for a heuristic matrix reordering strategy, so I think they were willing to accept some mistakes.
For most commercial embeddings (openai etc) this is not a problem as the embeddings are already normalized
In the intuitive explanation I gave, the distance from the origin doesn’t matter at all, unless you are trying to force “cosine disbtance” to mean “metric distance”
There are many many ways to quantify closeness. Metric distance i.e taking a ruler and measuring the distance between the 2 points, is one way.
Measuring the angles between the 2 lines that join the points to the origin is another way to measure closeness
The squared of the measured metric distance is another way
The absolute value of the ruler distance is another way
The question you asked is only relevant if you’re trying to force cosine distance into ruler distance. You don’t need to. Cosine distance is a sufficient way by itself to measure closeness.
But yeah generally speaking, costume distance correlates better with metric distance even the points are relatively more equidistant to the origin
For example, if I have a set of words and I want to consider their relative location on an axis between two anchor words (e.g. "good" and "evil"), it makes sense to me to project all the words onto the vector from "good" to "evil." Would comparing each word's "good" and "evil" cosine similarity be equivalent, or even preferable? (I know there are questions about the interpretability of this kind of geometry.)
The J-L lemma is at least somewhat related, even though it doesn't to my understanding quite describe the same transformation.
https://en.m.wikipedia.org/wiki/Johnson%E2%80%93Lindenstraus...
Imagine you have a simple bag-of-words model of a document, where you just count the number of occurrences of each word in the document. Numerically, this is represented as a vector where each dimension is one token (so, you might have one number for the word "number", another for "cosine", another for "the", and so on), and the magnitude of that component is the count of the number of times it occurs. Intuitively, cosine similarity is a measure of how frequently the same word appears in both documents. Words that appear in both documents get multiplied together, but words that are only in one get multiplied by zero and drop out of the cosine sum. So because "cosine", "number", and "vector" appear frequently in my post, it will appear similar to other documents about math. Because "words" and "documents" appear frequently, it will appear similar to other documents about metalanguage or information retrieval.
And intuitively, the reason the magnitude doesn't matter is that those counts will be much higher in longer documents, but the length of the document doesn't say much about what the document is about. The reason you take the cosine (which has a denominator of magnitude-squared) is a form of length normalization, so that you can get sensible results without biasing toward shorter or longer documents.
Most machine-learned embeddings are similar. The components of the vector are features that your ML model has determined are important. If the product of the same dimension of two items is large, it indicates that they are similar in that dimension. If it's zero, it indicates that that feature is not particularly representative of the item. Embeddings are often normalized, and for normalized vectors the fact that magnitude drops out doesn't really matter. But it doesn't hurt either: the magnitude will be one, so magnitude^2 is also 1 and you just take the pair-wise product of the vectors.
To be a bit more explicit (of my intuition). The vector is encoding a ratio, isn't it? You want to treat 3:2, 6:4, 12:8, ... as equivalent in this case; normalization does exactly that.
When I dabbled with latent semantic indexing[1], using cosine similarity made sense as the dimensions of the input vectors were words, for example a 1 if a word was present or 0 if not. So one would expect vectors that point in a similar direction to be related.
I haven't studied LLM embedding layers in depth, so yeah been wondering about using certain norms[2] instead to determine if two embeddings are similar. Does it depends on the embedding layer for example?
Should be noted it's been many years since I learned linear algebra, so getting somewhat rusty.
Consider [1,0] and [x,x] Normalised we get [1,0] and [sqrt(.5),sqrt(.5)] — clearly something has changed because the first vector is now larger in dimension zero than the second, despite starting off as an arbitrary value, x, which could have been smaller than 1. As such we have lost information about x’s magnitude which we cannot recover from just the normalized vector.
Pray tell, which dimension do we lose when we normalize, say a 2D vector?
>>> Magnitude is not a dimension [...] To prove this normalize any vector and then try to de-normalize it again.
Say you have the vector (18, -5) in a normal Euclidean x, y plane.
Now project that vector onto the y-axis.
Now try to un-project it again.
What do you think you just proved?
Before normalization, the vector lies in R^n, which is an n-dimensional manifold.
After normalization, the vector lies in the unit sphere in R^n, which is an (n-1)-dimensional manifold.
But the intuitive notion is that if you take all 3D and flatten it / expand it to be just the surface of the 3D sphere, then paste yourself onto it Flatland style, it's not the same as if you were to Flatland yourself into the 2D plane. The obvious thing is that triangles won't sum to 180, but also parallel lines will intersect, and all sorts of differing strange things will happen.
I mean, it might still work in practice, but it's obviously different from some method of dimensionality reduction because you're changing the curvature of the space.
> triangles won't sum to 180
Exactly. Spherical triangles have the sum of their interior angles exceed 180 degrees.
> parallel lines will intersect
Yes because parallel "lines" are really great circles on the sphere.
Yes, I guess it’s curious that the information lost doesn’t seem very significant (this also matches my experience!)
There's also some generalizations to higher dimensional notions of cosines that are kind of interesting.
https://www.falstad.com/dotproduct/
https://wordsandbuttons.online/interactive_mnemonics_for_dot...
In essence then, not as confusing to the beginner who might even know what a dot product 'is' operationally but not what it 'does'.
So level 1 'it's just a normalized dot product', level 2 more immediately intuitive: 'is arrow 1 pointing in the same direction as arrow 2?' or 'how close is arrow 1's direction to arrow 2's direction?'
Now what's left after that is 'Why is it so? Why did we decide on this in embeddings?'
Cosine similarity considers only the angle between vectors, not their magnitude. This is problematic when the magnitude carries important information. For example, if vectors represent term frequencies in documents, cosine similarity treats two documents with vastly different lengths but the same proportion of words as identical. Sensitive to High-dimensional Sparsity:
In high-dimensional spaces (e.g., text data), vectors are often sparse (many zeros). Cosine similarity might not provide meaningful results if most dimensions are zero since the similarity could be dominated by a few non-zero entries. No Sense of Absolute Position:
Cosine similarity measures the angle between vectors but ignores their absolute position. For example, if vectors represent geographical coordinates, cosine similarity won't capture differences in distances properly. Poor Performance with Highly Noisy Data:
If the data has significant noise, cosine similarity can be unreliable. The angle between noisy vectors might not reflect true similarity, especially in high-dimensional spaces. Does Not Handle Negative Values Well:
If vectors contain negative values (e.g., sentiment scores, certain word embeddings), cosine similarity may yield unintuitive results since negative values can affect the angle differently compared to positive-only data. Assumes Non-Negative Values:
Often, cosine similarity assumes non-negative values. In contexts where vectors have both positive and negative values (e.g., sentiment analysis with positive and negative sentiment words), this assumption can lead to misleading results. Not Ideal for Measuring Dissimilarity:
Cosine similarity can be unintuitive when measuring dissimilarity. Two vectors that are orthogonal (90 degrees apart) will have a similarity score of 0, but vectors pointing in opposite directions (-1 cosine similarity) might need a different interpretation depending on the context. Inappropriate Use Cases Data with Magnitude Importance:
When the magnitude of vectors is crucial (e.g., comparing sales data, where larger magnitudes indicate higher sales), using cosine similarity would ignore valuable information. Time Series Analysis:
For time-series data, the order and distance of data points matter. Cosine similarity does not account for these aspects and may not provide meaningful comparisons for temporal data. Geospatial Data:
When working with geospatial coordinates (latitude, longitude), cosine similarity does not account for Earth’s curvature or distance metrics like the Haversine formula. Data Representing Complex Structures:
For data representing graphs, trees, or other complex structures where connectivity or sequence matters, cosine similarity may not capture the intricate relationships between nodes or elements. Vectors with Negative Components:
In cases where vectors have meaningful negative components (like certain word embeddings or feature vectors in machine learning models), cosine similarity can yield misleading similarity scores. Suggestions for Alternatives Euclidean Distance: When absolute magnitude is important, or when interpreting actual distances between points. Jaccard Similarity: For binary or set-based data, where overlap or presence/absence matters. Pearson Correlation: For datasets where linear relationships are of interest, especially with normally distributed values. Hamming Distance: For comparing binary data, especially for bit strings or categorical attributes. Manhattan Distance (L1 Norm): For high-dimensional data where you want to measure the absolute difference across dimensions. Cosine similarity is effective for certain applications, such as text similarity, but its limitations make it unsuitable for other contexts where magnitude, distance, or data distribution play a critical role.
Having said this, nowadays the metric I like the most is SAD (sum of absolute difference) which is Sum(abs(x2i - x1i)), which is the L1-norm of the difference image. I find this oddly easy/simple to reason and implement, so I use it in any model where it works.
Most similarity metrics will be very low if vectors don't even point in the same direction, so cosine similarity is a cheap way to filter out the vast majority of the data set.
It's been a while since I've studied this stuff, so I might be off target.
"The important part of an embedding is its direction, not its length. If two embeddings are pointing in the same direction, then according to the model they represent the same "meaning"."
This can't be quite right. Any LLM transformer model looks at the embedding of the token sequence, (without normalizing, i.e. including its magnitude) for deciding on the next token. Why would you throw away that information, equivalent to throwing away one embedding dimension?
If I had to guess why cosine similarity is the standard for comparing embeddings I suspect it's simply because the score is bounded in [-1, 1], which you may find more interpretable than the unbounded score obtained by the unnormalized dot product or Euclidean distance.
In my experience, choice of similarity metric doesn't affect embedding performance much, simply use the one the embedding model was trained with.
EDIT: Here's a better treatment, and it is the case that they give the exact same orderings: https://ajayp.app/posts/2020/05/relationship-between-cosine-...
For anyone new to the topic, note that the monotonic interpretation of cosine distance is opposite to that of cosine similarity.
I've outlined some of the related issues here: https://github.com/ashvardanian/SimSIMD#cosine-similarity-re...
Regarding the speed, yes, I wouldn't use it with big data. Up to a few thousand items has been fine for me, or perhaps a few hundred if pairwise.
In other cases people prefer L2 distances for embeddings, where the magnitude can have a serious impact on the distance between a pair of points.
Cosine similarity still works though, since it only look at how aligned vectors are.
The thing that people tend to overlook is, that there is no need for embeddings to be a vector space endowed with an inner product.
Words don’t have this structure, we define it on the image of the mapping from words to n-tuples and the embeddings we use coevolved in such a way that we assume the cosine similarity to be meaningful.
the notation here is bad. the bottom of the division looks like a cross product
as a games and graphics programmer i find it amazing that this would be a mystery... understanding the dot product is utterly foundational, and is some high-school level basics.
Cosine similarity works if the model has been deliberately trained with cosine similarity as the distance metric. If they were trained with Euclidean distance the results aren’t reliable.
Example: (0,1) and (0,2) have a cosine similarity of 1 but nonzero Euclidean distance.
A dot product between two complex numbers naturally encodes confidence in the result in the magnitude.
I blame numpy