Is Cosine-Similarity of Embeddings Really About Similarity?
arxiv.org
arxiv.org
The other potential issue is for all the embeddings that I have seen the resulting space once you have embedded some documents is sort of "clumpy" and very sparse overall. So you have very large areas with basically nothing at all I think because semantically there are many dimensions which only make sense for subsets of concepts so you end up with big voids where really the embedding space is totally unreachable so distance doesn't have any meaning at all.
In spite of all that there are a few similarity measures which work well enough to be useful for many practical purposes and cosine similarity is one of them. I don't think anyone thinks it's perfect.
Was uniqueness ever a guarantee? It's a distance metric. It's reasonable to assume that two features can be equidistant to the ideal solution to a linear system of equations. Maybe I'm missing something.
The crossover point where the number of dimensions falls below the number of points is at 1113868. If you're willing to tolerate 10% error, it's at 7094.
Needless to say, dot products are directly supported in hardware via the FMA unit.
Add in nd indices and the costs tend to be very small.
a.b = |a||b|cos theta.
This means you get cos of the angle between the two vectors by just dividing the dot product by the product of their magnitudes. You don't actually take cos of the angle to get cosine similarity (for one because you don't know the angle) you just use "cos theta" (calculated as above) as a proxy for how narrow the angle is and therefore how close the two embeddings are.
The paper in TFA shows that if you construct an embedding space such that that angle isn't meaningfully measuring similarity then a low angle doesn't mean two things are very similar. I have a similar paper measuring bears and woods but I haven't got around to typesetting it for publication yet.
Then again, you could argue whether that is a problem when considering very high dimensional embeddings. Their conclusions seem to point in that direction but I would not agree on that.
But it doesn't differentiate further, so you can have "beautiful" and "ugly" embed very close to each other even though they are opposites - they tend to appear in similar places.
Another limitation of embeddings and cosine-similarity is that that they can't tell you "how similar" - is it equivalence or just relatedness? They make a mess of equivalent, antonymous and related things.
For modern embedding models which effectively mean-pool the last hidden state of LLMs (and therefore make use of its optimizations such as attention tricks), embeddings can be much more robust to different contexts both local and global.
The last one I have in mind is BERT and it's variants.
The not-quite-large embedding model I like to use now is nomic-embed-text-v1.5 (based on a BERT architecture), which supports a 8192 context window and MRL for reducing the dimensionality if needed: https://huggingface.co/nomic-ai/nomic-embed-text-v1.5
Bitext mining: Given a sentence in one language, find its translation in a collection of sentences in another language using the cosine similarity of embeddings.
Classification: identify the kind of text you're dealing with using logistic regression on the embeddings.
Clustering: group similar texts together using k-means clustering on the embeddings.
Pair Classification: determine whether two texts are paraphrases of each other by using a binary threshold on the cosine similarity of the embeddings.
Reranking: given a query and a list of potential results, sort relevant results ahead of irrelevant ones by sorting according to the cosine similarity of embeddings.
Etc etc.
These are MTEB benchmark tasks https://arxiv.org/pdf/2210.07316.pdf . If you have no need for something like that, good for you, you don't need to care how well embeddings work for these tasks.
All embeddings are first layer of DNN. In case of word2vec this is shallow 2-layer network. Selection of embedding is multiplication of embedding matrix by one-hot vector, which is usually optimized as array lookup.
Implementations that use a LLM require a full forward pass, but the many optimizations in inference speed make it not noticable for small applications.
We need embeddings to give relatedness across axes like synonymity etc.
Generic distance metrics can often be replaced with context-specific ones for better utility; it makes me wonder whether that insight could be useful in deep learning.
Claude Opus says "There are a few good distance metrics commonly used with latent embeddings to promote diversity and prevent model collapse:
1. Euclidean Distance (L2 Distance) 2. Cosine distance 3. Kullback-Leibler (KL) Divergence: KL divergence quantifies how much one probability distribution differs from another. It can be used to measure the difference between the distributions of latent embeddings. Minimizing KL divergence as a diversity loss would encourage the embedding distribution to be more uniform. 4. Maximum Mean Discrepancy (MMD): MMD measures the difference between two distributions by comparing their moments (mean, variance, etc.) in a reproducing kernel Hilbert space. It's useful for comparing high-dimensional distributions like those of embeddings. MMD loss promotes diversity by penalizing embeddings that are too clustered together. 5. Gaussian Annulus Loss: This loss function encourages embeddings to lie within an annulus (ring) in the latent space defined by two Gaussian distributions. It promotes uniformity in the embedding norms while allowing angular diversity. This can be effective at preventing collapse to a single point. But I haven't checked for hallucinations.
Beta diversity is one metric for examining diversity, define as the ratio between regional and local species diversity. https://en.wikipedia.org/wiki/Beta_diversity
This is a long-standing question for me. Theoretically, I should use the CS in my optimization and then also in the evaluation. But I haven't tested this empirically.
For example, there is sperical K-meams that clusters the data on the unit sphere.
It's my understanding that the delta between two word embeddings, gives a direction, and the magic is from using those directions to get to new words. The oft cited example is King-Man+Woman = Queen [1]
When did this view fall from favor?
[1] https://www.technologyreview.com/2015/09/17/166211/king-man-...
A vector is just a list of n numbers. Embedded into a n dimensional space, a vector is a distance in a direction. It isn’t ’the point you get to by going that distance in that direction from the origin of that space’. You don’t need as space to have an origin for the embedding to make sense - for ‘cosine similarity’ to make sense.
Cosine similarity is just ‘how similar is the direction these vectors point in’.
The geometric intuition of ‘angle between’ actually does a disservice here when we are talking about high dimensional vectors. We’re talking about things that are much more similar to functions than spatial vectors, and while you can readily talk about the ‘normalized dot product’ of two functions it’s much less reasonable to talk about the ‘cosine similarity’ between them - it just turns out that mathematically those are equivalent.
I think people skip over that the vectors are the result of the minimization of the objective.
That objective is roughly the same since word2vec. GLoVe is mathematically equivalent. LLMs are also equivalent.
For a LM, the objective function is still roughly the same. Maximizing probability of the next token conditional on previous tokens.
This means the embedding vector of a token minimizes distance to tokens that come before it often, and maximizes distance to those that don't.
https://github.com/tmikolov/word2vec/blob/master/questions-w...
https://github.com/tmikolov/word2vec/blob/master/questions-p...
and you could see how much better LLMs do on the same 20k examples.
The training process from the original Mikolov et al. paper only uses the analogy examples (questions-words.txt and questions-phrases.txt) to measure accuracy after training: https://github.com/tmikolov/word2vec/blob/20c129af10659f7c50...
Uncommon words have more information content than common words. So, common words having larger embedding scale is an issue here.
If you want to measure similarity you need a scale free measure. Cosine similarity (angle distance) does it without normalizing.
If you normalize your vectors, cosine similarity is the same as Euclidean distance. Normalizing your vectors also leads to information destruction, which we'd rather avoid.
There's no real hard theory why the angle between embeddings is meaningful beyond this practical knowledge to my understanding.
If you normalize your vectors, cosine similarity is the same as dot product. Euclidean distance is still different.
If all the vectors are on the unit ball, then cosine = dot product. But then the dot product is a linear transformation away from the euclidean distance:
https://math.stackexchange.com/questions/1236465/euclidean-d...
If you're using it in a machine learning model, things that are one linear transform away are more or less the same (might need more parameters/layers/etc.)
If you're using it for classical statistics uses (analytics), right, they're not equivalent and it would be good to remember this distinction.
> It's my understanding that the delta between two word embeddings, gives a direction, and the magic is from using those directions to get to new words... it's the directions and distances to nearby objects that matters most
Cosine similarity is a kind of "delta" / inverse distance between the represenation of two entities, in the case of these models.
Another thing you have to keep in mind is that these embeddings are in n dimensional space. Intuitions about the real world does not apply there
A direction can be given in terms of an angle measure, such as cosine.
In addition to the _quality_ of any proposed alternative(s), computational speed also has to be a consideration. I've run into multiple situations where you want to measure similarities on the order of millions/billions of times. Especially for realtime applications (like RAG?) speed may even out weight quality.
Ha interesting I wrote a blog post where I pointed this out a few years ago [1], and how we got around it for item-item similarity at an old job (essentially an implicit re-projection to original space as noted in section 3).
https://swarbrickjones.wordpress.com/2016/11/24/note-on-an-i...
I tried to create a Kaggle (TensorFlow Hub, TensorFlow Quantum) competition for motivating alternative formalisms but was unable to publish it because all Kaggle competitions must be evaluated with information retrieval metrics. Talk about a one-track mindset!
Today work in NLP advances by ``leaderboards'' and dubious, language-specific evaluation datasets that the same authors stand to benefit from when their proprietary model is praised for doing well on the evaluation criteria they invented a few months back. It validates the price hike for access to their proprietary models.
These formalisms that do work are at odds with Firth Mode, the preferred representation for Google (Stanford, OpenAI), so I guess we should be thankful they're still in the book. If you're interested in language, though, I'd suggest picking up a different book.
Given a function f(l, r) that measures, say, the logprobability of observing both l and r, and that the function takes the form f(l, r) = <L(l), R(r)>, i.e. the dot product between embeddings of l and r, then cosine similarity of x and y, i.e. normalized dot product of L(x) and L(y) is very closely related to the correlation of f(x, Z) and f(y, Z) when we let Z vary.
For example, here's the loss from the CLIP paper [1], which ensures cosine similarities are meaningful:
# joint multimodal embedding [n, d_e]
I_e = l2_normalize(np.dot(I_f, W_i), axis=1)
T_e = l2_normalize(np.dot(T_f, W_t), axis=1)
# scaled pairwise cosine similarities [n, n]
logits = np.dot(I_e, T_e.T) * np.exp(t)
# symmetric loss function
labels = np.arange(n)
loss_i = cross_entropy_loss(logits, labels, axis=0)
loss_t = cross_entropy_loss(logits, labels, axis=1)
loss = (loss_i + loss_t)/2
And Sentence Transformers [2] using CosineSimilarityLoss: train_loss = losses.CosineSimilarityLoss(model)
# Tune the model
model.fit(train_objectives=[(train_dataloader, train_loss)], epochs=1, warmup_steps=100)
[1] https://arxiv.org/pdf/2103.00020.pdfYou may not want to use cosine similarity as your only metric for rankings, however, and you may want to experiment with how you construct the embeddings.
There is a whole branch of mathematics called dedicated to other ways to measure distance.
If your embeddings aren't normalized it's worth trying. In our use cases it never made a substantial difference, but I imagine there are cases where it does.
E.g. the following triple
1: "Yes, this is a demonstration"
2: "Yes, this isn't a demonstration"
3: "Here is an example"
<1, 2>, Has "higher" cosine similarity than <1, 3>, structurally equivalent except for one token/word, <1, 2> semantically means the opposite of each other depending on what you're targeting in that sentence. While <1, 3> means effectively the same thing.
If this paper is about persuading people about efficacy with regards to semantic understanding, OK, but that was always known. If its about something with relation to vectors and the underlying operations, then I'll be interested.
You're probably best off using the commercial suggestion, and if its dot product, go for it. I am no expert in this area and my interest wanes every day.
Means squared error instead of dot product, it's not cheaper but it's close
If you want to go cheaper you could use sum of abs of differences.
For a lot of embeddings we have today, norm of any embedding vector is roughly of same size, so the angle between two vectors is roughly same size as length of difference that you are saying, and can be expressed in terms of 1 - dot product after scaling
I can't of course be certain in all cases, but dimensions are typically (past experience, and using knowledge from word2vec experiments from years ago) derivative of higher dimensions. The kernel still operates on the same concept by applying a norm along with whatever weightings to each dim.
Semantic understanding is still not there in my opinion, we might feign it by increasing specificity, but only so much. Largest contributor will likely still the determining factor rather than the series of smaller, more specific dimensions.
I tested this using similar sentences as my original comment and failing in more scenarios than passing. I of course am biased since it may be given I did not select the right dimensions or measures.
If you choose a straight linear mapping of tokens to a number, then you'd be right.
Extending that, if you choose any mapping which does not do a more extensive remapping from raw syntactic structure to some sort of semantic representation, you'd be right.
But hence why we increasingly use models to create embeddings instead of simpler approaches before applying a similarity metric, whether cosine similarity or other.
Put another way, there is no inherent reason why you couldn't have a model where the embeddings for 1 and 3 are identical even, and so it is meaningless to talk about the cosine similarity of your sentences without setting out your assumptions about how you will created embeddings from them.
I agree, but from generics POV, you have to settle on a few things to compare between models. If you can't, then benchmarks are useless too outside of extremely narrow measures.
I only address structure in the parent, and sure, it can be too generic of a statement by only touching on structure. But I would almost assert structure is still an important feature, and I would almost assert that it is required or otherwise a dominant feature when you want to deliver a product for general use.
I don't think I get too much more incorrect going beyond a few dimensions given this.
> Discrete entities are often embedded via a learned mapping to dense real-valued vectors in a variety of domains.
Already from that point, it is clear that a comparison based on the similarity of the textual version of the sentences is irrelevant to the evaluation in the paper. The paper consistently talk in terms of "learned embeddings" rather than simplistic direct mappings of words.