What's so special about vector embeddings
pinecone.io
pinecone.io
Shouldn't that "linear" constraint have more of a blunting effect on embeddings effectiveness? Why not use a different metric like mutual information that can account for non-linearities. Whenever correlation comes up in context of statistical comparison, there's always this warning that correlation is limited as a way of comparing data due to it's linearity so I'm confused why it seems so effective in the context of embeddings.
I found something for MI if you wanna check it out : https://www.aclweb.org/anthology/2020.acl-main.741.pdf
EDIT: As @rubatuga said,they have been avoided largely as embeddings are continuous random vars and MI is for distributions. Nonetheless I think there can be merit in exploring these
PS: I am the author and just wanted to talk about the familiar basic stuff on this one
More recent vector embeddings (e.g., BERT, ELMO) are harder to conceptualize neatly because one of the layers of the Neural Network includes a token's indexical position in the document.
Correlation (vs causation) and linearity are orthogonal concepts.