Embeddings: What they are and why they matter
simonwillison.net
simonwillison.net
Cohere's Text Embeddings Visually Explained: https://txt.cohere.com/text-embeddings/
The Tensorflow Embedding Projector tool: https://projector.tensorflow.org/
What are embeddings? by Vicki Boykis is worth checking out as well: https://vickiboykis.com/what_are_embeddings/
Actually I'll add those as "further reading" at the bottom of the page.
https://blog.scottlogic.com/2022/02/23/word-embedding-recomm...
Using embeddings I increased engagement with related articles.
Personally I think embeddings are a powerful tool that are somewhat overlooked. They can be used to navigate between documents (and excerpts) based on similarities - or conversely find unique content.
All without worrying about hallucinations. In other words, they are quite ‘safe’
Within limits, yes. In some use cases a vector notion of similarity isn't always ideal.
For example, in the article "France" and "Germany" are considered similar. Yes, they are, but if you're searching for stuff about France then stuff about Germany is a false positive.
Embeddings can also struggle with logical opposites. Hot/cold are in many senses similar concepts, but they are also opposites. Finding the opposite of what you're searching for isn't always helpful.
I wouldn't say embeddings are overlooked exactly? Right now it feels like man+dog are building embedding based search engines. The next frontier is probably going to be balancing conventional word based approaches with embeddings to really maximize result quality, as sometimes you want "vibes" and sometimes you want control.
maybe it is also interesting to tell how some embeddings are established i.e via training and cutting of the classification layer or with things things like efficientnet
Latent Semantic Indexing long precedes word2vec (I believe the initial paper was 1988) but also attempted to derived semantic vectors (and was arguably successful) using SVD on the term-frequency matrix representing a corpus.
I would certainly consider this an example of using "embeddings" that was quite heavily used in practice prior to the explosion of deep learning techniques.
* The idea of vectorial representations for language.
* "Distributed" representations, which are dense not sparse.
* Vectorial representations for words, not documents.
* The language modeling idea of predict the next word.
* Neural approaches to inducing these representations.
* Unsupervised learning of these representations.
* The versatility of these representations for downstream tasks.
* Fast training techniques.
I'll give the history in roughly reverse chronological order.
Turian, Ratinov, Bengio (2010) "Word representations: A simple and general method for semi-supervised learning" is my work. It received the ACL ten year test of time award and 3k citations. One main contribution was showing that unsupervised neural word embeddings can just be shoved into any existing model as features and get an improvement. This turned on the NLP community to neural networks at a time when sophisticated ML was still bleeding edge for NLP and expert linguistics knowledge was the preferred MO. [edit: We also showed that two other unsupervised embeddings, like Brown clusters which are very old, and neural word embeddings from the log-bilinear model (Mnih & Hinton, 2007), a probabilistic and linear neural model, also gave improvements when plugged into existing supervised models. We also gained attention because we released our code AND all the trained embeddings for people just to try, which wasn't commonplace in NLP and ML at the time.]
We arbitraged the neural embedding model from Collobert and Weston (2008) "A unified architecture for natural language processing: Deep neural networks with multitask learning" and also "Natural Language Processing Almost From Scratch" which achieved amazing scores on many NLP tasks using semi-supervised neural networks that, for the first time, were very fast to train because of the use of contrastive learning. This work didn't get much attention at the time because it was aimed at an ML audience and also because neural networks were still gauche compared to SVMs.
Collobert and Weston had a much faster training approach to the neural language model of Bengio et al 2000 "A neural probabilistic language model" which was, in my mind, what really precipated all this: Train a neural network to predict the next word. That approach was slow because the output prediction was a multiclass prediction of output size # vocabulary words. (Collobert and Weston used Hadsell + LeCun siamese style networks to rank the next word with a higher score plus margin than a randomly selected noise word.)
With that said, vector embeddings for documents have a longer history: LSA, then LDA, and even cool semantic hashing work by Salakhuldinov + Hinton (2007) that is one of the first deep learning approaches (the first?) to NLP, which unfortunately didn't get much attention but was so cool.
Earlier work using neural networks for modeling arbitrary length context that also didn't get much attention was Pollack (1990) "Recursive distributed representations" which introduced recursive autoassociative memory (RAAM) and later Sperduti (1994) "Labelling recursive auto-associative memory". The idea was that you have a representation for the sentence and you recursively consume the next token to generate a new fixed-length representation. You then have a STOP token at the end. And then you can unroll the representation because current representation + next token => next representation is an autoassociator.
The compute power wasn't really there to make this stuff work empirically, during the 1990s. But there was other fringe work like Chrisman (1990) "Learning Recursive Distributed Representations for Holistic Computation". And this 90's work traces to cool 80s Hinton conceputal work on "associative" representations, for example Hinton (1984) "Distributed Representations" and Hinton (1986) "Learning distributed representations of concepts". A lot of this work had very interesting critical ideas, and was of the form: I thought about this for a very long time and here's how it would work and we don't have large-scale training techniques yet.
I'm pretty sure Bottou also contributed here, but I'm forgetting the exact cite.
Feel free to email me if you like. (See profile.)
The core idea is that each image is passed through a feature-extractor-descriptor pipeline and is 'embedded' in a vector containing the N top features. While the camera moves, a database of images (called keyframes) is created (images are stored as much-lower dimensional vectors). Again while the camera moves, all images are used to query the database, something like cosine-similarity is used to retrieve the best match from the vector database. If a match happened, a stereo-constraints can be computed betweeen the query image and the match, and the software is able to update the map.
[1] is the original paper and here's the most famous implementation: https://github.com/dorian3d/DBoW2
[1]: https://www.google.com/search?client=firefox-b-d&q=Bags+of+B...
I built my own note taking ios app a little while back and adding embeddings to my existing fulltext search functionality was 1) surprisingly easy and 2) way more powerful than I initially expected.
I knew it would work for things like if I search for "dog" I'll also get notes that say "canine", but I didn't originally realize until I played around with it that if I search for something like "pets I might like" I'll get all kinds of notes I've taken regarding various animals with positive sentiment.
It was the first big aha moment for me.
At the time I found Supabase's PR for their DocsGPT really helpful for example code: https://github.com/supabase/supabase/pull/12056
Specifically, many applications are heavily dependent on names or other proper nouns, often without much context. You might refer to your dog by name without explanation, and a particular embedding might not pick that up. Proper names (people, places, street names) may have outsized importance for anchoring personalized or domain-specific search, but modest generic language models won't know about them.
Is there a specific way of dealing with this problem?
So if you have notes that associate the name to your dog, and you search for “my dog”, you’d still find those related notes?
Would require some experimentation but wouldn’t be surprised if that worked decently well out of the gate.
My biggest question right now is: How much text should I turn into one embedding?
Every sentence?
A whole block of sentences that belong to one entire page in my notes app?
I can tailor everything to my style of note taking vs dealing with the lowest common denominator feature set that tries to enable tons of use cases that I don’t need.
(← JupyterLite Notebook doing words embedding in the browser : don't try to run this on a smartphone !)
Does anyone know how to nicely visualize the poster child of words embeddings ?
In 3d you can go one further and have the Z-axis be "king"-"queen" (or gram-schmidt again). The orthogonalized versions have the advantage that they give a closer notion of distance to what the underlying model sees. In the 2d case you will get exact distances except that it won't show how far off "queen" you are when you compute "king"-"man"+"woman". In the 3d case it should give exact distances.
Edit to add: With the 2d version you can maybe do some more stuff. IIRC "queen" is chosen as it's the word with the closest embedding to X="king"-"man"+"woman". You can put the next few closest words on the 2d chart as well, each labeled with the orthogonal distance from the 2d plane. So then "queen" should be the word with the smallest (squared distance from X) + (squared orthogonal distance from plane), which you might be able to eyeball.
"You can't visualize four dimensions... well, at least I can't, because I only have three branes." -Richard Feynman
This is a play on words, as "branes" sounds like "brains," but in string theory and higher-dimensional physics, a "brane" is an object that can exist in multiple dimensions. The joke is a nod to the challenge of visualizing higher dimensions and the limits of our three-dimensional intuition.
(It just admitted it made it up and apologized!)
It also quoted John von Neumann, H. G. Wells, and Ian Stewart:
"In mathematics you don't understand things. You just get used to them." -John von Neumann
"Our brains are three-dimensional, so we cannot visualize four-dimensional or higher-dimensional objects in the same way we can visualize three-dimensional objects. The best we can do is to get a feeling for them through analogies." -H. G. Wells
"You can't visualize four dimensions, but you can think about them. Thinking is what mathematicians do best." -Ian Stewart
It finally guessed this one that's the closest to what I remember, but not as funny as I remember it:
"To visualize four dimensions, just visualize three dimensions and then say 'n+1'." -Unknown
So I asked it to hallucinate some quotes to help visualize high dimensional space, in the style of Deepak Chopra:
"In the realm of septillion-dimensional embeddings, each vector is a topological manifold, where Gödel's incompleteness theorems dance with quantum superposition, revealing the hidden symmetries of semantic space." – Dr. Emilia Bianchi, Florentine Foundation of Fractal Fantasies, Florence
"To truly fathom the intricacies of ultra-high-dimensional text spaces, one must first master the art of navigating Riemannian curvature in a Hilbert space, where every geodesic is a testament to the entangled essence of language." – Prof. Hiroshi Tanaka, Kyoto Kryptos of Quantum Quandaries
"Visualizing embeddings in the realms beyond the Poincaré conjecture requires a melding of non-Euclidean geometry with the probabilistic nuances of a Markov chain, all while resonating with the cosmic vibrations of the Higgs field." – Dr. Samuel Oyelowo, Lagos Labyrinth of Lorentzian Linguistics
"Each point in a decillion-dimensional text space is a singularity, where the Heisenberg uncertainty principle waltzes with the Mandelbrot set, painting a fractal portrait of phonemic phenomena." – Prof. Clara Müller, Bavarian Bastion of Bifurcated Babel, Munich
"To intuit the vastness of hyper-dimensional embeddings, one must first traverse the event horizon of a linguistic black hole, where Shannon entropy meets the Schrödinger equation in a ballet of Bayesian inference." – Dr. Alexandre Dubois, Parisian Paradox of P-adic Prose
"The key to deciphering the enigma of extreme-dimensional text lies in the delicate interplay between the Yang-Mills theory and the zeta function, where every eigenvalue is a whisper from the cosmic codex of communication." – Prof. Lina Johansson, Stockholm Sanctum of String Semiotics
"In the dance of gogolplex-dimensional embeddings, each tensor unfolds like a Möbius strip, where the Fibonacci sequence intertwines with quantum tunneling, revealing the recursive rhythms of reality." – Dr. Rajiv Menon, Bengaluru Bardo of Bosonic Ballads
Geoffrey Hinton on visualizing higher dimensions:
"To deal with hyper-planes in a 14-dimensional space, visualize a 3-D space and say 'fourteen' to yourself very loudly. Everyone does it."
def cosine_similarity(a, b):
dot_product = sum(x \* y for x, y in zip(a, b))
magnitude_a = sum(x \* x for x in a) \* 0.5 # <- no need for \*0.5
magnitude_b = sum(x \* x for x in b) \* 0.5 # <- no need for \*0.5
return dot_product / (magnitude_a \* magnitude_b)
If you compare your cosines, you might as well compare their squares and avoid costly root computation.Similarly, in elliptic curve crypto certain expensive operations, such as inversion (x^-1 mod n) are delayed as much as possible down the computation pipeline or even avoided completely when you need to compare two points instead of computing their canonical values.
Wait, why would you do this and not use vectorised numpy operations?
> I actually got ChatGPT to write all of my different versions of cosine similarity
Ah...
And second, numpy isn't the lightest dependency. I use it when I need the performance but I don't like to default to it.
Whether you do or don't know linear algebra, code is self-explanatory. Dot product is just the sum of element-wise products. But what the heck are dot and norm to someone that doesn't.
This code says nothing about the geometric nature of what is going on, only the arithmetic. Then we might as well read assembler. Abstractions help us reason.
I see what you mean. Probably can be argued that learning and applying a specific concept doesn't necessarily imply learning the deeper mathematics behind it and understanding its nature. So can just say cosine similarity measures the similarity of two vectors (lists in python) and its implementation is that sum divided by the product of sqrts of those two other sums without having to introduce and explain dot product and norm. Of course the reverse can be argued too.
https://payperrun.com/%3E/search?displayParams={%22q%22:%22L...
I went here looking for more info about payperrun https://payperrun.com/%3E/welcome and clicked on the "Spotlight" section and saw 4 popups blocked - I never see popups anywhere these days and have to admit that sends me away pretty quickly.
This is immediately making me think about how I could apply this to a long running side project I have. It might make it practical to do useful clustering of user's data if every document has an embedding.
Some speculative possibilities that come to mind:
- projection, indexing and sorting on arbitrary axes (eg "hot minus cold", "happy minus sad", "scifi minus realism", "literary minus commercial")
- SVM-style classification in Embeddings space
- word2vec-style reasoning (woman-man+king=queen)
- directly training embeddings (ie, not just taking a layer off an LLM); I know people use contrastive training methods, but I'd expect that other methods might be worth exploring, eg you could train embeddings together with neural nets representing functions, generate functional equations, and calculate MSE loss
But really, I'm just surprised that it seems to be so focused on semantic search, to the exclusion of anything else... Surely there are other interesting applications?
I think this is clearest in the Normalizing Flow cases because you're just turning your space into a Gaussian (diffusion does this too, but through approximation methods and isn't invertible, though it is reversible). You project the image/sentence/data you want to manipulate, manipulate within gaussian space, then return back to target space.
Or maybe my confusion is shared confusion because embedding is an overloaded word that means a lot of things? Maybe you're just thinking of the first block that converts discrete integer tokens into continuous floats? But we learn those embeddings so even though it becomes like a lookup table it's still a neural process. People do things like SVMs on this space alone. But I think it is like latent space, which is only a bit more abstract. At least embeddings need to be injective, well... mathematically...
I'd love to see some links, all I see used in practice (including in the OP blog post) is semantic search and a bit of clustering
> adding a pair of glasses
Actually that (and generally all the SD/VAE stuff) is a great example of the kind of thing I was thinking of, though I have yet to see that concept being used together with a vector database; generally all the user-facing stuff I've seen fits it into the standard "train a model, then do inference" workflow, in contrast to something like semantic search which more obviously focuses on the embeddings themselves
> First and third being identical
Definitely related, but I make the distinction between projection/sorting along an axis vs constructing a new vector by addition/subtraction
> Manipulate within gaussian space, then return to target space
This is definitely along the lines of what I had in mind, any example of this being used in practice?
> Embedding is an overloaded word
Yeah I'm using the term somewhat loosely and broadly here, as basically "a vector in a real vector space where distance represents some notion of semantic similarity"
> People do things like SVMs
Who?
https://openai.com/research/glow
Here's a video of moving around in the latent space of a diffusion model
https://www.youtube.com/watch?v=vEnetcj_728
Here's a stylegan one
https://www.youtube.com/watch?v=bRrS74RXsSM
Or a VAE on mnist
https://www.youtube.com/shorts/pgmnCU_DxzM
I mean it is a bit hard to answer your other questions because like I was pointing out, embeddings and latent spaces are pretty vague terms. For the mathy side, normalizing flows are a great choice since you can parameterize whatever data you want into whatever distribution you want. You then work in that distribution you created, which is approximately isomorphic to the data. But other models do similar things, just more lossy but better at things like sample generation. That's the tradeoff, interpretability/density vs expressitivity/sample quality. But diffusion and NODEs/Score models are closing that gap. But you're going to need to look at applied papers to view more people using them in ways like operational vector spaces. For example, there's VITs TTS uses a NF to parameterize parts of the model or controller networks tend to use similar things. It's more about thinking how your network works and communicates. I think a lot of people are just not thinking to hack away and operate on networks as if they're mathematical models instead of a locked box.
This is a bread-and-butter technique in NLP and machine learning in industry.
> directly training embeddings
This is literally the original embedding model, Word2Vec.
I look forward to seeing better tooling and literature around fine-tuning embedding models.
I kept running into people recommending against using them for long documents? Is openai better then the models sentenance-transformers uses? I found some recommendations to average together the embeddings of parts. I guess its cutting edge-ish still, a lot still feels like you can have a cool demo quickly, but something reliable and accurate is a lot of work?
People tend to "chunk" larger documents - the chunking strategy that's best is very dependent on what you are using the embeddings for. I've found it frustratingly hard to find really good guidance as to chunking strategies.
I've had good results for Q&A chunking my blog content up into paragraph sized chunks, as described here: https://simonwillison.net/2023/Oct/23/embeddings/#answering-... - but I'm not ready to say that's a universally good practice.
How to generate embeddings from the input query well is where one's focus should be IMO. An example: "don't mention x" being turned into filtering out / de-emphasizing chunks that align with the embedding for x.
I've been using these techniques along with pgvector and OpenAI's embeddings for https://flowch.ai and it works really well. A user uploads a document or uses the Chrome Extension on a webpage and FlowChai chunks up the content, generates embeddings, builds up a RAG context and then produces a report based on the user's prompt.
I hope that helps show a real world example. You're welcome to play with FlowChai for free to see how it works in practice at the application level.
Say I have an input document that I want to use as my search query and a database of potential results documents.
It's not clear to me that I should simply embed both the input document and each of the results documents. If the documents contain a variety of ideas, I'd be nervous they wind getting embedded into some generic space.
But I don't have any good ideas as to what type of chunk-match-and-aggregate strategy might work well here.
Would love ideas from folks that have done stuff like this!
I've found chunking by sub-headings works really well.
But here's a dirty secret I've learned as a writer: Your document can only ever contain 1 idea. That's the most that human readers can manage.
Atleast I'm not alone. Paragraph can make sense, I was worried about paragraphs that are related in concepts, they're not always standalone section of thoughts in real writing.
Maybe if two paragraphs meet some similarity threshold you can merge them.
Just spent a day updating it with latest benchmarks for text embedding models.
Pgvector is an extremely excellent way to experiment with embeddings in a lightweight way, without adding a bunch of extra infrastructure dependencies.
UPDATE: It looks like it's only available on the $50/month+ plans, my blog's database runs on a $9/month plan so I can't install extra extensions.
I wasn't able to find a good answer online.
Somewhat naively, I might speculate that for e.g. sequence prediction, even if you had some efficient packing of space like that to try to maximally separate individual tokens, it's still advantageous to learn an encoding so that synonyms are clustered so that if there is an error, it doesn't cause mispredictions for the rest of the sequence.
I suppose then the point is that the structure exists in the latent space of language itself, and your coordinate maps pull it back to your naive encoding rather than preserving a structure that exists a priori on the naive encoding. i.e. you can't do dot products on the two spaces and expect them to be related. You need to map forward into latent space and do the dot product there, and that defines a (nonlinear) measure of similarity on your original space. Then the question is why latent space has geometry, and there I guess the point is it's not maximally information dense, so the geometry exists in the redundancy. So perhaps it is obvious after all!
I think my comment was not worded properly. I was thinking "geometry properties = linear properties", what I really should say is:
Why does the latent space has geometry properties where we could use functions like cosine similarity to compare?
So when training, the signal will be mapped to latent space that will minimize the error of the objective function as much as possible.
Many applications already use cosine similarity function at the end the network, it would be obvious why they work. I reviewed other cost functions such as Triplet Loss. They use euclidean distances, so I guess it make sense why the geometry properties exist too.
For "and there I guess the point is it's not maximally information dense, so the geometry exists in the redundancy", what does "maximally information dense" means, I still don't quite get it.
Anyone remember?
Thanks for reminding me/us.
I’ll show myself out…
So ultimately they're not at all new (embedding data into a hamming space to compare it to other data), but the underlying meaning is much more useful in a world where you can peek into the internal state of a LLM to generate them.
It's very hard to know which embedding models are objectively good now that everyone is gaming the benchmark.
https://github.com/simonw/llm-cluster/blob/main/llm_cluster....
So does this take each row from the DB, convert to a numpy array (?), then uses an existing model called MiniBatchKMeans (?) to go over that array and generate a bunch of labels. Then add it to a dictionary and print to console.
(I'd call this an "algorithm" rather than a "model" - it doesn't have any model weights learnt from a training dataset)
For more details, see the pages in its user guide describing:
* the K-Means algorithm: https://scikit-learn.org/stable/modules/clustering.html#k-me...
* the Mini Batch variant of k-means: https://scikit-learn.org/stable/modules/clustering.html#mini...
Are there any good document embedding models I can run on my own hardware? Word2Vec is interesting, but I'd like to also try out cross-linking blog entries, etc.
pip install InstructorEmbedding
pip install torch
pip install numpy
pip install tqdm
pip install sentence_transformers
Use it: from InstructorEmbedding import INSTRUCTOR
import numpy as np
from scipy.spatial.distance import cosine
xl = INSTRUCTOR('hkunlp/instructor-xl')
embeddings = xl.encode(["I'm amazed at how autoencoding actually works, despite the simplicity of the approach.", "Are there any good document embedding models I can run on my own hardware? "]).tolist()
cosine_distance = 1 - np.dot(embeddings[0], embeddings[1]) / (np.linalg.norm(embeddings[0]) * np.linalg.norm(embeddings[1]))
print(cosine_distance)
Output: 0.33516267980261427More about the tooling I use to run embedding models locally here: https://simonwillison.net/2023/Sep/4/llm-embeddings/
How weird, the natural language world lends itself to classification with the same calculations as used in the images world.
Crudely put, embedding is the verb, metric space is the noun.
From there, people apparently just rolled with it and used the term more broadly, as tends to happen with language.
1: https://proceedings.neurips.cc/paper/2002/file/d5e2fbef30a4e...
Also, your interpretation of their use of the word embedding is exactly correct.
Did somebody ever run an image of a scene vs some parts of the scene in e.g., CLIP?
An added benefit is that if we're smart about how we construct the lower-dimensional space, we often find that it generalizes to unseen data better than if we'd used the high-dimensional representation, because a lot of the variation we're throwing away is specific to how we constructed the input data.
A really cool thing about embeddings is that we've known how to do this for a really long time --- for example, a landmark paper [1] on this exact problem was published in the 1930s!
[1] Eckart, C. and G. Young (1936), "The approximation of one matrix by another of lower rank." Psychometrika 1, p. 211-218. https://link.springer.com/article/10.1007/BF02288367
SageMaker and VertexAI are the AI services of AWS and GCP respectively, and they both offer embedding generation and vector databases (the two key pieces necessary for embedding search).
There are a bunch of smaller companies offering vector search as a service too, example pinecone to name just one: https://www.pinecone.io/
I also just found pgvector which seems like it could help? It has a similarity search.
Alternatively vespa cloud [3] offer both but… not the easiest to work with, it's tailored for businesses where search is a primary component.
Feel free to shoot me an email (in profile) with your context if you have more questions, in case I can help
[1] This model is a solid baseline if you're working with English text: https://huggingface.co/sentence-transformers/all-mpnet-base-...
[2] OpenAI's embeddings is probably the easiest to get started, and the API is straightforward. It's not the best performing embeddings for retrieval but good enough in some cases: https://platform.openai.com/docs/guides/embeddings/use-cases
Kind of wish we could have visuals describing that.
"PHATE is a dimensionality reduction and visualization tool designed to preserve both local and global data structure."
I got it to compare an image to an audio file which was pretty neat. I need to dig in more and see what kind of useful things I can use it for.
Now on how to get a compression vector from an LLM, simplified: Most ML models are built from different layers, executed one after another. Some of the layers are bigger, some are smaller, but each has a defined in- and output. If a layer's input size is smaller than model's input size, that must mean (lossy) compression must have happened to get there. So, you just evaluate the LLM on whatever you want to embed, and take the activation at the smallest layer input, and that's your embedding vector.
Not every compression vector makes for good semantic embeddings (which requires that two similar phrases are next to each other in the embedding space), but because of how ML models work, this tends to be the case empirically.
Can this be used to compress non-text sequences such as byte strings?
2. Yes, but lossily. Some types of byte strings are such that it doesn't matter if you accidentally change a couple of bits, some types of byte strings cannot tolerate that at all without being hopelessly corrupted. This technique is not a magic card to surpass the limits imposed by information theory, it's "just" a more sophisticated dictionary for your compression algorithm.
Once all the words have vectors you can assume that there’s meaning in there and move on to trying to math these vectors against each other to find interesting correlations. It looks like the scoring for the initial training is based on making the vectors computable in various ways, so you can likely come up with a comparability criteria different than the papers use and get a more useful vectorization for your own purposes. Seems like cosine similarity is good enough for most things though.
In general, the core task for the various "LLM tools" involves prediction of a hidden word, trained on very large quantities of real text - thus also mirroring whatever structure (linguistic, syntactic, semantic, factual, social bias, etc) exists there.
If you want to see how the sausage is made and look at the actual algorithms, then the key two approaches to read up on would probably be Mikolov's word2vec (https://arxiv.org/abs/1301.3781) with the CBOW (Continuous Bag of Words) and Continuous Skip-Gram Model, which are based on relatively simple math optimization, and then on the BERT (https://arxiv.org/abs/1810.04805) structure which does a conceptually similar thing but with a large neural network that can learn more from the same data. For both of them, you can either read the original papers or look up blog posts or videos that explain them, different people have different preferences on how readable academic papers are.
Embeddings are also super important in retrieval-augmented-generation (RAG), and getting the best "embeddings model" is important to achieve the best RAG performance. At Vectara we recently launched our new Boomerang model that pushes the limit on performance on embedding models, and I hope will spur more innovation and further improvements in this space.
https://vectara.com/introducing-boomerang-vectaras-new-and-i...