> - It seems like these embeddings are learned during the training phase. Presumably by backpropagating to the embedding during each epoch?
As far as I understand it, your first part is accurate.
> - Given that they are learned, doesn't that mean that they are completely context specific to a given trained model? Ie, they can't readily be shared on their own?
Yes, but you only have to use that context model in order to generate new embeddings to add or search the existing store. You can then use whatever is found wherever you want. Think of the vector store as a searchable database, where the vector is unique to the exact text embedding (like a hash function I believe, but I think without the possibility of collisions). You can search and compare vectors, then use the vector to determine what text was embedded (using simple relational queries or storing the text in the same row as the vector).
It means you can store chunks of information as embeddings in a vector database, search for similar content to a query and get chunks of text back related to what you searched.
> - What does "similar" mean here? Is there some emerging practice on how to determine how close is close "enough" between multiple vectors for the purpose of similarity searches? Is this too determined by the model weights somehow?
Cosine similarity search is the most used recently because it's what OpenAI recommends. There's also dot product. They are doing calculations to find the closest other vector(s) across the dimensions of the embedding vector. A 2D plane with a vector on it is 2 dimensions. Some models embed 700-800 different dimensions. OpenAI uses 1536 dimensions.
Imagine comparing a dozen 2D vectors to each other to find the three with the closest angle to a chosen other one. Somewhat easy to move them and overlay, and if they are pointed in the same direction they are likely related. Doing so in 1536 dimensions can't be imagined, but there is still an angle between the vectors with a dot product. It's this mathematical similarity which is used.
> I hope I'm not missing something fundamental with these questions. Or maybe I hope I am missing something and someone points out my errors. That's good too. :)
The use of the Postgres vector store is more from the use after an embedding model has been trained and there is a use for the data being embedded, like helping customers shop for similar items based on their search, and returning descriptions; or feeding a user question in to find context to feed to an LLM for few shot prompting.