> What if I have a billion or more embeddings? Can Postgres + pg_vector support a billion records with 1000’s or 10000’s concurrent users? Even if pg_vector is an easy add-on to Postgres, it might not be suitable for real-time similarity searching. You should use a database that is built for scale.
Has anyone tried pg_vector on an AWS Aurora cluster or Azure Citus cluster?
At the scale of hundreds of millions of embeddings (the most I have experience with) pgvector has some sharp edges. Vacuuming is a huge one; it’s difficult to vacuum indices that are hundreds of GBs in size. And reasoning about recall is therefore challenging. Pgvector’s HNSW implementation also lacks support for certain features (e.g. filtering), which is fine when silently falling back to brute force is acceptable but otherwise not so much. To be clear I think pgvector is an awesome project but it certainly has limitations.
As part of the massive shift into AI, vector databases have been increasing in popularity. Also known as vectorized databases, they play a crucial role in the context of AI, so it’s important to understand how they work. To do so, we’ll need first to understand what vectors are.
noob here. how does it know that the word description embedding matches with the image embedding? how does it know the color red is in that image?
The model is trained by giving pairs of examples. So that it learns to generate an internal representation of images, texts or whatever inputs.
There are other ways to train, but this is the basic example.