Supabase Vector, the Open Source Vector Toolkit for Postgres
supabase.com
supabase.com
The core of the API is this:
import vecs
DB_CONNECTION = "postgresql://<user>:<password>@<host>:<port>/<db_name>"
# create vector store client
vx = vecs.create_client(DB_CONNECTION)
# create a collection of vectors with 3 dimensions
docs = vx.create_collection(name="docs", dimension=3)
# Add embeddings
docs.upsert(vectors=[("vec0", [0.1, 0.2, 0.3], {"year": 1973})])
# Query embeddings
docs.query(query_vector=[0.10,0.21,0.29], limit=1)
This stores all the data in a new `vecs` schema in your database. It took me a while to grasp the nature of structured (eg, manage tables/data with migrations) with unstructured/nosql, but the use cases for vectors are very often geared towards data scientists/engineers. There is some more info about the interop between these approaches here: https://supabase.com/docs/guides/ai/structured-unstructuredI think the argument against pgvector would be "it doesn't scale", but we're using it against hundreds of thousands of rows of data and the performance is solid (< 100ms). Maybe we'll run into issues with tens of millions of rows, but we'll worry about that when it actually becomes a problem.
On "bad queries", we rely heavily on the users past SQL history to understand correct JOIN's and how they normally define metrics / use the database in general.
That said, using pgvector (or using other SQL databases with vector search support.. many have this capability) will let you do both ANN and your usual SQL filtering and joining (or full text, or.. etc) to produce more hand tuned results to a query. This is something the specialized vector databases don't have much support for.
It’s super useful for hybrid search: vector similarity search over unstructured data joined with filtering by structured data.
One caveat: It’s important indexes are added after a fair number of vectors have been persisted, which can make some production use cases challenging to implement.
managing indexes are the most challenging part of using pgvector right now. There is a bit of a "goldilocks zone" - you can't create the index with zero data, but if you create it too late then it can take a lot of memory. Andrew (and a few others, like AWS) are working hard on improving this, as well as researching other index types.
Without an index, I found it was too slow for data that size. I didn’t expect issues for data at this “small”, but it’s pretty new to me.
One of the benefits of using Postgres + pgvector, is it is much easier to delete records from a vector index versus using a vector database or Redis. While Redis has ttls for objects, there isn't an easy way to remove data from a vector index.