298 karma · joined August 7, 2018
If you'd like to get in touch, you can email my username at my username dot com.
Your point about HNSW being resource intensive is one we've heard. Our team actually built another extension called pgvectorscale [1] which helps scale vector search on Postgres with a new index type (StreamingDiskANN). It has BQ out the box and can also store vectors on disk vs only in memory.
Another practice I've seen work well is for teams use to use a read replica to service application queries and reduce load on the primary database.
To answer your third question, if you combine Pgai Vectorizer with pgvectorscale, the limitations around filtered search in pgvector HNSW are actually no longer present. Pgvectorscale implements streaming filtering, ensuring more accurate filtered search with Postgres. See [2] for details.
[1]: https://github.com/timescale/pgvectorscale [2]: https://www.timescale.com/blog/how-we-made-postgresql-as-fas...
You are correct in saying that that you can store embeddings and source data together in many vectordbs. We actually point this out in the post. The main point is that they are not linked but merely stored alongside each other. If one changes, the other one does not automatically change, making the relationship between the two stale.
The idea behind Pgai Vectorizer is that it actually links embeddings with underlying source data so that changes in source data are automatically reflected in embeddings. This is a better abstraction and it removes the burden of the engineer to ensure embeddings are in sync as their data changes.
However, FAISS is not a database. It can store metadata alongside vectors, but it doesn't have things you'd want in your app db like ACID compliance, non-vector indexing, and proper backup/recovery mechanisms. You're basically giving up all the DBMS capabilities.
For new RAG and search apps, many teams prefer just using a single app db with vector search capabilities included (Postgres, Mongo, MySQL etc) vs managing an app db and a separate vector db.
Right now the system only supports OpenAI as an embedding provider, but we plan to extend with local and OSS model support soon.
Eager to hear your feedback and reactions. If you'd like to leave an issue or better yet a PR, you can do so here [1]
> 1. Required and Negated Words
This is not a bug of pgvector but a bug in embedding models and simple similarity search itself. You'd run into this issue on when doing RAG with any vectordb or ANN search library. You could probably solve this with query expansion, running multiple similarity searches in parallel, and doing filtering of results containing the required theme vs trying to get this all in a single search.
> 2. Explainability With Highlights
This is a misunderstanding of the purpose of embeddings based semantic search. The point of semantic search is not to match on exact keywords but to match on the /meaning/. If you want exact keyword matching, use full text search. This is something that can also be solved by hybrid search, combining full text search and semantic search and using a re-ranker. PostgreSQL has built in FTS with tsvector.
> 3. Performant Filters and Order By’s
The authors do not disclose any details about what index they use here or what kind of filtering they are trying to do. As other commenters point out, the StreamingDiskANN in the pgvectorscale extension [0] (complement to pgvector, you can use them together) improves on performance and accuracy of filtering vs pgvector HNSW (see details in [1]).
>4. Support for sparse vectors, BM25, and other inverse document frequency search modes
This is probably the most fair point in the post. But there exists projects like pg_search from ParadeDB which bring bm_25 to Postgres and help solve this [2]
Lastly, I respect companies trying to provide real-world examples of the trade-offs of different systems, and so thank the authors for sharing their experience and spurring discussion.
Disclaimer: I work at Timescale, where we offer pgvector, and also made other extensions for AI/ vector workloads on PostgreSQL, namely pgvectorscale, and pgai. I've tried to be as even in my analysis as possible but as with everything on the internet, you can make up your own mind and decide for yourself.
[0]: https://github.com/timescale/pgvectorscale/ [1]: https://www.timescale.com/blog/how-we-made-postgresql-as-fas... [2] https://github.com/paradedb/paradedb
Here's an overview of the methods [1] (all except separate database of course)
We're excited to release pgvectorscale. Our team built this extension to make PostgreSQL a better database for AI and to challenge the notion that PostgreSQL and pgvector are not performant for vector workloads.
pgvectorscale is open-source under the PostgreSQL license and free to use on any PostgreSQL database.
Here are two helpful companion reads to the post linked by OP: A benchmark of how PostgreSQL with pgvector and pgvectorscale performs against Pinecone [1], and a technical deep dive into pgvectorscale's StreamingDiskANN index and Statistical Binary Quantization implementations [2].
Questions and feedback welcome!
[1]: https://www.timescale.com/blog/pgvector-vs-pinecone/ [2]: https://www.timescale.com/blog/how-we-made-postgresql-as-fas...
[0]: https://www.timescale.com/blog/how-we-made-postgresql-the-be... [1]: https://github.com/timescale/python-vector [2]: https://www.timescale.com/ai/#resources
[0]: https://www.timescale.com/blog/how-we-made-postgresql-the-be...
To your question about RAM usage, we provide a graph of index size. When enabling PQ, our new index is 10x smaller than pgvector HNSW. We don't have numbers for HNSWPQ in FAISS yet.
Totally fair about cloud-only. Many developers prefer developing on cloud, but some prefer local dev. YMMV.
By the way timescale vector offers pgvector as well so it's easy to test and compare. (Note: I work for Timescale)
[0]: https://www.timescale.com/blog/when-boring-is-awesome-buildi... [1]: https://www.timescale.com/blog/40-million-to-help-developers...
One clarification. While TimescaleDB is open-core, our community version is source-available and 100% free to use. We do not "hold back features for customers". You do not need to pay to use any of TimescaleDB's best features, it's all free via the Timescale Community license.
You only pay if you'd like to use our hosted offerings (and save the hassle of self-managing your DB): Timescale Cloud or Managed Service for TimescaleDB.
For more see: https://www.timescale.com/products
(Disclaimer: I work at Timescale)
There's also comparisons of TimescaleDB vs MongoDB[1] and AWS Timestream [2].
[0]: https://blog.timescale.com/blog/timescaledb-vs-influxdb-for-...
[1]: https://blog.timescale.com/blog/how-to-store-time-series-dat...
[2]:https://blog.timescale.com/blog/timescaledb-vs-amazon-timest...
Within the core database, we offer features that are carefully marked as "experimental", which we discuss at length in this blog post [1].
Beyond TimescaleDB, we also offer other products that are more SaaS-y in nature. While they're all based on the rock-solid foundation of TimescaleDB, we are also able to ship new features more quickly because they are UI components that make using the database even easier.
Finally, some of our "launches" are more textual in nature, such as this benchmark, which we have spent months researching and compiling.
[0]: https://blog.timescale.com/blog/when-boring-is-awesome-build...
[1]: https://blog.timescale.com/blog/move-fast-but-dont-break-thi...
From the blog post: "So today’s serverless data platforms are not familiar or flexible. But further, black boxes are never truly easy and worry free: you never know if there are any skeletons lurking in the proverbial closet, just waiting to cause your service to fall over."
(Timescale employee here)
That said, I'd be curious to hear about other folks experiences.
Disclaimer: I work at Timescale.
[1] https://hasura.io/blog/using-timescaledb-with-hasura-graphql...
Many users actually find Timescale Cloud to be cheaper because it replaces both a time-series and relational db with a single database that can store (and analyze) both kinds of data.
[1] https://www.timescale.com/cloud [2] https://blog.timescale.com/blog/announcing-explorer-a-better...