Which vector database should I use? A comparison cheatsheet
navidre.medium.com
navidre.medium.com
It also looses out on qualitative attributes that distinguish some of them from the others. E.g. Weaviate has a lot better DX (in my opinion) than any of the others as, as it handles integration of different vectorizers etc. a lot better, which makes it stand out.
As an aside, I’m not sure if I’ve just made a low value comment. If someone comes to the comments first, I hope they’re informed of my conclusion and take that into consideration before clicking through. I wonder how dang feels about these sorts of comments.
- necessary functions / use cases (eg prefiltering, dense search)
- embeddings version management
- anticipated embedding size (the article only considers glove-100 on ANN-benchmarks, which is quite different from openai-ada-002 1536 - both in terms of their output distribution and the vector size)
- required precision / recall
- required ingestion speed
- required ingestion throughput / time to ingest periodic updates
- required query speed (percentiles, not average!)
- required query throughput
- required RBAC, data privacy, active-active, etc
...and so much more. ANN-benchmarks is a good start for thinking about this but remember that actual throughput is quite different from whatever you see in the algorithms benchmarking!For many use cases, being able to put a proof of concept out of the door in hours vs days vs weeks is the top selection criterion if everything else is "good enough".
https://devblogs.microsoft.com/azure-sql/vector-similarity-s...
Scaling can come later, after the solution has proven its worth.
Also having some benchmark to compare performance would help.
[0] https://www.elastic.co/guide/en/elasticsearch/reference/curr...
OpenSearch specifically has an edge over Elasticsearch because it supports vectors up to 10k dimensions, whereas ES maxes out at indexing 1024 dimensions, which isn't enough to support OpenAI's 1536 dimension vectors.
And then there's the benefit of it being well documented / Q&A'd, and able to support regular searching, faceting, etc. as well.
You need to do dimensionality reduction before indexing. Basically it's fine to just pick n first components if you don't want anything fancy.
Given how well OpenSearch works and scales, I would find it hard to justify a specialized vector-specific database unless it brought A LOT of new benefits to the table. And I am not currently aware how any of them would actually do that.
Also, OpenSearch provides all of that out-of-the-box. You just configure a vector field mapping and start inserting your data. No need for an add-on plugin/extension. It just works.
This may not matter if you are not doing high throughput / have tight latency requirements, but in my case, it did. Of course you should weigh that versus the convenience of preexisting ES/OS clusters and so on. You can also use ES/OS together with a separate vector DB. (these tradeoffs are, of course, what make a static benchmarking post like this one so hard to think about).
I find vector search more convincing as a feature of an existing database than as justification to design an entirely new database - it's basically a new type of index.
I really hope Pinecone doesn't become the defacto vector DB. They're getting all the attention, but they're closed and crazily venture funded. That's going to turn into an Oracle situation fast.
I understand wanting to keep Amazon out of your business, but licences exist that allow that.
In practice, as long as search latency meets requirements (like, say, 100ms p95), the deciding factors tend to be things like cost for a given scale, amount of engineering overhead required or saved, reliability, features that affect search quality such as filtering and hybrid search, and so on.
Everyone has different workloads and different need. For example, I wouldn't recommend Pinecone to someone who just needs a pure ANN index like Faiss or HNSW on a single machine. Try out a few options and see what works for you... We make Pinecone easy + free to try for exactly this purpose, so you don't have to rely on a barebones "comparison" table from a third party.
In my mind the whole embedding / vector DB craze will come crushing down.
In other words, they published top-1 accuracy from top-2 accuracy calculations.
I would not over-index on that paper. However, I would err in favor of simpler methods.
Even so, driving the cost down for large workloads like that is a priority for us. We recognize the GenAI / RAG stack is a completely new line item in most companies' budgets so anything to keep that low can help these projects move forward.
I was playing around with making my own UI for interfacing with chatgpt. I saved the chat transcripts in a normal postgres DB, along with the open AI embeddings for each message in a vector db, with a pointer to the message id in postgres in the vector DB metadata.
Then as you chatted, i had chatgpt continuously creating a summary of the current conversation you were having in the background and doing a search in the vector db for previous messages about whatever we're talking about, and it would inject that into the chat context invisibly. So you can do something like say: "Hey do you remember when we talked about baseball" and it would find a previous conversation where you talked about so and so hitting a home run into the context and the bot would have access to that, even though you never mentioned the word "baseball" in the previous conversation -- home run is semantically similar enough that it finds it.
If you're using openai embeddings as your vectors, it's _extremely_ impressive how well it finds similar topics, even when the actual words used are completely different.
I recommend starting at https://www.pinecone.io/learn/vector-database/
This enables use cases like semantic search and Retrieval-Augmented Generation (RAG) as mentioned in the article.
Semantic search is: I search for "royal" and I get results that mention "king" or "queen" because they are semantically similar.
RAG is: I make a query asking, "tell me about the English royal family", semantically similar information is fetched using semantic search and provided as context to an LLM to generate an answer.
However, for large dataset deployment, cost becomes more critical since vector search is computation intensive. Anything like es, mongodb and redis can not even share their results in the benchmark.
Also, if you are looking for more fancy features rather simply ANN, purpose built vector database has faster iterations than traditional databases
You're also missing a lot of details. For example, Milvus and Zilliz are actually a little different, check this out for more details: https://github.com/zilliztech/VectorDBBench (of course run it on your own stuff, don't blindly trust companies just because their product is open source)
Also if you want to throw some more comparisons in their checkout elastic search
https://clickhouse.com/docs/en/engines/table-engines/mergetr...
https://docs.datastax.com/en/astra-serverless/docs/vector-se...
It would be nice to have another column to compare the largest scales these DBs can support.
Disclosures: None
It lets you run you run the benchmarks using your own API keys. Although it is made by Zilliz (maintainers of Milvus), you can take a look and see what is going on and judge if its fair.
Disclaimer: I am the author of txtai
"Redis can be a simple store, either with the embedding as the entire value, or as a value in a hash along with other metadata, or their newer vector search functions. Overall this works, but is more work than necessary, and not ideal for this use case."
Disclaimer: I'm the author (and work at Pinecone).
Etienne Dilocker, The Co-founder/CTO of Weaviate and Ram Sriharsha, the VP of R&D at Pinecone are both presenting at The AI Conference.
Lots of other smart people are presenting including Nazneen from Hugging Face, Harrison from Langchain, Jerry from Llamaindex, Ben the co-founder of Anthropic and many more.
A hackathon is happening in the evening at the event as well.
If you can't make the event, we'll put up all the talks on YouTube post-event.
More info at https://aiconference.com
Here are 5 free tickets to the event: www.eventbrite.com/e/487289986467/?discount=hack4free
Please only take one ticket each. They are first come, first served.
*This is my event -- Shameless plug *
Happy Monday!