Vector databases: analyzing the trade-offs
thedataquarry.com
thedataquarry.com
We then launched vector search in Jan 2023, and just last week we launched the ability to generate embeddings from within Typesense.
You'd just need to send JSON data, and Typesense can generate embeddings for your data using OpenAI, PaLM API, or built-in models like S-BERT, E-5, etc (running on a GPU if you prefer) [2]
You can then do a hybrid (keyword + semantic) search by just sending the search keywords to Typesense, and Typesense will automatically generate embeddings for you internally and return a ranked list of keyword results weaved with semantic results (using Rank Fusion).
You can also combine filtering, faceting, typo tolerance, etc - the things Typesense already had - with semantic search.
For context, we serve over 1.3B searches per month on Typesense Cloud [3]
[1] https://github.com/typesense/typesense
[2] https://typesense.org/docs/0.25.0/api/vector-search.html
I'm surprised the original post doesn't benchmark Typesense.
Let's say your dataset has the words "Oceans are blue" in it.
With keyword search, if someone searches for "Ocean", they'll see that record, since it's a close match. But if they search for "sea" then that record won't be returned.
This is where semantic search comes in. It can automatically deduce semantic / conceptual relationships between words and return a record with "Ocean" even if the search term is "sea", because the two words are conceptually related.
The way semantic search works under the hood is using these things called embeddings, which are just a big array of floating point numbers for each record. It's an alternate way to represent words, in an N-dimensional space created by a machine learning model. Here's more information about embeddings: https://typesense.org/docs/0.25.0/api/vector-search.html#wha...
With the latest release, you essentially don't have to worry about embeddings (except may be picking one of the model names to use and experiment) and Typesense will do the semantic search for you by generating embeddings automatically.
SOOOO much more matters than that! Any production database is going to have a huge medley of concerns and constraints. A reasonable recall at reasonable speed, but much easier to integrate and maintain, is going to be far far preferred. Not to mention a good retrieval system needs a broad range of features than just dense vector retrieval.
It's a sign the space is maturing away from being an academic/"benchmarking" competition space to one with actual industry concerns.
Agreed. My team recently settled on Qdrant because it was fast and painless to set up and get started using.
txtai is an all-in-one embeddings database for semantic search, LLM orchestration and language model workflows.
Embeddings databases are a union of vector indexes (sparse and dense), graph networks and relational databases. This enables vector search with SQL, topic modeling and retrieval augmented generation.
txtai adopts a local-first approach. A production-ready instance can be run locally within a single Python instance. It can also scale out when needed.
Vector search still feels like more of an index type feature than a separate product to me.
txtai is more than just a vector database. It also has a built-in graph component for topic modeling that utilizes the vector index to autogenerate relationships. It can store metadata in SQLite/DuckDB with support for other databases coming. It has support for running LLM prompts right with the data, similar to a stored procedure, through workflows. And it has built-in support for vectorizing data into vectors.
For vector databases that simply store vectors, I agree that it's nothing more than just a different index type.
I assume that is the rationale behind a dedicated database but I generally feel the same way as you
It's similar in that it writes data to vector index formats such as Faiss, Hnswlib. It has metadata filtering via SQLite/DuckDB to filter on additional fields.
It's different in that it can use other vector databases for it's file format. And it has significant logic via workflows for data transformation. Then there is the graph component for topic modeling.
So that's where I came up with the term "embeddings database" which I consider a vector database and much more.
I’m just not sure if I’d consider it a database. It’s just a long lived cache for us.
I'm working on something that needs a similar, small and fast, vector search implementation. Crucially we also need fast indexing speed for our usecase, but a bottlneck we're hitting is the time it takes to generate vector embeddings for larger documents in our dataset (a few megabytes in our case). Wondering what's the fastest way to approach that?
Another option if you just want search (and aren't training or tuning your own models) is a managed search offering where you aren't responsible for generating embeddings.
Naively I guess, at first we hoped to get by using a 3rd party API. We're hosted in GCP and tried using the Vertex AI `textembedding-gecko` model initially. But now we're investigating running models on our own infra, although not sure where we've got with it yet as someone else is working on that.
The full RAG workflow - using ADA to embed the user input, deserialize embeddings, run cosine similarity, and call gpt-3.5-turbo - is about 3 seconds end-to-end to get a result.
OpenAI embeddings are 1 per request payload, right? Have you hit any rate limits doing that?
We have a performance budget of ~1 second for the generate-index-search pipeline, which may or may not be feasible. I discounted OpenAI because it seemed like we're guaranteed to hit the rate limit if we flood them with concurrent requests for embeddings. Typical corpus size that we need to work with is 20 concurrent documents ranging from ~100kb to ~2mb. Chunking those documents to fit the 8k token context window balloons the request count further.
When I work in Python, my favorite is absolutely Chroma embedded.
The article is wrong about Chroma embedded being in memory only. It also works fine with write through to a local disk, including maintaining index files.
pip install milvusIt can run embedded in a single Python instance and has no issues running in production that way.
While it’s great to see efforts to make sense of the (admittedly noisy) vector database market, I’m struggling to grok large chunks of this. For example I can’t tell what the author means by “serverless”, but given that they put a whole bunch of open-source, self-hosted solutions in that part of the diagram it’s definitely not the commmonly understood meaning.
For anyone diving into the topic, here is another introductory article to help you: https://www.pinecone.io/learn/vector-database/
edit: pinecone is mentioned in the earlier parts of the blog series, though.
I don't even really love Qdrant or chroma (haven't tried weaviate, but it's at least in the right region according to the article), but at least they are embedded.
I pretty much refuse to use any DB that requires using API keys, putting data off premises, and even if it requires setting up ACL. I don't even use postgres much for the complex ACL and having to set up ports reason.
SQLite and DuckDB are truly incredible, can store gigantic databases (>2TB is still perfectly quite performant) and you can just hand any collaborator the entire DB on a disk, without having to worry about complex password junk.
I am also a bit surprised at the hostility to the pinecone employee here. I see a bunch of other companies and projects jumping in and everything is cool. But for some reason the pinecone dude catches a lot of grief.
What’s going on? Is there some weird toxic subculture in the vector database space that’s got it’s knives kit for pinecone for some reason?
(I work for Weaviate)
Unfortunately, some players in the space (who are on this list) are cheating and playing an unfair game. I guess that this comes with a rapidly growing space.
I hope we can all quickly go back to focussing on our respective communities and educating the market (together) on the awesome things one can do with vector DBs.
This is false
> […] and most popular vector database
Based on what?
Some unique features of MyScale:
1. This solution is built on ClickHouse and offers comprehensive SQL support. Our users leverage vector search for a wide range of interesting OLAP use cases.
2. We utilize a property vector search algorithm called the multi-tier tree graph (MSTG). This algorithm is significantly faster than HNSW for both vector index building and filtered vector searches.
3. We utilize NVMe SSDs for the vector index cache, which greatly reduces the cost of hosting millions of vectors.
Most frameworks, like Haystack, can wrap embeddings generation for you.
I would say that most people seem to prefer an engine that embeds and stores things as a service, but using Instructor is only a few lines of code and runs locally.