Every database will become a vector database sooner or later
nextword.substack.com
nextword.substack.com
As the traditional db move forward in the space the need for dedicated vector databases will likely shrink, except for some very specific implementation that offer unique enough features (I.e. deeplake does vector search over object storage, which is very convenient for certain specific scenarios)
Even if you could make it perform well, it would not do what you want.
https://github.com/asg017/sqlite-vss
Not associated with the project, just love SQLite and find it very useful.
The benefit is that you don't have to pay for the compute part of a database, and the storage layer is as cheap as it could be on the cloud.
retrieve the embedding index and to run an indexed search to identify the data to be retrieved. Please bear with the layman like questioning -
So if the data is {obj: "obj1, "data": {"name": "atlas", "embedding": "1123124234" } What is an embedding index ? Is it something like {"1123124234": "obj1"} ?
From what I understand the query will be "geography" whose embedding will be "12311111" and now you have to run a KNN for a match which will return {"name": "atlas", "embedding": "1123124234"}
Not sure where the embedding index comes into play here.
Now if you want to do similarity search you have to measure the distance between 2 or more vectors and that's independent of the indexing. No ?
So any database with sufficient memory should be able to accomplish this as evidenced by the vector similarity search feature of Redis. ( I don't know how Redis folks have implemented vector similarity but they do support KNN search )
There are alternative index types, of course, or you could index the hash of the vector. These both come with tradeoffs.
Aside, you can increase the size of tuples you can index in a postgresql btree by increasing the postgresql page size (requires a recompile and creating a new database instance).
Of course ACID scales to well into the Fortune 500 scale so...
They mostly seem a tarted up associative array, sure, but a key-value store is a thing.
DynamoDB underpins much of AWS which in turn underpins a ridiculous number of web services.
So definitely more than just a thing.
NoSQL has been around for over 20+ years.
Since then Cassandra, DynamoDB, FoundationDB, MongoDB, Neo4J, Redis etc are not only still around but widely used and powering many of the services you use today.
Even looking at HN's reactions to that video show a few comments that did not age very well (although most of them did): https://news.ycombinator.com/item?id=1636198
The sensible takes were that NoSQL would supplement and enhance RDBMSs, but the hype was much more than that.
RDBMS is such a mature and powerful technology, and with the vast power in modern single-node hardware, they will scale to considerable sizes. But there is a limit.
Once a table or set of tables hit a certain scale, it needs to be distributed. Once you get to those scales, you are likely dealing with the threat of exponential data growth, so doing the "distributed Postgres" will bandaid the problem ... but you are starting to run into the CAP theorem's problems.
You'll need to AP-scale the biggest data tables, and since that almost always means a big coding change lift and introduction of an AP-scaling database, it will almost always be a six months to a year transition, and possibly adding an entirely new DB technology (Cassandra / DynamoDB / maybe FoundationDB).
I have yet to have anyone explain how joins scale on an AP distributed database except in limited situations where the joined data is somehow node-local to the other tables, usually some hierarchical situation. Otherwise you are pulling data from lots of nodes and aggregating and comparing the different sets to account for node drift / partitions / network failures. Cassandra and Dynamo basically say "you are scaling a single table/query/update/pre-joined data table".
Which really isn't fun for RDBMS folks. Because it is a shitload of denormalization on top of all the AP headaches and distributed transactions / updates.
As I said, megahuge machines really make that an outside case unless your use case really emphasizes the "A" in CAP, where write speed can be tolerated to be low and you want to tolerate entire cloud or datacenter outages.
But yeah, nosql was never going to kill the rdbms. The functionality/power of rdbms/SQL is so so so so much higher. Just be aware when you're about to shoot over the limit and prepare for the code changes to handle true scale in the (unlikely) event you're going to need it.
Massive RDBMS instances are awful to manage. Hard to backup/restore/fork. Hard to migrate tables and schemas without causing downtime. Accidental downtime happens all the time due to locks, bad indexes, bad query plans , etc etc. at large scale they are capricious beasts, care and feeding and most importantly changing them becomes a dark art.
Don’t let them get too big you’ll regret it :)
Firstly, even for non vector data, read/write transactional database vs read-optimized store purely for fast serving are already markedly different. Then, the shape of data that is used to generate embeddings is markedly different than the shape of data that is ready to transact or serve.
So, no matter where it is stored, it has to leave that store, get transformed and enriched and then run through an embeddings generator (ML inference).
Then, it has to be stored in a manner that is optimal for retrieval ranking. If you are doing ANN that's one thing, but if you are doing attribute based filtering while retrieving and you wish to accelerate it through GPU to do more exhaustive search, that's another thing altogether.
All these lead to fairly sophisticated optimized implementation. Sure, a singular database product solution that has all these different optimized engines can emerge over time but surely it is too early today to converge like this.
Same for pretty much all the data in your database. It comes from an app, that does validation/editing/transforms. The benefit is that it's all together, can be atomic, only requires one query to get, update, and delete.
Vector databases are just normal databases with a vector index. There's no reason for you to have a specialized DB for it.
Also, embeddings aren't inference, it's a token lookup. There's no forward pass.
Now imagine you are using that database for storing transactions and other day to day business ops that will still be storing millions of records but with small indexes. This would have ideally only required a single DB instance with a replica for redundancy. Now if you integrate Vectors into the equation, you will have to needlessly scale this DB both horizontally and vertically just to maintain a decent query/write performance to your DB (which would have ideally been extremely fast without embeddings in the mix). You will eventually separate the embeddings out as it makes no sense for the entire DB to be scaled just for the sake of scaling your embeddings. I am not even accounting for index generation for these vectors which will require nearly 100% of all CPU cores while the index is being generated (depending on type of ANN you are using) and which in turn would slow your DB to a crawl.
Someone makes the example in another comment, but it’s analogous to OLTP vs OLAP
Now you do NOT want to run such a setup on the same hardware that you use for your transactional systems, of course. But you CAN use the same software (like Oracle), which means that you do get some reduction in tech complexity.
Clustering, load balancing, aggregating queries etc are quite different for a vector database in comparison to traditional OLTP databases.
It's the same as difference between OLAP vs OLTP. Both have different underlying architectural differences which make it incompatible for both to run in an integrated fashion.
For instance, in a traditional DB the index is maintained and rebuilt alongside data storage and for scaling you can separate it into read/write nodes. The write nodes typically only focus on building indexes while the read nodes for querying eventually consistent indexes (eventual consistency is achieved by broadcasting only the changed rows rather than sending entire index).
Now it's similar in vector dbs too. You can seperate the indexer from query nodes (which access eventually consistent index). However, the load is way higher than a regular db as the index is humongous/takes a long time to build and sharing the index with query nodes is also more time consuming and resource/network intensive, as you won't be sharing few rows but the entire index itself. It requires a totally different strategy to get all query nodes to be eventually consistent.
The only advantage of traditional DBs also implementing vector extensions is familiarity for the end user. If you are already familiar with postgres you wouldn't want to leave your comfort zone. However, scaling a traditional DB is different from scaling a vector DB and you'll encounter those pain points only in production and will be forced to switch to proper vector databases anyways.
For basic CRUD we use the Supabase endpoints directly but none of that involves querying a vector db :P
I'd really love to know what kind of insane scale justifies that tradeoff...
Of course if I sounded incredulous it's because I didn't think you had that scale, and it sounds like I was correct?
We don't also "pipe" our data to Supabase, we use a couple different data stores depending on the best use case. For example we also use R2 and Durable Objects.
Just because you have a hammer doesn't mean everything is a nail.
Also don't you think it's funny calling out somebody else's tech choices when you have zero insight into it and when it's worked out perfectly for us?
By the way, how many TBs of vector data are you storing in Postgres and needing to retrieve with minimal latency?
Literally irrelevant lol
Previous uiua discussion: https://news.ycombinator.com/item?id=37673127
Brainfuck: https://en.wikipedia.org/wiki/Brainfuck
But a question for true DB experts here:
1. Is there any real advantage to building a dedicated vector DB from scratch?
2. Is vector DB something that can be just 'tacked on' to a normal DB with no major performance penalties?
We know from history, that data warehouses are genuinely different from databases, and cloud data warehouses are overwhelmingly superior to on-prem ones. So that emerged as a distinct, enduring category with Snowflake/Databricks/Bigquery.
Most vector databases use one of a few different vector indexing libraries - FAISS, hnswlib, and scann (google only) are popular. The newer vector dbs, like weaviate, have introduced their own indexes, but i haven't seen any performance difference -
Reference: https://ann-benchmarks.com/
Some advantages of having a separate index is that it can work with different backends, it can be independently scaled, and it can index data for more than 1 database server.
Some disadvantages are increased latency, increased complexity, and distributed system problems.
The operational pains if you need to self host this stuff are real, split brain, backup/restore not really considered (compared to a normal databases features), things like replication and sharding _exist_ but often are a buggy mess.
OLAP is definitely distinct from OLTP, and most of these vector queries have some aspect of both - they are similar to OLAP in that they need a decent amount of preprocessing to be useful (inferrence) and they are similar to OLTP in that they are often used for serving point queries or tiny lookups.
Yet OLAP databases continue to thrive alongside OLTP databases, the nascence of NewSQL hybrid (HTAP) databases notwithstanding. Different needs dictate different design choices for optimality.
I guess it makes sense because the infra is so different, but I’m not sure whether it need be.
How does that relate to the OLTP vs OLAP dimension? Are they not both primarily OLTP dbs still?
"HTAP" is the buzzword you're looking for. It's promising, but also complex and nascent. It'll be interesting to see how much traction it gets over time, but things like TiDB and Unistore are pretty early on to call a nail in the coffin for redshift/bigquery/clickhouse etc.
Could not agree more. Even for time series, which could be seen as a subset of OLAP, trade-offs and design choices inherent to time-series data are necessary. As an example of a TSDB that I know well, QuestDB: Data is always ordered by time once it lands on the disk, the data is partitioned by time, and the ingestion protocol is conceived to stream large volumes of data, which can be either continuous or in bursts.
You pair a vector db with a metadata store (can be anything, but ideally you want low gravity with the vdb and disk for fast retrieval... i.e. leveldb, sqlite equivalents... or hell a traditional db - and the author is right traditional dbs dont need to work too hard to create a ANN extension)
In general, the greater challenge is the data engineering and overhead of managing retrieval stores (in terms of the data-integration/model pipelines) ... so I'm bull on the solutions addressing opportunities here.
It's why data engineering is a thing in our industry. We move and prepare data for a set of tools, and we pay good money to do so, because we believe we derive value from those tools.
Let's say MySQL offers it, anyone already using MySQL is likely to fence the MySQL instance(s) focused on vector stuff off for various reasons (resilience, different read/write patterns, security, etc.)
MySQL as the (imaginary) basis only offers some transferable skills, because this DB will require different care and feeding.
Like the difference between Postgres and PG with cstore_fdw, similar, but sufficiently different.
The view that everything needs to support direct input for generative AI is short sighted. There are other use cases as well. Even if ultimately these will become just building blocks for whatever AGI there comes. Horses for courses
> PlanetScale has forked MySQL to add vector storage and search! You’ll be able to support your AI and ML applications with the world’s most scalable database platform. This unifies the reliability and functionality of MySQL with the ability to store vectors and perform similarity search.
I’m not criticizing the specialized case for a true vector database, but for most workloads I agree that the big database players will be the right choice for many users.
My starting point for pretty much any storage problem is "have you tried throwing it in postgresql" but this is one reason pulling something else off the shelf might be a good idea.
This sounds a problem more similar to what 3D physics engines can do, but generalized to higher dimensions, as opposed to traditional text and key-based database stuff.
The algorithms and data structures in physics engines (bounding volume hierarchies, kd-trees, etc.) are quite different from how a traditional database index (B-trees and skip lists are popular there, if I remember correctly) is searched and stored.
So, by using databases which can efficiently answer geometric nearest-neighbor questions, you can quickly search for chunks of text that are similar to each other.
For example, the text "chocolate milk" is all the same characters as "milk chocolate", but likely have very different usage within the context of retailers or cooks. So, their vectors should be very different.
Word2Vec is NLP that uses a neural net to build these vectors: https://en.wikipedia.org/wiki/Word2vec
Vector databases are a gimmick at the moment. Ultimately conversational AI agents should be able to extract information from a diverse set of sources with a diverse set of tools. The approach that is currently taken is hit-and-miss at best. How often do you searched something and the first result happens to be the thing you are looking for? Why should it be any different with vector DBs? Obviously the query matters a lot no matter how the information is searched.
Embeddings work really well to store semantic meaning and are great for searching. Or, at least, a 1st stage of searching to filter out the non-relevant content.
I'm working on my own "notes" app, based on embeddings because I 'm tired of never finding what I need due to bad search/tagging/categorizing
Searched where? Every search powered by a large tech company has almost certainly been using vector search for years. Then combined that with other non-vector results. Then run that through numerous ranking models. Then showed that to you. It's far from perfect but absurdly better than the average ElasticSearch results you might get elsewhere.
I bring this up because the ad hoc analytical use case for graph stores is so niche most engines haven't even seen enough demand to introduce it because you can always store graph relationships and offer retrievals in those engines to a limited depth which is typically sufficient for most operational / transactional use cases.
I can help you put right now. You don't even need a database to have a "vector database". FAISS is an in-memory "vector database" that runs on the data that you happen to have. So: if you data is stored in a .txt file, load it into memory, index it with FAISS... bam a vector database from a data file.
Can we take X arbitrary "real" database and implement KNN search on top of 1000 indexed columns, I'm sure it is possible - but I'm also pretty sure most databases will die under the pressure (source: I've asked some of my favorite DBAs if I could do this and they said "no")
They allow you sort results by the cosine similarity[1] between vectors. The idea is that you can attach a vector to documents in the database and then when you pass a vector in the query and get back documents that most match the query vector.
The function that creates these vectors(string -> vector) is called an embedding and is constructed in such a way that "semantically similar" strings have vectors that are close together.
It's not a very complicated idea, but complicated and powerful are orthogonal concepts.
They are useful in AI(LLMs) when you would like to include documents in your prompt that are relevant to your instruction. The best way to describe this is by example.
Imagine your query is "What is the capital of France?" Rather than requiring your LLM to encounter this fact during its training, you can embed the question("What is the capital of France?" and retrieve documents (say you've indexed all of Wikipedia in your vector db) and return some snippets from articles that include this information(context).
You then pass the the prompt+context to an LLM and given that it now has relevant information, it can answer the question.
You can also imagine that it's much easier to update a vector db with new information than it is to retrain a model to ingest new facts.
One can imagine there is some optimal way to lay out the data in memory such that it would be relatively quick to do that. One can imagine a naive way to do it which probably wouldn't be fast. A vector DB does the first thing - lays out the data in a way which then enables it to do that fast.
And then it does all the other stuff a DB does - persisting to disk (which means data needs to be laid out in a sensible way on disk too), handling multiple queries, updates, and the 50 other complicated things databases tend to do. Users generally want to do other operations on vector as well so a vector DB does those too.
For a small number of vectors you can build a vector DB yourself. Write a list of vectors to a file, load them into memory in no particular order, then for your "n closest" function, just iterate through the list calculating the difference in direction one by one and keeping the top n. Your simple system will work just fine for a toy demo.
Here:
https://chat.openai.com/share/9e557a90-e127-4654-9271-7c51fd...
Let's say you have a chatbot and stored in its database is the usual info like session id, timestamp, message etc. To "vectorize" this db then would be to vectorize all the messages. Is this too simple an understanding?
Once the db has been vectorized then we can do semantic search on the messages and create more informative graphs based off the semantic similarity for messages within a given timeframe or other criteria.
btw I'm working in a DB startup - https://hyper-space.io/
(Disclaimer: I work for Pinecone, so obvious bias ahead but also perspective of 3 years since launching the Vector DB category and actually seeing billion-scale vector search deployments.)
> Basically, having separate vector DBs can add to cost and complexity. Imagine you were a MongoDB shop, with over 500m documents stored cross-region. If you are using a separate vector DB, say Pinecone, that may require moving potentially billions of embeddings between two databases, cross regions. This costs a lot, not to mention complex, since you are responsible for generating the embeddings... It’s faster, cheaper, and simpler if one database (Mongo, Elastic) just supported vector search.
If you want, say, 100ms search latency on just 100M vector embeddings in Elastic that'll already cost you $12,600 per month at minimum. And if you regularly write new or updated data to the index then your latencies will creep up until eventually you have to run a "force merge" which will grind your vector search to a halt for several hours (so much for easy and simple). I don't know how much it is on Mongo but given that it's bolting on the same vector index I would guess it's in the same ballpark. The cost grows sublinearly with more embeddings. (Pinecone is around 60% less than that, and will be even less soon.) The suggestion that having "billions" of embeddings in a traditional DB is easier and less costly shows you exactly why you should run your own tests and see for yourself.
When traditional database companies bolt-on vector indexing libraries such as HNSW[0] on top of their existing architecture, it's to meet demand from their existing users that have a relatively basic need for vector search.
For very basic and small-scale use cases, like <10M vectors with a relaxed data freshness requirement, you should just use whatever is the most convenient. Sometimes that's Pinecone, and sometimes that's the database you already have. (And if your current DB doesn't offer basic vector search, just wait two days).
When it comes to larger scale, like 100M+ vectors, if you want any hope of meeting performance, cost, and data freshness requirements then you should look at a purpose-built vector database. As GenAI workloads start to enter production and scale, a lot of people will see find this out the hard way.
This has been true for every unique data structure and querying pattern for the past 40 years and it’s true for vector embeddings and vector-based retrieval. You can't blame the proliferation of different database types on hype and VC funding alone.
But don't take my word for it either. Go and run some tests that resemble your production workloads, then do what makes sense for your use case!
This is not even a dig at Elastic. The problem is deeper than that… It’s an issue with the underlying vector index they (and many others) chose to bolt on, HNSW, which was not designed with frequent live updates in mind.
We have a post coming soon that covers the technical parts of this in more detail. You asked a good question.
I'd love to see someone who has some expert knowledge of elastic chime in to hear if the characterization here seems right? But admittedly not everyone's going to be a power user so if this isn't easy there's definitely a problem.
It definitely makes sense that the faster the rate of change the more likely the engine has to combine results from multiple places until optimal placement on disc is found in steady state later, and hence query latency would rise.
You should probably try it too before blogging about it.