Do we really need a specialized vector database?
modelz.ai
modelz.ai
It is a bit as if Postgres has become what Oracle wanted to be - a one-stop solution for all your data storage needs.
The need for specialized time series databases is explained by the basis of any software development process: reuse ready-made high-level abstractions instead of having to reinvent the wheel, poorly.
There are a few key features of time series databases that you either go without (i.e., data compression) or have to reinvent the wheel (querying, resolution-based retention periods, statistical summary of time periods,etc).
You might live without those features,but in some applications there's a hefty price tag tied to not having them. If you adopt a time series database, you benefit from all those features without having to implement a single one of them.
Perhaps that has no value to you, but it does have for the world.
https://www.timescale.com/blog/timescaledb-2-3-improving-col...
https://www.timescale.com/blog/massive-scale-for-time-series...
Their Slack is very responsive too, top service.
Time series databases should be viewed in the context of being a fairly niche product however. The archetypal use case is in finance where you have a constant firehose of tick data and trades across a wide array of instruments you need to be able to index instantaneously and aggregate easily. The sheer volume of data can't be overstated here. To make matters worse, the records are often both wide and sparsely populated.
A non-specialized really DBMS doesn't cut it in this case. You need to build the DBMS from the ground up to cater to this sort of a usecase. Slapping a columnar backing store onto something existing does not cut it.
A columnar database engineered from the ground up for time-series should have better foundations to allows fast ingest (to the tune of million of rows/sec per server) and also be very efficient for time-based queries that can be done via languages such as SQL.
Using QuestDB as an example: Data is stored in chronological order and is optimized for sequential ingestion with a timestamp component (re-ordering data on the fly if it comes out-of-order). The data is also partitioned by time. The InfluxDB Line Protocol is better suited for streaming type of ingest versus transactional inserts via Postgres.
Application is storage of a diverse set of (irregular) events. Lifecycle management and archival is more important than compression. Volume is moderate.
In the end it depends how much data you're dealing with. If you're only ingesting dozens to hundreds of records per second per writer you might actually like AWS Time stream. If you're dealing with a million + rows per second from a low number of writers (e.g. scientific data) then you probably want Influx or Timescale. If you want to host it yourself and you have less than 1TB of data then Influx is okay or if you want it cloud hosted then Influx may be a good option.
If you need full control, scalable, TB or PB of data, etc. Then Timescale is great but it's definitely more work.
It might sound paradoxical, but specialized time-series databases, such as TimescaleDB or InfluxDB, are not good at time-series, compared to normal (boring) OLAP databases.
If it weren't coming at the cost of everyone else in the industry, i wouldn't mind this preference, not everyone has to be a hacker. As is, apple's existence is a net negative.
I real life, though, having a proper Linux kernel at hand is useful.
I've used windows machines for years doing the same thing elsewhere and I had none of these problems. I run cygwin,wsl or a Linux VM and I had everything I wanted locally run on a Windows laptop.
Now having a Linux laptop (based on rhel9) many things are really nice. For example: I cun run gui apps via x11 over LAN on my desktop, I can share many dotfiles from my desktop to have the laptop configured exactly how I like it. But there is a price to pay. Unsurprisingly having to use outlook and ms teams via a browser is a pain.
Many linux users seem very bitter that people value Apple's offerings over linux based solutions. That this preference is somehow taking something away from them. I don't think anyone should be looked down upon because they don't value the same things that someone with a hacker/tinkerer mentality values. I think that attitude is partly what puts a lot of people off who are half in/half out of the ecosystem. I personally find it a bit tiring, and in general I'm supportive of open source and related things.
What is common though is having to use tools you are not used to or tools that are subpar e.g. a Mac. You as a Mac user might call that being bitter. It is like forcing an Excel jockey to only use OpenOffice. Those people turn bitter fast.
If those same mac-using developers were allowed to use Asahi Linux (or one of the other up-and-coming ones like Fedora on Apple Silicon), it'd be interesting to see how many make that choice.
Either way, it's usually much easier to negotiate hardware in smaller companies: they don't have stock, hardware policies, security software to support, b2b contracts to handle.
It's the bigger guys that want unification.
at least until wasm gets more and more traction
Which was sort of fine in a world where nobody cared about data security. But it’s almost impossible to build a “monolith” database that lives up to the modern legislation requirements, at least within the EU. So we’re basically living in a world where the way we’ve done databases for the past 50 years, doesn’t work for the business. Unless you separate everything into business related domains in small kingdoms with total sovereignty over who comes and goes. Which is basically what micro services are.
At the same time we’re also living in a world where it’s often cheaper (and performant enough) to put 95% of what you do in a container using volumes for persistence. If SQLite was ready for it, or maybe, if all the ORMs were ready for SQLite, you could probably do 90% or your database needs in it. Since that’s not the case, Postgres reigns supreme because it’s the next best thing (yes, you’re allowed to think it’s better). I view Vector DBs sort of similarity to this. We operate almost all of our databases in containers, the one exception is for the one services everyone calls where the performance loss of the container is too much. Well, that’s not entirely true, our dataware house isn’t operated in containers, at least not by us, but I rarely interact with the BI side of things except for data architecture so I’m actually not sure how their external partners do Ops. Anyway, I think Postgres will do 90% or (y)our vector needs and then we will need dedicated vector databases for those last bits that won’t fit into the “everything” box. I don’t think this is bad either, maybe Oracle wanted it to be theirs, but then Oracle should not have Oracles governance. I’m sure they didn’t foresee this future when they bought MySQL though, I know I didn’t. Because who could’ve foreseen that that way we used to do Databases would become sort of obsolete because of containers and EU legislation… and other things?
Not sure I agree. It seems far simpler to comply with most regulations if you store your data in a centralised data store with good management features.
If each and every microservice handles its own data, how would you implement rules affecting all of them, such as retention periods, deletion requests, permissions, access logs, etc? It's not just that the data is stored all over the place. You may also have to use microservice specific APIs to access it.
If you're referring to laws that require data storage in a particular jurisdiction, it would still be easier to centralise data management in each jursdiction rather than distribute it across scores of services within each of those jurisdictions.
I don't think we disagree, it's just that we soon won't be allowed to physically store our data on the same servers because the EU deems it too risky. A couple of years ago you could operate 600 energy plants from the same few servers, but by January 2024 that will no longer be legal.
Aside from that, we're now entering a world where you need to do comply with regular audits on data field levels as well as an explosion of access roles. I'm not sure how you would manage that in a traditional centralized data store, you probably could, but we do it with OPA and once you do that, there is no difference between using a monolith or micro services because it all goes through the central policies.
You could argue that the energy sector is a bit of a special case, but it's even sticter in anything related to the financial sectors.
I agree, because we never left it. People who say otherwise have an extremely distorted view. Most companies are not Google or Meta.
This is an extremely narrow view. There are a zillion uses of databases that are not storing user data from web apps or whatever.
Who said anything about "user" data?
I work in the energy sector. The GDPR rules are tiny issues compared to the requirements we face to be able to submit ourselves to audits of exactly whom accessed exactly what. We're at the point where the role you get when you log on a system, needs to define whether you have access to each column in a row or not, and we need to log it too.
For authorization, policies and logging we have frameworks like OPA, but for the actual data protection... Well...
I do some “analytics” in a postgres database, with data and metrics stored in the same db, but I guess it’s not a huge amount - amount of data swamps the metrics by orders of magnitude. Seems to work ok for me, queries are 5ms or less, and it is only one thing to learn/deploy/maintain.
It’s on a pretty powerful server, though.
Disk I/O. Most analytical workloads only need a subset of columns for the tables under query, but Postgres has to read in every column due to how data is organized. Columnar and hybrid columnar can significantly reduce disk I/O by orders of magnitude, which makes a big difference beyond a few dozen gigabytes.
Even if all your data can fit into memory, Postgres will still be slower because it needs to loop over values for aggregations. Columnar databases can use SIMD instructions and better utilize the CPU cache across multiple aggregations in the same query.
Postgres is also built for lookup queries and point updates, so every row needs a separate entry in b-tree indexes. Analytical databases can generally use sparse indexes that resolve to groups of rows, which means indexes take up both less space in memory and on disk that make range queries more efficient.
When I've used Postgres JSON fields in the past, I've used it for unstructured data that doesn't need to be indexed, as if there's anything that does need to be indexed, then it's important enough to extract out to a field for efficiency.
It was Google's BigTable paper (2006) and Amazon's Dynamo paper (2007) that led to the "NoSQL revolution" in the late 2000s. The goal was to make it easier to scale a lot of de-normalized data horizontally with the tradeoffs being lack of RDBMS-like querying and schema enforcement.
In hindsight it was bizarre to market a whole category of horizontally scalable data systems based on one of their main negative trade-offs: Lack of SQL querying.
Since then many of these data systems have added SQL-like querying.
C certainly has its faults, and while I have no real experience with Rust, I'm willing to believe that it's significantly better as a language.
But pgvector, at least from a quick scan, looks like a well-written, easily comprehensible C codebase with decent tests. There are undoubtedly lots of hard problems that the developers have solved in implementing it. Putting time and effort into reimplementing all of that in another language because of an aversion to C feels like a waste of effort that could be put into enhancing the existing extension.
Maybe there's something I'm missing? Is the C implementation not as solid as it looks at first glance?
Do we know that tokio's concurrency strategy is optimal for database access?
From a consultant perspective, Postgres as a requirement is a much easier sell then some new, hot but unknown dedicated vector DB.
here are a few benchmarks from this week's commit: https://jkatz05.com/post/postgres/pgvector-hnsw-performance/
What we won't be doing is figuring out how to scale to petabytes of data distributed across multiple data centers in a massive active-active cluster. We've spent the last 14 years perfecting that, and still have work to do. With the benefit of hindsight, if you have a database that is less than 10 years old, all I have to say is good luck. You have some challenging days ahead.
1. https://cwiki.apache.org/confluence/display/CASSANDRA/CEP-30...
If it takes quantization + buying a 1 TB RAM server ($4k of RAM + parts), do that in memory with the raw tensors and shed a small tear -- both for cost and the joy of the pain that you are saving yourself, your team, and everyone around you.
If you need more, then tread lightly and extremely carefully. Very few mega LLM pretraining datasets are even on this order of magnitude, though some are a few terabytes IIRC. If you are exceeding this, then your business usecase is likely specialized indeed.
This message brought to you by the "cost reduction by not adding dumb complexity" group. I try to maintain a reputation for aggressively fighting unnecessary complexity, which is the true cost measure IMO of any system.
For example what is the difference between a "chunk vector" and a "prompt vector"? Aren't they essentially the same thing (a vector representation of text)?
How do we "search the prompt vector for the most similar chunk vector"? A short code snippet is shown that queries the DB, but it's not shown what is done with what comes back. What format is the output in?
I suspect this works by essentially replacing chunks of input text by shorter chunks of "roughly equivalent" text found in the vector DB and sending that as the prompt instead, but based on this description I can't be sure.
For example, let’s say you have a prompt that lets you answer questions about a book. If the book is long enough, you won’t be able to include it as is in the prompt, so you have to figure out what are the most relevant passages you must include to answer a given question. What you usually do is find the passages that are the most semantically similar to your question.
Chunk vectors are the vectorized passages of the book (i.e., a numerical vector that represents a passage), and the prompt vector is usually the vectorized question.
To find the most similar vectors you need a distance measure, cosine similarity being the most popular.
The output of finding the most similar vectors is the vectors + it’s metadata (chunk, page, chapter, etc)
I think I understand how it works now for many kinds of prompts where the information to be extracted is contained in one(or more) of the chunks of much bigger whole.
I'm not sure about prompts where it is required to "understand" entire input to answer properly. For example summarising a book. Although even with this vector search could perhaps help by looking for things not near "please provide a summary", but certain hand crafted values such as "important to the plot" etc.
I guess I need to do some experimenting with It. I found some open source alternatives to (not at all)OpenAI in form of "Sentence Transformers" to create embeddings.
However, what would be really neat is to have a large open source dataset of embeddings already created on some general purpose collection of texts, to try searches etc.
There are also open datasets of this, eg https://huggingface.co/datasets/kannada_news for news, or https://sites.google.com/eng.ucsd.edu/ucsdbookgraph/reviews?...
Also worth mentioning that vector databases are really useful in Resource Augmented Generation - aka find an answer from an existing corpus and utilize it to supplement the LLM.
Does this mean that a question is changed to a similar question? Doesnt it decrease the quality of answers significantly?
If someone asks "who is the best student in California?" will the question be changed to "what is the best school in California?".
This would explain the terrible drops in quality of amswers that we saw. The underlying technology is changed for easier scaling, but is much worse.
It's like the current google (what has multiple problems), where it also sometimes doesnt search your keywords - even when they are in quotation marks - because it knows "better"..
No. It is retrieving the most semantically similar document(s) to work with. For a question about students information about students will be more semantically similar than information about schools.
>It's like the current google
Google has used vector search for years. It's why you can just type questions in and Google will understand what you are talking about.
>doesnt search your keywords - even when they are in quotation marks
Quotation marks work. If there are no results it will show results without the quotes. The mobile site doesn't make this clear when it happens though.
https://neon.tech/blog/pg-embedding-extension-for-vector-sea...
txtai (https://github.com/neuml/txtai) sets out to be an all-in-one embeddings database. This is more than just being a vector database with semantic search. It can embed text into vectors, run LLM workflows, has components for sparse/keyword indexing and graph based search. It also has a relational layer built-in for metadata filtering.
txtai currently supports SQLite/DuckDB for relational data but can be extended. For example, relational data could be stored in Postgres, sparse/dense vectors in Elasticsearch/Opensearch and graph data in Neo4j.
I believe modular solutions like this where internal components can be swapped in and out are the best option but given I'm the author of txtai, I'm a bit biased. This setup enables the scaling and reliability of existing solutions balanced with someone being able get started quickly with a POC to evaluate the use case.
There are no 3rd party benchmarks for txtai as of now. You would have to compare how it does on your own data to judge.
Is that right? That sounds like, 2 OOMs higher than I would've expected. Is he doing something wrong like loading the model from a cold start?
That is a fine argument if you don't mind that pgvector is second-to-worst amongst all open-source vector search implementations, and two orders of magnitude slower than the state of the art [1].
The author also makes the argument that traditional DBs are better because they are battle-tested, and then goes and rewrites the pgvector plugin from C to rust.
That's why I think that outside some very specific use cases vector databases are not very useful.
Edit:
Ah I see pgvector will soon support HNSW as well (from 0.5.0): https://github.com/pgvector/pgvector/issues/181#issuecomment...
Lack of consistency is not a big deal and you often don't have it at scale anyways.