As the traditional db move forward in the space the need for dedicated vector databases will likely shrink, except for some very specific implementation that offer unique enough features (I.e. deeplake does vector search over object storage, which is very convenient for certain specific scenarios)
Even if you could make it perform well, it would not do what you want.
https://github.com/asg017/sqlite-vss
Not associated with the project, just love SQLite and find it very useful.
The benefit is that you don't have to pay for the compute part of a database, and the storage layer is as cheap as it could be on the cloud.
retrieve the embedding index and to run an indexed search to identify the data to be retrieved. Please bear with the layman like questioning -
So if the data is {obj: "obj1, "data": {"name": "atlas", "embedding": "1123124234" } What is an embedding index ? Is it something like {"1123124234": "obj1"} ?
From what I understand the query will be "geography" whose embedding will be "12311111" and now you have to run a KNN for a match which will return {"name": "atlas", "embedding": "1123124234"}
Not sure where the embedding index comes into play here.
Now if you want to do similarity search you have to measure the distance between 2 or more vectors and that's independent of the indexing. No ?
So any database with sufficient memory should be able to accomplish this as evidenced by the vector similarity search feature of Redis. ( I don't know how Redis folks have implemented vector similarity but they do support KNN search )
There are alternative index types, of course, or you could index the hash of the vector. These both come with tradeoffs.
Aside, you can increase the size of tuples you can index in a postgresql btree by increasing the postgresql page size (requires a recompile and creating a new database instance).
Of course ACID scales to well into the Fortune 500 scale so...
They mostly seem a tarted up associative array, sure, but a key-value store is a thing.
DynamoDB underpins much of AWS which in turn underpins a ridiculous number of web services.
So definitely more than just a thing.
NoSQL has been around for over 20+ years.
Since then Cassandra, DynamoDB, FoundationDB, MongoDB, Neo4J, Redis etc are not only still around but widely used and powering many of the services you use today.
Even looking at HN's reactions to that video show a few comments that did not age very well (although most of them did): https://news.ycombinator.com/item?id=1636198
The sensible takes were that NoSQL would supplement and enhance RDBMSs, but the hype was much more than that.
RDBMS is such a mature and powerful technology, and with the vast power in modern single-node hardware, they will scale to considerable sizes. But there is a limit.
Once a table or set of tables hit a certain scale, it needs to be distributed. Once you get to those scales, you are likely dealing with the threat of exponential data growth, so doing the "distributed Postgres" will bandaid the problem ... but you are starting to run into the CAP theorem's problems.
You'll need to AP-scale the biggest data tables, and since that almost always means a big coding change lift and introduction of an AP-scaling database, it will almost always be a six months to a year transition, and possibly adding an entirely new DB technology (Cassandra / DynamoDB / maybe FoundationDB).
I have yet to have anyone explain how joins scale on an AP distributed database except in limited situations where the joined data is somehow node-local to the other tables, usually some hierarchical situation. Otherwise you are pulling data from lots of nodes and aggregating and comparing the different sets to account for node drift / partitions / network failures. Cassandra and Dynamo basically say "you are scaling a single table/query/update/pre-joined data table".
Which really isn't fun for RDBMS folks. Because it is a shitload of denormalization on top of all the AP headaches and distributed transactions / updates.
As I said, megahuge machines really make that an outside case unless your use case really emphasizes the "A" in CAP, where write speed can be tolerated to be low and you want to tolerate entire cloud or datacenter outages.
But yeah, nosql was never going to kill the rdbms. The functionality/power of rdbms/SQL is so so so so much higher. Just be aware when you're about to shoot over the limit and prepare for the code changes to handle true scale in the (unlikely) event you're going to need it.
Massive RDBMS instances are awful to manage. Hard to backup/restore/fork. Hard to migrate tables and schemas without causing downtime. Accidental downtime happens all the time due to locks, bad indexes, bad query plans , etc etc. at large scale they are capricious beasts, care and feeding and most importantly changing them becomes a dark art.
Don’t let them get too big you’ll regret it :)