TL;DR - They store vector embeddings and index them for efficient nearest-neighbor retrieval. Some also allow CRUD operations, metadata filtering, fault-tolerance, and all that jazz. Useful for things like semantic search, similarity search (of unstructured data), and recommender systems.
I am cofounder of ZIR AI (https://zir-ai.com/). I researched neural information retrieval at Google, before starting ZIR in 2020. (Note: Vespa, who appear in your article, reference some of my work in [1])
To give you some historical perspective, embedding based retrieval on large text corpora became viable only after the introduction of transformers in 2017. Google Talk to Books (https://books.google.com/talktobooks/) is the first such system I'm aware of --- I designed the neural network that powers that system.
It's tricky to compare with BM25 on extremely large datasets (like the BioASQ challenge), because the ground truth is collected using keyword search, hence a chicken-and-egg problem.
However, it's been shown in recent research (post 2018) that neural retrieval can outperform BM25, as measured by mean average precision (MAP), by large margins in the following cases:
1. On natural language queries (e.g. longer than 5 words; spoken queries, etc.) 2. On queries, where the seeker doesn't know the answer in advance (information seeking, rather than reference lookup). 3. When searching user generated content. 4. On smaller corpora (non-web) sized.
As a side note, our platform might fit well into your comparisons, although we provide a more comprehensive solution for text indexing, specifically. I consider platforms like Amazon Kendra and Microsoft Semantic Search to be stronger contenders in our space.
Please correct me if I am wrong, with genetic algorithm, DL, etc. The vector embedding happen on the software layer (inside the network). I imagine it would be somewhat tricky to separate the vector embedding and put it inside a vector database. It could also be a performance hit during the training process.
An example in NLP world is BERT-like NNs, that allow you to embed your text into a dense vector representation.
I might say transformer-based NNs instead. The problem with cross-attentional models like BERT is that they won't scale to large datasets. They are more often used in reranking results within an IR pipeline. However, even for that use they require distillation.
Genetic algorithm usually deals with chromosomes. I have used binary and hexadecimal chromosomes in the past and found that the binary chromosomes are more flexible, especially with bitwise operations.
Let's say a chromosome with four chunks of 4-digit binary, with each digit as dimension of true/false value; we end up with something like 1011 1100 0010 0101. Then I stored these four chunks in four documents in a NoSQL database. Each document then also has the records of other 16-digit chromosomes, so that I can refer to those 16-digit chromosomes contain an exact/similar 4-digit chunk. This was the fastest method I could come up with the last time I worked on it; I am sure that there are more efficient methods out there.
Hopefully, this can shed some light on how the genetic algorithm works.
In a data scientist’s perspective, this may seem to be a hack. I would be very appreciated to learn more from you about this. I can tell you that the weakness on this method is the multiple writes to the database. I assume vector database can gets this implemented with less writes.
In the past, I have used this for supervised training and yielded very good results. However, I think this would be inefficient in large scale networks. I am planning to use Go + Clickhouse to improve the performance in the next project.
You have 4 out of 6 open source DBs, and 2 commercial ones give you the managed service. All 6 can scale to quite large numbers of vectors.
My goal is to systematically study each DB through the lens of a specific search task, which will not be (only) text based.
If you think we could collaborate in some way on your dataset, that would be fantastic and probably a learning experience to both sides.