Vald: A Highly Scalable Distributed Vector Search Engine
github.com
github.com
- Native, first class support for dense vectors first, no need to hack around the inverted index like this talk by Simon Hughes: https://haystackconf.com/2019/vectors/
- Using native programming languages, avoiding the JVM and all the memory and GC overhead
- Hopefully native languages might make it easier to deploy code that uses GPUs!
- A chance to focus on info retrieval when current Lucene solutions are either a bit crufty (Solr) or very focused on logs/analytics (Elasticsearch)
- a chance to have one system that unifies search and recommendations in one backend, not need to split them between two systems
- a chance to rethink the language used to specify query ranking factors. Solr has a very unique DSL that’s very write-only. Elasticsearch is very verbose to do anything non trivial.
Still I do wonder even if Lucene catches up a little, it’s current mindshare, reliability and install base will continue to dominate. For example see Lucene 9004 for adding dense vector search: https://issues.apache.org/jira/plugins/servlet/mobile#issue/...
The holy grail is a system that does a good job at classic full text search AND dense vector search. I wonder what system has the best chance of delivering that.
(context: I'm maintaining a recipe search engine, and it strikes me that recipes could be considered as part of a vector space, but I'm wary of adopting new technologies with something that works 'well enough' for now (Elasticsearch, in this case))
PS: I work at Pinecone and you can see if any of the examples here help you : https://www.pinecone.io/docs/examples/
The examples are helpful and well documented, thank you. I'd mention that I'm constraining the technology options to self-hostable and FLOSS software, and with a strong bias towards concepts that are already well-understood by most software engineers, but I'll keep a look out for Pinecone.
I assume that Vlad takes care of the reindexing in the background, which would save a lot of the work.
In Weaviate, any imported vector is immediately searchable, you can update and delete your objects or the vectors attached to the objects and all results are immediately reflected. In addition every write is written to a Write-Ahead-Log, so that writes are persisted, even if the app crashes.
We wanted to make sure that Weaviate really combines the advantages of AI-based search with the comfort and guarantees you would expect from an "old school" database or search engine.
- https://github.com/spotify/annoy
- https://ai.googleblog.com/2020/07/announcing-scann-efficient...
I did something naively like this using array indexes of about length 20 vectors on Cloudant (CouchDB) in order to implement a text similarity search, based on something simple like LDA auto tagging and keyword frequency clustering vectors.
I would query based on input text vectorized the same way, and simply offset each element of that input vector by a fixed small value, using those offsets as start and end keys for a couch view query.
I didn’t test this on larger vectors but have always been curious on how far it could go.
Probably the matching algorithms in this project are more accurate, but this naive method worked pretty well for my use case and on a maintenance free DBaaS.
Would love to hear more about concrete usage examples in this area if anybody likes to share - thanks
Edit - found my slightly more detailed blog post about this http://splatcollision.com/page/fast-vector-similarity-querie...
There are also plans to further optimize this based on how restrictive the boolean filter is. For example, if the filter still matches a lot of of ids (e.g. 50% of all candidates) the approach outlined above works really well. But if if your filter is extremely restrictive and maybe only matches 1k out of 100M results, then this approach is no longer ideal, as the vector index will discard most of the ids it sees. However, in this case it would be more efficient to actually perform a brute-force search on those 1k vectors. This and other optimizations around combined vector and scalar search are on the roadmap.
This thread has been a wealth of information. I’ve had great success with Annoy so far; eager to kick the tires on Vald.
for Kubernetes https://vald.vdaas.org/docs/tutorial/get-started/ for Docker https://vald.vdaas.org/docs/tutorial/agent-on-docker/
enjoy!!!
Indexing them is really hard: The local neighborhood's volume to be searched grows by radius^dimension; you need a specialized engine.
Now, say you are a three letter agency and have Facebook's image data (2B faces, 100B images). You train a deep learning image classification network to identify faces. But for training time reasons you only can do it on a dataset of about 10k different faces, 10M images. The NN will identify characteristic visual features to make its classification, available at the n-1 layer as activations.
But you need to have the rest of the dataset searchable! You now index the activations of all the 100B remaining images. And with this you can search by similarity of visual features. If you have an image of someone you want to to track, get its activations and search for other images that have similar vectors.
This works for everything that a NN can model: speech, words, words in sentences, videos, etc.
----
On a another note Facebook has a feature-rich, GPU-powered, scalable indexation engine called FAISS:
ML-models output vectors, you can use a vector search engine to store those vectors and quickly search through them. Weaviate also stores the data objects btw like a DB: https://db-engines.com/en/blog_post/87