Milvus – An Open-Source Vector Similarity Search Engine
milvus.io
milvus.io
(you can see the VERY "research quality" code on Github, here's a decent starting place https://github.com/hyperstudio/spectacles/blob/master/specta...)
EDIT: from the insertion docs https://milvus.io/docs/guides/milvus_operation.md#Insert-vec... it seems that they still ask you to re-build your indices after you insert vectors, although in some cases they can tell that they need to re-build the indices for you. Looks like the major value adds here are potentially shifting computation to the GPU and building multiple indices. I'll certainly evaluate this next time I'm building a project around vector search.
We are now working on the vector deletion. Hopefully will be ready by the end of 1Q this year.
EDIT: from reading the linked article, it seems like newly inserted vectors will be queried using brute force. Very interesting idea!
Tests show that FAISS is bit better than annoy on retrieval of both small and million items indexes. It also includes ind x compression techniques that in our tests do fair very well, with very low loss on mid size 500k image indexes.
Again, hopefully be ready by the end of 1Q this year.
[1] https://gnes.ai/
It explains how Milvus managing vectors.
Why are you not quantizing the vectors when you insert them? Bolt [1] and Quicker-ADC [2] make 10-100x compression basically free for approximate search (and also get you ~100x compression roughly 10x faster querying within a partition....)
[1] https://github.com/dblalock/bolt
[2] https://github.com/technicolor-research/faiss-quickeradc
Based on our users experience, SQ8 is the most balanced one at this moment. SQ8 provides 8x compression, higher accuracy and better performance.
Most people told us running Milvus on arm looked cool but they were not sure if they want to do this...
Please tell us your requirements and scenarios on arm. It will really help.
It's about the ML scenarios. If you want to search thru a huge amount of unstructured data after vectorization tech (like deep learning), Milvus will help you a lot.
Our users use Milvus in below scenarios: 1. Chemical molecules analysis, searching SMILE format vectors 2. Image retrieval type application, for example shopping website 3. NLP 4. Recommendation system 5. and more, we are collecting users' feedback
At this moment, the IVF indecies are based on FAISS. So the performance is the same as Faiss.
IVF_SQ8H is the reconstruction from Faiss IVF SQ8. Performance is much better, but you need GPU for it.
We provide benchmark test procedures and tools.
Please check this: https://github.com/milvus-io/bootcamp/tree/master/EN_benchma...
1) Most of these datasets have extremely correlated dimensions. If you plot the covariance matrices, you'll see dense blobs of entries close to 1 all over the place. This makes the ANN task much easier than it would be with, say, high-quality DNN features. As an example, I've compressed MNIST digits down to 1 byte representations with vector quantization and still gotten nearly perfect retrieval accuracy.
2) 1M vectors is not that many. You can get easily get 1k queries per second in a single thread at a decent precision/recall just brute-force scanning through them with a SIMD approximate distance function like Bolt or Quicker ADC [1]. Also worth noting that the FAISS paper (along with a lot of other work since then) focuses mostly on 100M to billions of vectors.
3) Related to (2), I think most of these methods aren't incorporating state-of-the-art approximate distance functions yet (though I haven't dug into all of their source code). AFAICT FAISS+Quicker ADC [2] is the actual leader on x86 CPUS. Can't comment on the production-readiness of their code though.
[1] The latter is a bit faster for ANN search, though the code is more complex IIRC.
[2] https://github.com/technicolor-research/faiss-quickeradc
I think the Ann benchmark should pay more attention on
1. The index building speed, as this is very important in some production scenarios. Now it only says I will give 5 hours to build the index on that 1 million vectors.
2. The memory footprint, as 1m vectors are not that many. We will have to deal with billion s of vectors for chemical molecules, images and word vectors. The memory consumption will definitely impact how many servers you need.