HNHacker News
TopNewBestAskShowJobs

gujun720

5 karma · joined October 21, 2019

submissionscomments
gujun720··on [dead]
I'd like to hear any feedback from who has tried this out.
gujun720··on MindSpore WebAssembly Backend
It will be helpful, if the author could compare the performance of WASM back end.

Will it be better than Model runtime with HTTP support?

gujun720··on Milvus – An Open-Source Vector Similarity Search Engine
Good points. I want to add one more.

I think the Ann benchmark should pay more attention on

1. The index building speed, as this is very important in some production scenarios. Now it only says I will give 5 hours to build the index on that 1 million vectors.

2. The memory footprint, as 1m vectors are not that many. We will have to deal with billion s of vectors for chemical molecules, images and word vectors. The memory consumption will definitely impact how many servers you need.

gujun720··on Milvus – An Open-Source Vector Similarity Search Engine
It's not about the ML platform.

It's about the ML scenarios. If you want to search thru a huge amount of unstructured data after vectorization tech (like deep learning), Milvus will help you a lot.

Our users use Milvus in below scenarios: 1. Chemical molecules analysis, searching SMILE format vectors 2. Image retrieval type application, for example shopping website 3. NLP 4. Recommendation system 5. and more, we are collecting users' feedback

gujun720··on Milvus – An Open-Source Vector Similarity Search Engine
200 GB is the size of original vectors. When creating index, Milvus supports IVF SQ8 and IVF PQ ADC.

Based on our users experience, SQ8 is the most balanced one at this moment. SQ8 provides 8x compression, higher accuracy and better performance.

gujun720··on Milvus – An Open-Source Vector Similarity Search Engine
Milvus could run on arm CPU. We ported it to Nvidia Jetson NANO and Raspberry PI 4 (4GB mem) so far.

Most people told us running Milvus on arm looked cool but they were not sure if they want to do this...

Please tell us your requirements and scenarios on arm. It will really help.

gujun720··on Milvus – An Open-Source Vector Similarity Search Engine
Correct, new vectors will first be searched thru brute force until the index is created on that file slice.
gujun720··on Milvus – An Open-Source Vector Similarity Search Engine
Please check https://medium.com/@milvusio/managing-data-in-massive-scale-...

It explains how Milvus managing vectors.

gujun720··on Milvus – An Open-Source Vector Similarity Search Engine
We are working on this feature which allows use to perform hyper search (attributes plus feature vectors). And you can code your scoring rules.

Again, hopefully be ready by the end of 1Q this year.

gujun720··on Milvus – An Open-Source Vector Similarity Search Engine
We have some test reports in https://github.com/milvus-io/milvus/tree/master/tests

At this moment, the IVF indecies are based on FAISS. So the performance is the same as Faiss.

IVF_SQ8H is the reconstruction from Faiss IVF SQ8. Performance is much better, but you need GPU for it.

We provide benchmark test procedures and tools.

Please check this: https://github.com/milvus-io/bootcamp/tree/master/EN_benchma...

gujun720··on Milvus – An Open-Source Vector Similarity Search Engine
You may check our Medium site. We will post more tech details.

https://medium.com/@milvusio

gujun720··on Milvus – An Open-Source Vector Similarity Search Engine
Milvus allows users to append vectors. Vectors are stored in multiple file slices. When a file slice reaches the threshold, Milvus will build the index for that file slice, and new data will be inserted into a new file slice. For details, please refer https://medium.com/@milvusio/managing-data-in-massive-scale-...

We are now working on the vector deletion. Hopefully will be ready by the end of 1Q this year.