Show HN:Wikipedia vector search. 36M passages embeddings in just 2.54 GB
speech-kws.ozonetel.com
speech-kws.ozonetel.com
https://speech-kws.ozonetel.com/wiki
Thanks to @CohereAI for releasing the Wikipedia embedding dataset. I saw that for 36 million passages the embedding size would be around 120 GB. So if I had to host the embeddings and enable neural search on this dataset I will have to do LLM ops and runs vectorDB clusters. This was a good dataset to test out the efficacy of the Alpes KE Sieve algorithm. We built an embedding space based on @huggingface all-mpnet-base-v2 model. We created two models, one with 540 dimension bit embedding(small) and another with 2200 dimension bit embedding(large). We were able to embed all the 36 million passages in 2GB and 10GB respectively.So you can now have local wikipedia vector search locally. Since the size is so small, we just used np.array and no vector databases. You can test it out in the link above and share your feedback.
Interesting the small model generates more unique results than the large where I just get two very close hits but they are repeated. Search : "Memory Palace and physical spatial metaphors for 3D information organization in history"