127 karma · joined May 9, 2019
https://github.com/raphaelsty/knowledge
I did not find yet a solution that suit me for knowledge which is not an online webpage / pdf.
Documents and queries embeddings can be obtained using .encode_documents and .encode_queries methods
I save most of my embeddings (python dictionnary with documents id as key and embeddings as values) using joblib in a Bucket in the cloud. I don't really know if it's a good pratice but it does scale fine to few millions documents for offline (no real-time) applications.
I actualy use it at least twice a week to retrieve content I bookmarked, so I'm happy to have created such a tool.
The app: https://raphaelsty.github.io/knowledge/?query=bayesian
The Github: https://github.com/raphaelsty/knowledge
Search implements a wrapper of the Python ElasticSearch client that is scalable and dedicated to corpora composed of tens of millions of documents.
1) The dependency on the Elasticsearch python client allows Elasticsearch to be used as a retriever. The same goes for Lunr. It might be interesting to separate the different dependencies.
2) Of course I'll update it.
I just published a library dedicated to knowledges graphs embeddings. The Mkb API is inspired by Scikit Learn. I provide modular tools for building latent graph representations.