1. Install sentence-transformers [1]
2. Initialize the MiniLM model - `model = SentenceTransformer('all-MiniLM-L6-v2')`
3. Embed your corpus [2]
4. Embed your queries, then search the corpus
This runs on CPU (~750 sentences per second), and GPU (18k sentences per second). You can use paragraphs instead of sentences if you need more text. The embeddings are accurate [3] and only 384 dimensions, so they're space-efficient [4].Here's how to handle persistence. I recommend starting with the simplest strategy, and only getting more complex if you need higher performance:
- Just save the embedding tensors to disk, and load them if you need them later.
- Use Faiss to store the embeddings (it will use an index to retrieve them faster) [5]
- Use pgvector, an extension for postgres that stores embeddings
- If you really need it, use something like qdrant/weaviate/pinecone, etc.
This setup is much simpler and cheaper than using a ton of cloud services to do embeddings. I don't know why people make semantic search so complex.I've used it for https://www.endless.academy, and https://www.dataquest.io and it's worked well in production.
[2] https://www.sbert.net/examples/applications/semantic-search/...
[3] https://huggingface.co/blog/mteb
[4] https://medium.com/@nils_reimers/openai-gpt-3-text-embedding...