The steps would then be: 1. Embed your private data in chunks and store the resulting embeddings in a vector database 2. In your prompting workflow, when a user queries the chat model, embed their query using the embedding model 3. Retrieve the most similar chunks of text from your vector database based on cosine similarity 4. In the chat response, provide it the context of those chunks of text
For example, if you asked "who have I discussed Ubuntu with?", it might retrieve emails that have similar content. Then the model will be able to answer informed by that context.
Both have extensive examples in their documentation for almost identical use cases to the above.
The computing power we're requiring is simply what's available in any M1/M2 Mac, and the resource usage for the indexing and search is negligible. This isn't even a hard requirement, any modern PC could index all your emails and do the local hybrid search part.
Running the local LM is what requires more resources, but as this project shows it's absolutely possible.
Of course getting it to work *well* for certain use cases is still hard. Simply searching for close sections of papers and injecting them into the prompt as others have mentioned doesn't always provide enough context for the LM to give a good answer. Local LMs aren't great at reasoning over large amounts of data yet, but getting better every week so it's just a matter of time.
(If you're curious my email is in my profile)
How do you encode the private data into the vectors? It is a bunch of text but how do you choose the vector values in the first place? What software does that? Isn’t that basically an ML task with its own weights, that’s what classifiers do!
I was surprised everyone had been writing about that but neglecting to explain this piece. Like math textbooks that “leave it as an exercise to the reader”.
Claude with its 100k context window doesn’t need to do this vector encoding. Is there anything like that in open source AI at the moment ?
You use sentence transformers (https://www.sbert.net/).
You use a strong baseline like all-MiniLM-L6-v2. (Or you get more fancy with something from the Massive Text Embedding Benchmark, https://huggingface.co/spaces/mteb/leaderboard)
You break your text into sentences or paragraphs with no more than 512 tokens (according to the sentence transformers tokenizer).
You embedding all your texts and insert them into your vector DB.
Yes, turning words into vectors is it's own class of machine learning. You can learn a lot on the NLP course on hugging face https://huggingface.co/learn/nlp-course/chapter1/1 (and on youtube).
But even at 100K, you do eventually run out of context. You would with 1M tokens too. 100K tokens is the new 64K of RAM, you're going to end up wanting more.
So techniques like RAG that others have mentioned are necessary in the end at some point, at least with models that look like they do today.
How do you know that Claude doesn't do this? If you have multiple books, you end up with more than 100k context, and running the model with full context takes more time so it is more expensive as well.