Using Llamafiles for embeddings in local RAG applications
future.mozilla.org
future.mozilla.org
I know they add valuable input to it, but CC-BY-NC is really rubbing me the wrong way.
Whether weights can be copyrighted at all (which is the basis of these licenses) is also unclear. Though again I think they should be, they are just as creative a work as a computer program for any nontrivial model release (though really I think it's important to be able to enforce copyleft on them more than anything).
Also, all these laws work to the benefit of who ever has the deepest pockets, so salesforce will win against most others by virtue of this, regardless of how the law shakes out.
From what I can see, they both have the ability to archive / FTS your bookmarks.
But in terms of API access, historious only allows WRITE access (ugh), where at least pinboard allows read/write.
What else am I missing?
Also, don’t discount plain old BM25 and fastText. For many queries, keyword or bag-of-words based search works just as well as fancy 1536 dim vectors.
You can also do things like tokenize your text using the tokenizer that GPT-4 uses (via tiktoken for instance) and then index those tokens instead of words in BM25.
https://ollama.com models also works really well on most modern hardware
And as to sidestepping inference, I can totally do that. But I think it's so much better to be able to ask the LLM a question, run a vector similarity search to pull relevant content, and then have the LLM summarize this all in a way that answers my question.
Here is the list of technological problems:
1. When is a page ready to be indexed? Many websites are dynamic.
2. How to find the relevant content? (To avoid indexing noise)
3. How to keep an acceptable performance? Computing embeddings on each page is enough to transform a laptop into a small helicopter with its fans. (I used 384 as the embedding dimension. Below, too imprecise; above, too compute-heavy).
4. How to chunk a page? It is not enough to split the content into sentences. You must add context to them.
5. How to rank the results of a search? PageRank is not applicable here.