LLM Python/CLI tool adds support for embeddings
simonwillison.net
simonwillison.net
Don't miss the new llm-cluster plugin, which can both calculate clusters from embeddings and use another LLM call to generate a name for each cluster: https://github.com/simonw/llm-cluster
Example usage:
Fetch all issues, embed them and store the embeddings and content in SQLite:
paginate-json 'https://api.github.com/repos/simonw/llm/issues?state=all&filter=all' \
| jq '[.[] | {id: .id, title: .title}]' \
| llm embed-multi llm-issues - \
--database issues.db \
--model sentence-transformers/all-MiniLM-L6-v2 \
--store
Group those in 10 clusters and generate a summary for each one using a call to GPT-4: llm cluster llm-issues --database issues.db 10 --summary --model gpt-4If I want 100 close matches that match a filter, is it better to filter first then find vector similarity within that, or find 1000 similar vectors and then filter that subset?
... although on thinking about this more I realize that a better approach may well be to just filter down to ~1,000 and then run a brute-force score across all of them, rather than messing around with an index.
Can be simplified to:
jq 'map({id, title})'You almost certainly want a graph like structure (overlapping communities rather than clusters).
But unsupervised clustering was almost entirely ineffective for every use case I had :/
I mainly like it as another example of the kind of things you can use embeddings for.
My implementation is very naive - it's just this:
sklearn.cluster.MiniBatchKMeans(n_clusters=n, n_init="auto")
I imagine there are all kinds of improvements that could be made to this kind of thing.I'd love to understand if there's a good way to automatically pick an interesting number of clusters, as opposed to picking a number at the start.
https://github.com/simonw/llm-cluster/blob/main/llm_cluster....
Alternatively, there is a Bayesian GMM in sklearn. When you restrict it to diagonal Covariance matrices, you should be fine in high dimensions
One question, I saw a comment here that doing RAG efficiently entails some more trickery, like chunking the embeddings. In your experience, is stuff like that necessary, or do I pass the returned documents to GPT-4 and that's it for my RAG?
For an example of what I'm doing now, I bought an ESP32-Box (think basically an OSS Amazon Echo) and want to ask it questions about my (Markdown) notes. What would be the easiest way to do that?
The absolute easiest approach right now is to use Claude, since it has a 100,000 token limit - so you can stuff a ton of documentation into it at once and start asking questions.
Doing RAG with smaller models requires much more cleverness, which I'm only just starting to explore.
LLM works as a library already, but there's definitely room for improvement there:
https://llm.datasette.io/en/stable/python-api.html
https://llm.datasette.io/en/stable/embeddings/python-api.htm...
What I've come up with is either a) ask an LLM for the common label from samples from the grouped set after indexing (what keyterm best describes the relationship between these documents), or b) determine the label (or keyword) while indexing (by having the LLM find the keyterms ahead of time), then use set overlap on the grouped set's keyterms after to determine a label for the group.
I wonder how well it would work for building an "apropos" tool for searching man pages (and infotext)?
https://github.com/simonw/llm-gpt4all/blob/0046e2bf5d0a9c369...
https://github.com/simonw/llm-mlc/blob/b05eec9ba008e700ecc42...
https://github.com/simonw/llm-llama-cpp/blob/29ee8d239f5cfbf...
I'm not completely happy with this yet. Part of the problem is that different models on the same architecture may have completely different prompting styles.
I expect I'll eventually evolve the plugins to allow them to be configured in an easier and more flexible way. Ideally I'd like you to be able to run new models on existing architectures using an existing plugin.
Do I have to run it against my own corpus? Are there "standard" embeddings that many people use that I could use with llm?
They've embedded all papers on Arxiv and made the results available for anyone to use.
I believe they used the Instructor XL model - I need to build an LLM plugin for that.
It's a higher layer abstraction that runs on top of multiple libraries like HF transformers, in some case calling plugins that use transformers directly.
Python Library "llm" now provides tools for working with embeddings
I initially was trying to parse that, thinking "is this an open AI thing?". Of course the answer is just a click away, but people might miss this if they are interested in Python coding and AI.I put a bunch of work into getting it into Homebrew so that people who aren't Python developers can "brew install llm" and start using it.
Details on the CLI here: https://llm.datasette.io/en/stable/usage.html and https://llm.datasette.io/en/stable/embeddings/cli.html
The fact that it's written in Python is, I think, one of the least interesting aspects of the project.
It's a great name from a "glad you got there first" perspective but it's also so general as to be ungoogleable. Like if I want to find documentation for this library/CLI in Google, what would I search? I'd probably end up putting "Python" in the query just to disambiguate it from all the results about LLMs in general. So IMO you may as well include "py" (or some other disambiguator) in the name, since people are going to need to include it in their search queries anyway.
I'm taking a bit of a risk here, but I actually think there's a tiny chance that the acronym is still obscure enough that I'm in with a chance. I'm on the second page of Google already.