Language models on the command line
simonwillison.net
simonwillison.net
What other CLI tools are people using to work with LLMs in the terminal?
There one comment here about https://github.com/paul-gauthier/aider and Ollama is probably the most widely used CLI tool at the moment: https://github.com/ollama/ollama/blob/main/README.md#quickst...
I use aichat: https://github.com/sigoden/aichat
I especially like the terminal integration where I can type a natural language request at the terminal and press Alt+E to have it converted to a command to run.
It is used as Bash widget, where you press CTL-F, describe your command, then press Enter, the generated command will be inserted into your shell prompt.
I need to integrate distil whisper large v3, aider, and shell_gpt to tidy up a lot of my disjointed LLM use. As someone else mentioned, the commits created by aider allow me to "skip" some intermediate steps that would be required when working on coding tasks with other frameworks.
$ openai api chat.completions.create -g 'user' 'say hello' -m 'gpt-3.5-turbo'
$ Hello! How can I assist you today?I looked at llm but it doesn't appear to have a mechanism for multi-shot prompting, where you provide both prompts and responses within your query. (Ref https://platform.openai.com/docs/guides/prompt-engineering/t... .) Maybe take this as a feature request?
It feels like the 'template' system in llm might be able to encompass this but the docs don't appear to provide a reference for the yaml format, only examples. I guess that is another feature request, sorry!
(BTW if you haven't seen https://docs.divio.com/documentation-system/ it really changed how I think about documentation)
With LLM you can do that using the undocumented Python Conversation API, but it's undocumented for a reason (I don't think it's good enough yet). You could also fake a previous conversation through the CLI tool but that is VERY undocumented - you would have to write fake rows into the SQLite database!
I also want to support the Claude thing where you can prefill the start of the response - amazing for things like forcing HTML by refilling an HTML doctype.
And what I expected of llm is for the template file to (optionally) contain an array of prompt/response pairs. You could even imagine the --save flag constructing one from an ongoing conversation of llm -c maybe.
We built http://github.com/robusta-dev/holmesgpt/ to investigate Prometheus/Jira/PagerDuty issues. We're able to get pretty good results (we benchmark extensively) because we use function-calling to give the LLM read acess to relevant data. I think we're the only open source AIOps tool, and the only AIOps tool period that does something more complex than RAG + summarization.
for example: I wrote a short bash script which uses yt-dlp and ffmpeg to download a song, convert it to my preferred format and then uses gemini to add metadata.
artist=$(gemini -s "Please respond with the name of the Artist based on
the songs title. do not use any other words, just the artist name.
example:
'Bruce Springsteen - Old Dan Tucker [S-GHbDFrwlU].opus'
Bruce Springsteen" -p "$opus_file" | tr -d '\n')very similar, although the ability to continue a conversation like you can in yours is a killer feature I wish it had.
Chaos.
You can find out more here: https://github.com/ferrislucas/promptr
npm i mktute
You can select between local model (ollama), claude 3.5 sonnet, or gpt-4. I've been surprised to find sonnet much better in performance and price for this task.
Open Interpreter lets LLMs run code on your computer to complete tasks. Eg "fix my python" or "make a movie out of the pictures in this directory."
I do think that LLMs have the potential to fundamentally change the way we interact with our computers. There's a lot of edge cases (especially when combining it with the inaccurate science of screen readers) but it's pretty mind-blowing when it works. I'm working on a blog post, but here's my little proof of concept working on both Windows in a web browser[1] and MacOS in the Finder [2].
[1] https://vimeo.com/931907811
[2] https://dvt.name/wp-content/uploads/2024/04/image-11.png
$ bashy find large files over 10 gb
find / -type f -size +10G
print -z $command
And the command will appear on your cli as if you had written it.I don't think you can do this in bash. Interestingly, this is something that seems quite difficult to both google and ask GPT for help. Both get confused and are thinking different questions are being asked. Probably because there are similar more common questions but the subtleties of possible wordings makes it difficult to differentiate.
Copilot is pretty good, but the forced change > commit > QA process that Aider forces you through is really powerful.
I love this. Simple and effective. RAG is just search leveled up with LLMs. Such an obvious thing to do. We know how to do search and can use it to unlock vast amounts of knowledge. Instead of letting LLMs dream up facts by compressing all knowledge into them, a better use of them is letting them summarize and reason about the facts it finds. IMHO the art is actually going to be in letting them come up with the right query as well. Or queries. It could be a lot more exhaustive in its searches than we could be.
The issue with the current generation of models is that they can't reason, they may do very well at pretending to reason, but they can't [1]. Reasoning requires the ability to identify and reuse patterns, and while there has been some advancement in this area [2] with getting models to learn the underlying pattern and rule, it doesn't generalize. This results in models that will happily tell you that a statement is both true and false, and be unable to identify the logical problem with that.
Even creating summaries is difficult, and LLM's are more than happy to hallucinate even when summarizing documents, providing incorrect, or entirely made up facts [3]. The general workaround is multiple runs with the same work and averaging the response, but that's a lot of work, and energy.
[0] https://win-vector.com/2024/05/21/i-want-flexible-queries-no...
[1] https://medium.com/@konstantine_45825/gpt-4-cant-reason-2eab...
[2] https://arxiv.org/abs/2405.15071
[3] https://community.openai.com/t/gpt-4o-hallucinating-at-temp-...
Anyway, I enjoy playing with this and because I am using my own little Python scripts, I can switch models and hack in it easily.
I say this because the scraper demo bit looks very neat, but I've been down this path before and I don't want to waste my time getting bad or deceptively incorrect results.
PS: I appreciate the work put into llm tho, it's a neat program I used with my Automator scripts to bring LLMs to macOS before Apple Intelligence was announced. I just wish the stability was not a concern.
I personally love using x-cmd. Small size (1.1MB), open source, interactive operation