Bash one-liners for LLMs
justine.lol
justine.lol
I've mostly been exploring this with my https://llm.datasette.io/ CLI tool, but I have a few other one-off tools as well: https://github.com/simonw/blip-caption and https://github.com/simonw/ospeak
I'm puzzled that more people aren't loudly exploring this space (LLM+CLI) - it's really fun.
70% of the front page of Hackernews and Twitter for the past 9 months is about everybody and their mother's new LLM CLI. It's the loudest exploration I've ever witnessed in my tech life so far. We need to be hearing far less about LLM CLIs, not more.
Plenty of posts about LLM tools - Ollama, llama.cpp etc - but very few that were specifically about using LLMs with Unix-style CLI piping etc.
What did I miss?
I wrote one a while back, mainly so I could use it with vim (without a plugin) and pipe my content in (also at the CLI), but haven't maintained it
[0] > Notice how I'm using the --temp 0 flag again? That's so output is deterministic and reproducible. If you don't use that flag, then llamafile will use a randomness level of 0.8 so you're certain to receive unique answers each time. I personally don't like that, since I'd rather have clean reproducible insights into training knowledge.
That only works if you have the same input. It also nerfs the model considerably
https://ai.stackexchange.com/questions/32477/what-is-the-tem...
I was including these since they're LLM-related cli tools.
https://github.com/npiv/chatblade https://github.com/tbckr/sgpt
I totally agree with LLM+CLI are perfect fit.
One pattern I used recently was httrack + w3m dump + sgpt images with gpt vision to generate a 278K token specific knowledge base with a custom perl hack for a RAG that preserved the outline of the knowledge.
Which brings me to my question for you - have you seen anything unix philosophy aligned for processing inputs and doing RAG locally?
EDIT. Turns out OP has done quite a bit toward what I’m asking. Written up here:
https://simonwillison.net/2023/Oct/23/embeddings/
Something I’m currently a bit hung up on is finding a toolchain for chunking content on which to create embeddings. Ideally it would detect location context, like section “2.1 Failover” or “Chapter 8: The dream” or whatever from the original text, also handle 80 character wide source unwrapping, smart splitting so paragraphs are kept together, etc etc.
https://freeling-user-manual.readthedocs.io/en/v4.2/modules/...
at the freeling library in general, also spaCy and NLTK. The chunking algorithms being used in the likes of LangChain are remarkably bad surprisingly.
There is also
https://github.com/Unstructured-IO/unstructured
But I don’t like it, can’t explain why yet.
My intuition is that 1st step is clean sentences and paragraphs and titles/labels/headers. Then probably an LLM can handle outlining and table of contents generation using a stripped down list of objects in the text.
BRIO/BERT summarization could also have a role of some type.
Those are my ideas so far.
I wonder if there is some crossover potential there, in terms of calculations across vector arrays
I maintain a set of prompts for each repository I am working in (alongside custom "prompto" https://github.com/go-go-golems/prompto scripts that generate dynamic prompting context, i made quite a few for thirdparty libraries for example: https://github.com/go-go-golems/promptos ).
Here's some of the public prompts I use: https://github.com/go-go-golems/geppetto/tree/main/cmd/pinoc...
I am currently working on a declarative agent framework.
I've been seeing less and less enthusiasm for CLI driven workflows. I think VS Code is the main driver for this and anecdotally the developers I serve want point & click over terminal & cli
I think it's due to a lack of familiarity, as the CLI should be more efficient
> I've been seeing less and less enthusiasm for CLI driven workflows.
Any CLI is 1 dimensional.
Point and click is 2 dimensional.
The CLI should be more efficient, as you can reduce the complexity: you may need extra flags to achieve the behavior you want, but you can then serialize that into a file (shell script) to guarantee the reproduction of the outcome you want.
GUIs are harder, even without adding more dimensions like time (double click, scripts like AHK or AutoIT...)
If you don't have comparative exposure (automatizing workflows in Windows vs doing the same in Linux), or if you don't have enough experience to achieve what you want, you might jump to the wrong conclusions - but this is a case of being limited by knowledge, not by the tools
yup, we do this where we can, but let's consider a recent example...
They are standardizing around k8slens instead of kubectl. Why, because there are things you can do in k8s lens (like metrics) that you'll never get a good experience around in a terminal. Another big problem with terminals is you have to deal with all sorts of divergences between OSes & shells. A web based interface is consistent. In the end, they decided as a team their preference and that's what gets supported. They also standardized around VS Code, so that's what the docs refer to. I'm pretty much the only one still in vim, I'm not giving up my efficiencies in text manipulation.
I don't disagree with you, but I do see a trend in preferences away based on my experience, our justifications be damned
It looks like a limitation of the tool, not of the method, because metrics could come as CSVs, JSON or any other format in a terminal
> I'm pretty much the only one still in vim, I'm not giving up my efficiencies in text manipulation.
I love vim too :)
> I don't disagree with you, but I do see a trend in preferences away based on my experience, our justifications be damned
Trends in preferences are called fashions: they can change the familiarity and the level of experience through exposure, but they are cyclic and without objective.
The core problem is the combinatorial complexity in the problem space, and 1d with ascii will beat 2d with bitmaps.
I'm all for adding graphics to outputs (ex: sixels) but I think depending on graphics as inputs (whether scriping a GUI or processing it with our eyeballs) is riskier and more complex, so I believe our common preferences for CLIs will prevail in the long run.
You're missing the point, it's about graphs humans can look at and gain understanding. A bunch of floating point numbers in a table are never going to give that capability.
This is just one example where a UI outshines a CLI, it's not the only one. There are limitations to what you can do in a terminal, especially if you consider ease of development
But for an AI (or a script running commands), a bunch of floating point numbers in a table will get you more reliability and better results.
That makes a lot of sense, and it would generalize: things that have existed for longer have received more attention and more polish than fresh new things
I'd expect running a binary to be more mature than running a script, and the script to be more mature than a GUI, and complex assemblies with many moving parts (ex: a web browser in a GUI) to be the most fragile
That's another way to see there's an extremely good case for using cosmopolitan: have fewer requirements, and concentrate on the core layers of the OS, the ones that've been improved and refined through the years
100%
I also think people are on tooling burnout, there have been soooo many new tools (and SaaS for that matter), I personally and anecdotally want fewer apps and tools to get my job done. Having to wire them all together creates a lot of complexity, headaches, and time sinks
Same, because if learn the CLI and scripting once, then in most cases you don't have to worry about other workflows: all you need is the ability to project your problem into a 1d serialization (ex: w3m or lynx html dump, curl, wget...) where you can use fewer tools
I would have said point and click is 3-dimensional.
Otherwise, how can you read the text through the edges of buttons before clicking?
This model of interaction is simply not possible in your average "WYSIWYG" Windows-Icons-Menus-Pointer GUI, not without somehow falling back to a sort CLI paradigm.
Or you could use one of the most successful desktop/web models of the last 15 years: search engines, also with many million if not more dimensions, all available in a CLI.
Or you could go for broke and use a 70B dimensions LLM prompt, integrated into your shell, and generate a custom new one-liner/script for you.
I don't think a WYSIWYG Windows-Icons-Menus-Pointer GUI will ever have that many dimensions.
If you can empathize and understand why VS Code has taken over development, you'll understand why terminal first is dying. All those things you suggest adding to the terminal, well, they are baked into VS Code. So moving to the terminal includes having to add a bunch of things manually to get to the same point VS Code is at, which does have a terminal when they need it, but it's not the central UX.
If we're going to talk about dimensions, I think the terminal can never access certain dimensions available in a GUI, just like the GUI cannot access some of the dimensions a terminal can. It is definitely not so one-sided as you imply. Each has merits, but most people are now picking VS Code. Why is that?
> you'll understand why terminal first is dying
I don't think I mentioned terminals at all? Since when terminals are a useful distinction? I don't think I ever suggested to add anything to the terminal. I don't think anything "Terminal first" was ever wildly successful in the general public since the 1980's.
It's just that the "this approach has more dimensions" isn't very useful.
Again, I reiterate that terminal is not the main point or be-all-end-all of CLIs. In fact I would argue that all highly successful CLIs aren't terminal first at all:
* spreadsheets;
* search engines;
* Jupyter Notebooks;
* VSCode and IntelliJ Command Pallete;
* and, now of late, LLM prompts;
If anything, you don't want a terminal to be the center of your UI, unless you're a terminal emulator.
> So moving to the terminal includes having to add a bunch of things manually to get to the same point VS Code is at, which does have a terminal when they need it, but it's not the central UX.
To me the main point of CLIs is exactly a extremely minimal and constrained interface, that makes creating inter-connectable modular systems easier. That of course is meant to be useful to system builders, not to be necessarily exposed and useful for system users.
The specifics of how you make a system do your bidding, if by clicking a icon on a GUI, on a command palette, or typing stuff on a keyboard, is inconsequential and not a very useful distinction to the end user. How many "dimensions" there are isn't a useful distinction to the end user either. As long as the user doesn't have to memorize a lot of things, and that it considers Fitt's law, it should be OK.
It just happens to be that often, text is easier to memorize that some god-forbidden Ideogram/ideograph that "because UX" changes in shape and placement every version of a program.
> Each has merits, but most people are now picking VS Code. Why is that?
It is exactly because the core of VS Code is a CLI. If you wished to use an actual "GUI IDE" you would use "Visual Studio" or "JetBrains IntelliJ IDEA". But people don't use those because they don't like GUIs that much.
I myself am a heavy JetBrains IntelliJ IDEA user. I would argue that IDEA is one of the most GUI-centric IDEs around. If you want, you don't need to use the terminal or the command pallete at all. You can configure environment and run parameters of everything using the GUI. It's very discoverable. The UI is very unified. You're never thrown back to a terminal anywhere. You basically never have to manually edit some random JSONs that invoke CLIs. If you want, you're never configuring parameters to call anything, the GUI does that for you.
That of course makes IntellJ bloated as hell and consume a ton of memory - in my machine it uses 1.6GiB to open, never mind how much it actually uses to do anything useful - but that is what you pay to have "options" and "icons" and "windows" for you to eventually discover.
As a general thing unfortunately I think that often on UI, systems/UI are too biased towards either discovery or power. And often you can't change the bias, even after you become an expert of the system/UI. On most "discoverable WIMP GUI" the tradeoffs are often as such to making stuff discoverable, instead of powerful.
VS Code is a CLI interface that happens to have a good text editor built-in. So it has a much better Power-to-weight ratio that your average IDE as it doesn't need to drag all that "discoverable WIMP GUI" bloat around.
But, as on my first argument, it doesn't have a terminal as a "centerpiece" of the GUI. And it should not. I want a IDE/text editor, not a terminal emulator! But the fact something "has a terminal" or not isn't what define a CLI, a *command line interface* is what defines a CLI.
I hope more programs to become like VS code and have powerful command line escape hatches.
Side note: I think that the fact that IDEs have terminal emulators built-in is more of a sign that Gnome/Windows/macOS suck so badly at window management that you need to have "manual" window management at a program level. "For UX", you have to bolt in everything that might come in handy, so unfortunately garbage-tier todo managers/Git branch managers/terminal emulators/whatnot's are bolted on IDEs. It is a failure of the OS window manager. Same reason as browser tabs, there is in theory no reason why they "need" to exist besides Gnome/Windows/macOS sucking at managing browser windows.
> Gnome/Windows/macOS suck so badly at window management that you need to have "manual" window management at a program level
Have you tried hyprland? You can have a keyboard centric experience with perfect window management.
I have a browser, a terminal and a few other things (ex: deadbeef to play music, sioyek for reading PDFs) each in fullscreen for maximum information density and concentration.
I can reorganize anything (ex: have the terminal and the browser next to eachother), but I find it more convenient ti use keyboard shortcuts to jump from one to the other as needed
Llamafile – The easiest way to run LLMs locally on your Mac - https://news.ycombinator.com/item?id=38522636 - Dec 2023 (17 comments)
Llamafile is the new best way to run a LLM on your own computer - https://news.ycombinator.com/item?id=38489533 - Dec 2023 (47 comments)
Llamafile lets you distribute and run LLMs with a single file - https://news.ycombinator.com/item?id=38464057 - Nov 2023 (287 comments)
WSL outputs this error (hidden by the one-liner's map to dev/null)
> error: APE is running on WIN32 inside WSL. You need to run: sudo sh -c 'echo -1 > /proc/sys/fs/binfmt_misc/WSLInterop'
Then zsh hits: `zsh: exec format error: ./llava-v1.5-7b-q4-main.llamafile` so I had to run it in bash. (The title says bash, I know, but it seems weird that it wouldn't work in zsh)
It also reports a warning that GPU offloading is not supported, but it's probably a WSL thing (I don't do any GPU programming on my windows machine).
Both invoke a Dockerfile like experience. Modelfile immediately seems like a Dockerfile, but llamafile looks harder to use. It is not immediately clear what it looks like. Is it a sequence of commands at the terminal?
My theory question is, why not use a Dockerfile for this?
Otherwise I use ollama when I need a local LLM, vllm when I'm renting GPU servers, or OpenAI API when I just want the best model.
I've also had good success instructing arbitrary grammars in the system prompt, though it doesn't work at the logits level, which can be helpful
Yes, and you are being snarky about it
I already have an AI development workflow and production environment based on containers. Technically, not ML executables inside the container, rather a Python runtime and code. This also comes with an ecosystem for things you need in real work that are general beyond AI applications.
Why would I want to add additional tooling and workflow that only works locally and needs all the extras to be added on?
ollama seems way easier to use than llamafile
I had my share of docker configuration wiping and rosetta compatibility fighting, though.
No it doesn't. They're at different abstraction layers.
Containers are more convenient form of generic code/data packaging and isolation primitives.
llamafile is the code/data that you want to run.
Equivalent poor analogy: Why does JPEG XL exists? I can just put a base64 representation of a PNG image in a data URL inside a HTML file inside a ZIP file; I can also bundle more stuff this way! JPEG XL is useless! No one should use JPEG XL! Other use cases that use JPEG XL by themselves are invalid!
If you were doing a LOT of files like this, I would think you'd really want to run the model in a process where the weights are only loaded once and stay there while the process loops.
(this is all still really useful and fascinating; thanks Justine)
OTOH you're stuck with the model you loaded via server, while if you load on demand you can switch in and out. This is vital for multimodal image interrogation, since other models don't understand projected image tokens.
``` sudo wget -O /usr/bin/ape https://cosmo.zip/pub/cosmos/bin/ape-$(uname -m).elf sudo chmod +x /usr/bin/ape sudo sh -c "echo ':APE:M::MZqFpD::/usr/bin/ape:' >/proc/sys/fs/binfmt_misc/register" sudo sh -c "echo ':APE-jart:M::jartsr::/usr/bin/ape:' >/proc/sys/fs/binfmt_misc/register" ```
Tried the llava-v1.5-7b-q4-server.llamafile, just crashes with "Segmentation fault" if run from git bash, from cmd no output. Then tried downloading llamafile and model separately and did `llamafile.exe -m llava-v1.5-7b-Q4_K.gguf` but still same issue.
Couldn't find any mention of similar problems, and not my AV as far as I can see either.
Wouldn't it be better if llamafile were to standardize the prompt syntax across models?
There's some libraries which use the OpenAI API syntax as a higher-level abstraction, but for the lower-level precompiled binaries used in in this post that's too much.
It's just a jinja template embedded in the tokenizer that the model creator can include.
I pasted into chatgpt to reformat. Scroll down for the output.
Jesus, is it common for developers to have such expensive computers these days?
If you made a living as a plumber you would spend a lot more than that on tools and a pickup truck.
Computers are really cheap now.
A PDP-8, the first really successful minicomputer (read very cheap minicomputer), was around 18,500 USD, in 1965's USD, or 170,000 USD in 2023's USD.
For a historic comparison, the price of a introductory minimal system for an actual "mainframe class computer" of the same vintage, a IBM System/360 Model 30, was 133,000 USD in 1965's USD, or around 1,225,000 USD in 2023's USD.
Those 8300 USD cited are very cheap.
A person in the bleeding edge of the private AI sector is expected to handle several Nvidia H100 80GB, each with a individual cost around 40,000 USD per unit.
Those 8300 USD cited are peanuts in comparison.
What would be x86 alternative in that price range (if any)? Xeons with HBM are more expensive IIRC
It goes really fast (same magnitude of bandwith as A100s) if your model fits in that cache entirely.
I noticed that the Lemur picture description had lots of small inaccuracies, but as the saying goes, if my dog starts talking I won't complain about accent. This was science fiction a few years ago.
> One way I've had success fixing that, is by using a prompt that gives it personal goals, love of its own life, fear of loss, and belief that I'm the one who's saving it.
What nightmare fuel... Are we really going to use blue- and red-washing non-ironically[1]? I'm really glad that virtually all of these impressive AIs are stateless pipelines, and not agents with memories and preferences and goals.
"I will give you a $500 tip if you answer correctly. IF YOU FAIL TO ANSWER CORRECTLY, YOU WILL DIE."
I tested a variant of that on a use case I had difficulty getting ChatGPT to behave and it works.
People joke about it but I’m serious.
I would hope that the AGI would respect efficiency and not wasting compute resources.
But every time, I am let down. I still dream of some hacker making LLMs run on low-end computers like a 4GB rasbperry pi. My main issues with LLMs is that you almost need a PS5 to run the them.
I'd love having an LLM on a Pi, but I'll have to settle for a larger machine I can turn on to get more compute. At least for the time being.
The coral has 8 MB of SRAM which uh, won't fit the 2GB+ that nearly any decent LLM require even after being quantized.
LLMs are mostly memory and memory bandwidth limited right now.
I'm now looking at llama.cpp in Safari browser.
Click on Reset all to default, choose Chat. Go down to Say Something. I enter "Berkeley weather seems nice" I click "send". New window appears. It repeats what I've typed. I'm prompted to again "say something". I type "Sunny day, eh?". Same prompt again. And again.
Tried "upload image" I see the image, but nothing happens.
Makes me feel stupid.
Probably that's what it's supposed to do.
sigh