In fact, I've had much easier time maintaining LLM assisted programs in Scheme and Clojure than other languages I've tried using because functional style naturally leads to low coupling. And that makes controlling context far easier than the rats nest of shared state that you have in imperative languages.
I have been trying Scheme and CL with LLMs for the last three years or so, and in recent months, I have finally decided that they are good enough.
My idea is that well, it's good enough that I can now produce more training data just by using Autolith with the most basic claude/gpt subs, haha
The niche language thing is really not a problem at all any more. If you're working in some esolang it doesn't take more than a 1-2k token primer in the context to get great results, and lisp is popular enough to not even need that.
The benefit of having the agent directly in the image like with Autolith here is that it can directly inspect all defined symbols and explore and orient itself automatically. Really doesn't need much guidance to get great results.
This all correct, I'd also add that in my experience, the GPTs are even better at Lisp, namely in the counting parentheses department.
Which is not an issue that much per-se because in Autolith, the harness detects Lisp file edits (CL, Scheme, Clojure) and gives hints when the edits lead to unbalanced files
(The heuristic is pretty simple, we detect if there's a mismatch, and if yes, it provide hints where the extra/missing might be based on indentation)
Do you have benchmarks on non-trivial tasks (say, generating zstd) that show it does any better than rust?
This is explicitly called out as only weakly supported in that blog post:
- You should use a popular language
- There's weak support for this statementIt makes sense, needing to train the model on things that aren't already in its weighs takes up valuable context. Until we have models that update their weights based on what they've seen in their recent sessions and learn like people, this will be a problem.
For now, though, between the results I'm seeing here, and the lack of need to look at code, I think this kills off any reason for me to use less popular languages.
I don't think there's really ever a downside to leaning in and making use of a language or system that works for you. Trying to tell people they should just use the popular thing is, imo, bad advice to turn hackers and experimenters into boring people.
AI changes the constraints here for now, since it can't permanently learn things. I'm waiting until that changes, but right now it's better to use what it knows out of the box if you want good results.
A better language doesn't buy me anything other than performance; the reason to stick an AI in here is to remove interactions with the code. I don't care what the AI chooses to use, as long as it gets results.
coincidentally, "good code" in popular lang is rarely directly attributed to only that part; and it's also about the underlying principles it tries to follow in the code... another example; is it typescript that's good, or are "types" inherently making things/feedback loops easier to reason about in llms? (only using ts here for all example because it's probably one of the most "trained on" pl)
I think there's an optimal ratio somewhere
These days I moved up the ladder of abstraction, so I don't really look; the main criteria I have is how the LLM gets things done.
If you give that to an LLM, it is then also able to iterate and develop faster.
The best that people have said about lisp is that evidence LLMs perform worse with it is weak.
You can’t tell that with a few uncontrolled runs
Autolith can spawn managed Lisp REPLs either from saved images (so it can do checkpoints) and triage changes before committing them to files, and then run test suites in the same REPL, it's been very useful for this.