1. Have you had much success w/moar macros in the age of the LLM? I've been impressed by the models' ability to write good ones, but I can tell my taste/judgement for macros isn't quite there. But they tend to be pretty good at writing gnarly ones, and I would love to work more macros into my workflow. Would love your thoughts. (Have I taken "Simple Made Easy" too far and left macro value on the table?)
2. Do your models ever get confused with image-based dev, and state? It's seemed dumb to me to have models keep running `sed` to change files, but it is nice to have a human-readable, filesystem-backed record of definitions. Would love to hear your experience here.
I've seen a big improvement in LLMs writing macros since Opus 5.5 came out. What really helps I think is that I've written skill files with my own examples and instructions.
Same thing for image-based dev. without an agent.md file with good instructions on how to work with a live image it will do dumb things. this sort of workflow is jsut so far off the training distribution.
I think what changed recently isn't that LLMs got better at CL, they just got way better at taking my skill/agent.md files and reasoning through them.
Also, in most languages an error will crash your program. So if you’re writing code with an LLM it will have to read your crash logs to make some changes and run your program again. In Common Lisp your program won’t crash, it’ll stop and open a debugger with the whole stack and all the variables. You can just point your LLM at the debugger, and it’ll make its fix and resume the program.
What does this look like in a real example system that you're maintaining? I can't imagine you'd always be able to resume like that if it's something like a webserver.Now, if your entire server is taken down because one connection threw an exception, that's bad design. But pretty much no major language works that way. All of them allow you to set things up so that an exception handling connection A won't affect connection B. And if you've done that in Common Lisp, then connection A halting and waiting for the debugger won't affect connection B either.
Anyway it sounds like this isn't really the target use case for the debugger since restarting a webserver is supposed to be easy. Maybe there's a different use case in mind?
Also of course the default is to not get the debugger but let the server thread crash and print the backtrace. With a user setting, you can choose to get the debugger, and have this request wait (the connection may time out, which isn't an issue during development). Another setting is to print the backtrace in the browser (= dev mode).
Python REPL and CL's are very different!
In CL: you can install new dependencies from the REPL. You can change a class definition, the existing objects will take the changes at the next invocation. You can control how this happens, etc. TLDR; CL is built around live programs. Buuut we can also do it the dumb (and safe) way following the industry's best practices.
Yes. I code Python in Emacs with the ancient Python support for running a REPL, loading buffers, editing a function and just hot reloading the modified function, etc.
I am an old man, and I like old man tools :-)
> can be useful in extreme situations
it's already useful in development, it's one of those things that shorten the dev loop and make development in CL a breeze. For production, you choose. Connect to the live image to inspect the state without modifying anything or debug while it's live.
> recorded nowhere
you can connect to a running program and have the source under VCS. The thing to do isn't to copy-paste new function definitions in the live image's REPL, but to connect to the image, make changes to the source files, and re-compile them (a C-c C-c in Slime) (sending changes to the running program).
> safer
yep, some things are safer. Advanced features are useful even for simple things (introspection, etc).
https://comp-348.github.io/lisp-debugging.html has an example of what it looks like. A toy example, to be sure, but the basics of a real example would still look the same. You're given a choice of several options, very reminiscent of the "Abort, Retry, Ignore?" choice that used to be oh-so-familiar in the days of DOS. Except this one is more useful, because it offers ways to specify how to resume. E.g., the toy project is halting on a `(print X)` call where the value of X is not defined. And the choices are:
0. Continue. (Retry using X).
In the toy example, this would fail, because nothing else has defined X. But in real code, the name might have been undefined because the data needed to define it hadn't arrived yet, from the database or the filesystem. In which case retrying the statement might work the second time.
1. Use-value. (Use specified value).
This one prompts you to enter a value for the undefined variable, and continues, but it does not modify the value of X in the program. The next time the program tries to use X, it will halt again with another "unbound variable" error.
2. Store-value. (Set specified value and use it).
This one, just like Use-value, will prompt you to enter a value to use... but then it will set X to that value and continue running the program. Next time the program tries to read the value of X, it will have one, and the program won't halt.
3. Abort. (Exit debugger, returning to top level).
This is what you would choose if there's no good way to fix the error, and you just have to quit the program and restart. Though note that choosing this option isn't going to exit the program you're debugging, just take you out of the debugger. You'll still need to kill-and-restart it some other way... or come back an hour later when the value is finally available, and then choose options 1 or 2.
Hopefully that gives you a taste for what the CL debugger is like to use in practice.
And it's not editing files on the server, it's actually reaching into the running code and tweaking its values.
That, I think, is the difference here. In many languages, the debugger can pause on the exception and let you inspect the code. But in every other language I've used, once you edit the code to fix the bug, you can't resume from where the debugger paused. You have to recompile the code and resume from the top. In CL, you can resume from exactly the state you were in when the debugger paused, only this time with the correct data in place. (Or even with a code fix having been applied, live, to the code).
Second, other languages can also do that, Java for example has facilities to recompile functions and then rewind the stack to run the function again from the beginning of it. Dart also does that even better and people use that extensively in Flutter with hot reloading. I do that in a debugger during tests though, never seen it done in prod.
So, in practice, it might look like you hit any kind of runtime failure, and then the LLM writes some code to fix it, and the user's request completes successfully with no errors.
I'm the author and I share the same philosophy, hahaha, my users told me about this post
DSL presumes agreement on semantics, and that's often the most difficult part.
- economies of scale no longer work, and you end up doing a custom ERP for your business from scratch.
- your business changes might invalidate your model quickly. You sell through distributors, but open an online shop -- and suddenly your customer is not one of few dozen well-known businesses with a known address and tax number, but user2252 who bought something late at night last night. And you want to understand the needs and behaviors of both.
- for the economies of scale, you might develop your custom solution 20x faster now, but you're still in competition with the established provides with templates for most of the cases (who btw have the same LLM capability at their disposal)
- for the change in business -- you can ask an LLM which changes this induces, and it will give you most typical impacts. And then you're back at square deciding if it's better to roll your custom DSL and the custom system downstream, or just use off-the-shelf stuff that covers 95% of it from day one (well, maybe day two or three)
When it comes to back office business programming, there’s just a lot of code tasked with copying a litany of bits of data from one structure to another.
Whether it’s copying a web form into a database, or converting Their JSON to Your JSON, it’s a lot of detail that does not abstract well. It’s all shapes and sizes and formats, and it almost always has to be enumerated in excruciating detail and, typically, twice.
Sure, there’s logic and whatnot involved, but it, too, is specific to some subdomain of the larger system and it, too, does not abstract well. Not in the large context of the overall system.
Accounts Payable and Accounts Receivable, at 10,000 feet look almost identical. They’re almost literally the same thing with the sign flipped. But in practice, they don’t share code well. You end up with two similar systems, but not similar enough where sharing is actually worthwhile.
At best they can leverage a common API to the GL.
Turns out a lot of languages can manifest a decent level of abstraction. But even then, folks push back.
Consider the love/hate relationship with ORMs. Or the annotation driven markup in Java programs and the underlying “magic” that they enable. Like scribing mystic runes onto things.
Those are both very powerful, yet folks experience that and toss their hands in the air and throw out the baby with the bath water and jump into something “magic free” like Go.
Just because you can use something like CL to “make your own magic”, doesn’t mean it’s a good idea. Doesn’t mean it scales. Doesn’t mean it communicates well to others. AI or no.
It’s not the AIs world yet. We already know that if the AIs want a better language suited to AI efficiency, they’ll come up with their own. I’ve already seen crass examples of “code only an AI could love”. Completely impenetrable, at least to me. May as well have represented it as a color image and collection of RGB values. Opaque to me, but the AI could “read” it.
There is much more to programming and systems than token density, and AI is still getting cheaper by the day, so less reason to even pre-optimize for it anyway.
While I am extremely taken with Lisps and the lisp way of doing DSLs, I would probably go with an OCaml to make a DSL for a company specific ERP. It seems a better way to go about the problem.
Lisp, on the other hand, I have found to be extremely good at domains which seem the same but which are tremendously different. For example, a workout app is a surprisingly complex domain. Different exercises have different storage models and functions, as do different training sessions and different programs. Rather than try to build a monoprogram, one training app to rule them all, I find lisp wonderful for making "microprograms".
This bears resemblance to Accounts Payable and Accounts Receivable but I don't think Lisp would be a good fit for those. Perhaps a Lean or a Rocq, something with proofs.
> We already know that if the AIs want a better language suited to AI efficiency, they’ll come up with their own.
My agents seem to really like Tree Calculus and have bullied me into working on a language which uses it.
There shall be one minimal, ultra-hardened, tiny attack surface, "majority gate" picking the answers that most implementation agrees on.
This shall not only detect a great many implementation issues but also it'll help find security issues and platform defects (say the Common Lisp, Haskell, Rust and Python all agree but the Java one fails: in rare case it'll be due to a JVM bug and finding that out shall be simplified).
Code shall be generated from specs in n languages and ran on n stacks. The gate shall return the answer as soon as a quorum is met and, later on, any bogus answer arriving shall be cause for enquiry.
We'll have such systems, it's just a matter of time.