How can we have a conversation and reach a shared understanding if one of the participants is avoiding attempting to understand the problem?
I have no problem with AI generated code, provided it’s reviewed, but I might as well be talking to a brick wall if the person hasn’t understood.
Easily solved - call that contributor into a meeting to explain their PR. If they cannot explain it verbally, it would be made clear to them why it cannot be approved without you having to actually say anything.
In the good old days (SVN), code reviews at places I worked at were done by booking one of the many small rooms and throwing the diff up on a projector. The author could defend, explain, etc.
Maybe we'll have to return to that type of code review from now on.
I wish we had a language that was targeted specifically for LLMs to write and humans and LLMs to inspect:
- Simple robust syntax
- One obvious way to do things
- Static type checking
- Purely functional encouraged, escape hatches for performance
- Inspect-able, testable, and reviewable in small pieces
- Something like formal predicates, preconditions, post-conditions, assertions, or effects typing
Giving LLMs all the surface area of Python, JavaScript, TypeScript, or C++ seems like a huge mistake. It's amazing it works as well as it does. Well written Haskell is beautiful, but there are way too many ways to write Haskell:
https://people.willamette.edu/~fruehr/haskell/evolution.html
I review a lot of Go code, which means I review a lot of LLM Go code. Even though the human vs LLM authorship distinction has strong signals, Go's simple nature seems like a useful constraint on how LLMs can express themselves.
- Type system is too concrete, can't express sets or unions.
- Not much support for functional programming except passing closures
- Can't express immutability. Well, there's const, but it's crippled
These things mean Go falls short of what GP wants.
But Go has very distinct and (honestly kinda weird) design goals, it isn't really supposed to be "the simple applications language". It's supposed to be "C with NewSqueak", very much a systems programming language. You see that in its aggressively concrete type system
Elixir's exhaustive pattern matching and gradual type checking is also on my radar.
I guess I don't know what you mean by effects typing.
These allow the model to cheat in ways that are harder to detect. In actually pure FP you can be sure that it's not counting function invocations to bypass tests.
Maybe these well known cases could be hidden in the API, but it's easy to come up with other examples. I forget what Clojure calls it, but they have some notion about things where they're mutable during "birth" and then locked down.
Someone smarter than me knows how to do this right.
Amusingly, a friend had Claude write some BrainF*ck the other day. Non-trivial algorithm, and it made working code on the first try.
> - Inspect-able, testable, and reviewable in small pieces
This is more a property of the architecture than a property of the language, although some programming languages support it better than others.
- Simple robust syntax
> - One obvious way to do things
That always leads to more verbosity. Much syntactic sugar is a more specialized way to do a subset of a more general thing. Why have `a + b` when you could just write `a.add(b)`? Because it's easier to read. The general functionality needs to account for all cases, the specialized one can cut that down just the things that matter to a specific common use case.
- Static type checking
\<meme>Which one?\</meme>
That's a very, very deep can of worms. One could argue that of an LLM is generating the code, then the type system doesn't need to be understandable to humans, so throw in all the features you'd ever want and just have the LLM change the code if it's not valid. Unions of higher order generic functions, sure! Or one could argue that there should be minimal magic, because the LLM understands the language only by its source, so everything should be explicit. If the LLM can prove that something is sounds to the compiler, accept it. Give ways to give extra evidence of soundness, like declaring invariants and contacts.
- Purely functional encouraged, escape hatches for performance
"One obvious way to do things", except when you need two. Purely functional except when it matters.
Why doesn't performance always matter? (And how will an LLM know if it does?)
Being purely functional is nice for data, but not for data structures that are updated in place. If all you do is stream data from one DB query into another, then your mutable state is the database. Otherwise might as well accept that it's a multiparadigmatic language with both imperative, functional and OO features. Just like all the others.
- Inspect-able, testable, and reviewable in small pieces
Good modularity and abstraction. No global scope. Maybe something like dependency injection to decouple from dependencies? (That does not make code readable, though.)
- Something like formal predicates, preconditions, post-conditions, assertions, or effects typing
That! LLMs look at the source. The more explicit the source is, the less it has to infer from context or existing knowledge. If the LLM can create its own predicates, accepted by the static type/analysis system, to prove that it's code is sounds, that allows more flexibility than having to fit into any fixed type system. Do we want or need that flexibility? Maybe. Most LLM-generated code is directly inspired by existing idiomatic code in the same language. Some is translated from other languages, trying to match up idioms. Some is just blindly trying to make unit tests pass. We could end up with generated predicates tailored to specific unit tests, not the actual concept.
How would this help? LLMs are fine with existing languages' syntax, semantics, idioms, etc. Where they fall over is in understanding intent of a prompt within the context of architecture, systems, roadmap, etc.
Currently Claude Code and Codex are primarily interfacing with code directly mostly using shell tools, while humans typically use IDEs.
Yes, I realize that there are lots of IDE-integrated LLM solutions, but this isn't exactly what I'm thinking about...
I'm thinking more of an API/MCP that the agent would need to use to work with the code which provides all the necessary interfaces, and designed specifically for AI agents.
When I observe its reasoning I see it doing this all the time.
You can delegate the responsibility of understanding to an LLM, but that is not the same.