How to tackle unreliability of coding assistants
martinfowler.com
martinfowler.com
While coding assistants seem to do well in a range of situations, I continue to believe that for coding specifically, merely training on next-token-prediction is leaving too much on the table. Yes, source code is represented as text, but computer programs are an area where there's available information which is _so much richer_. We can know not only the text of the program but the type of every expression, which variables are in scope at any point, what is the signature of a method we're trying to call, etc. These assistants should be able to make predictions about program _traces_, not just program source text. A step further would be to guess potential loop invariants, pre/post conditions, etc, confirm which are upheld by existing tests, and down-weight recommending changes which introduce violations to those inferred conditions.
ChatGPT and tab-completion assistants have both given me things that are not even valid programs (e.g. will not compile, use a variable that isn't actually in scope, etc). ChatGPT even told me that an example it generated wasn't compiling for me b/c I wasn't using a new enough version of the language, and then referenced a language version which does not yet exist. All of this is possible in part b/c these tools are engaging only at the level of text, and are structurally isolated from the rich information available inside an interpreter or debugger. "Tackling" unreliability should start with reframing tasks in a way which lets tools better see the causes of their failures.
In fact, GPT is really good at correcting its code given the compiler output. We just need to automate that.
I think that this is not done yet because LLMs costs are O(n+1) so you really don’t want it to be stuck in some loop.
There's a lot of headroom for sophistication as a few more research insights are made and an experienced tooling team commits a few years to making rich multimodal/ensemble code assistants that can perform smart analyses, transformations, boilerplating, etc on your project instead of just adding some text to your file.
But it'll take a years of insight and labor to build that system up, adapt it to different industries/uses, and prove it out as engineering-ready.
People get caught up in the novelty of Copilot and ChatGPT and then imagine that the revolution arrives when some new paper comes out tomorrow (or that none will arrive because of today's limitations), but the far more likely reality is that the revolution paces as something more like the internet's -- real and profound, but unfolding gradually over the span of decades as people work hard to make new things with it.
Not necessarily good for Nvidia. If the new algorithm is branch heavy, don’t parallelize well, … etc. and operate better on CPUs rather than GPUs, it will be Intel, AMD, and ARM that enjoy the windfall.
Just a few days ago there was a (somewhat cryptic) report about OpenAI developing a model that could do simple math. I'm certain they're pursuing quite a lot of research that's not directly related to LLMs.
I don't think we're heading towards AGI any time soon, but I can imagine a complex system in the next couple years which uses an LLM for text generation, but offloads its serious "thinking" on to predictable, specialized subsystems. It's easy to imagine a lot of possibilities for code in particular.
Alpha fold.
The leap that area took thanks to it was a leap.
It's more than just something potentially better
(Well, not artificial, but you get my point.) It couldn’t design ASICs in minutes.
Nobody wants to hear this, but AGI could very well be the next invention after practical cold fusion and faster-than-light travel.
Yes, we're moving faster, but does anyone know how long the road is?
A space rocket is moving very quickly (50k kmph?), much faster than any other technology such as airliners, maglev trains, but the closest star is still several light years away.
I disagree with this conclusion. Crude as it is, Copilot _does_ make me faster in practice, even though I have to proofread its output. I think this also agrees with your second statement, that progress will come in small but useful improvements. If you treat Copilot like improved auto-completion, then it is very useful today.
Actually, while using syntax trees as the underlying model is a very interesting approach, there is still a lot of low-hanging fruit to make Copilot much more useful than it is today, just in terms of performance (waiting 1-2 seconds to see if there is a completion, and if it makes sense, interrupts my typing flow) as well as low-level features (e.g. mark and insert part of a suggestion, instead of inserting the whole and then removing the wrong parts).
The first tech giving me internet reliable was DSL or 60kbytes/sec.
This tech could only be replaced with replacing or adding cables.
Nvidia is making a ton more petaflops per month than last year and a lot of money and people have moved/shifted in a very short period of time.
Even my company saw the writing and the only/primarily thing we talk about is ml.
I also don't have to wait for services like a bank or anyone else supporting emails or internet services.
We have everything already in place for a much faster speed than at any other tech before.
-Arthur C. Clarke
You'd need something that actually understands the language. What is a lifetime, what is scope, what are types, functions, variables, etc. Something that can contextually look at correct, incomplete, or broken programs and reason about what they're doing by knowing the rules that the language operates on and following them to a conclusion. It would also need to understand high-level design patterns and structure to not just know what's happening in the literal sense "this variable is being incremented" but also in a more abstract sense "this variable keeps track of references because this is mixed C/Python code that needs that to handle deallocation". Something that recognizes patterns both within and outside of the code, with appropriately fuzzy matching where needed.
And I think importantly you'd need to be able to query it at a variety of levels. What's happening on a line-by-line basis and what's happening at the high level at a given point of code. One OR the other isn't sufficient.
That is not a simple ask. We're a long way away from something that smart existing.
A step in this direction, which I've been trying to figure out as a side project, and which I would love someone to scoop me on and build a real thing, is to stitch an ML model into into a (mini)kanren. Minikanren folk have built relational interpreters for small languages but not for a "real" industrial language (so far as I'm aware). These small relational interpreters can generate programs given a set of relational statements about them (e.g. "find f so that f(x1) == y1, f(x2) == y2, ..."). Because they're actually examining the full universe of programs in their small language, they will eventually find all programs that satisfy the constraints. But their search is in an order that's given by the structure/definition of the interpreter, and so for complex sets of requirements, finding the _first_ satisfying example can be unacceptably slow. But what if the search were based on the softmaxed outputs of an ML model, and you do a least cost search? Then (roughly) in the case that beam-search generation from the ML model would have produced a valid answer, you still find it in roughly the same time, but you _know_ it's valid. In the case where a valid answer requires that in a small number of key places, we have to take a choice that the model assigns low probability, then we'll find it but it takes longer. And in the case that all the choices needed to construct a valid answer are low-probability under the model, then we're basically in the same place that the vanilla minikanren search was.
It's too easy to fall into infinite loops for something with only a naïve understanding of the questions.
This is where the massive corpus of source code available on the Internet can help generate a "LSM" (large software model) if you can expose the tokens as the lexer understands them in the training set.
If your LSM sees a trillion examples of correct usage of lifetime and scope and types and so on, then in the same way that an LLM trained on English grammar will emit text with correct grammar as if it understands English, your LSM will generate software with correct syntax as if it understands the software. Whatever the definition of "understands" is in the context of an LLM.
- natural language is flexible, computer languages are less so.
- "pretty decent English" still includes hallucinations. I've seen companies whose product demo for generating marketing copy just makes up a plausible review. Hallucinating methods, variables, other packages/modules yields broken code.
- the human thought behind natural language is not feasible to directly provide to a model. An IR corresponding to the source of the program is feasible to provide. A trace of the program executing is feasible to provide. Grounding an LLM in the rich exterior world that humans talk about is hard; grounding an LSM in the rich internal representations accessible to an IDE or a debugger is achievable.
Indeed, Chat GPT 4 and Copilot can generate "pretty decent code" that will look fine to the average human coder even when it's incorrect (making up methods or getting params wrong or slighly missing requirements or similar).
The level of precision required for "pretty decent non-trivial code" is much higher than prose that looks like it was written by an educated human, so I share the idea that if it was augmented - even in really stupid ways like asking the IDE if it would even compile, in the case of Copilot, before suggesting it to the user - it would work much better at a much lower effort than increasing it's understanding implicitly by orders of magnitude.
I would have thought babies have been showing this beyond a doubt since time immemorial.
So far it really looks like a sufficiently large LLM with a sufficiently large / high-quality dataset can learn basically anything (given a slightly generous interpretation of the word "sufficiently").
In case of code completion, I think just training on program output in addition to source code would already unlock huge capability boosts.
Why is that more important for programming languages than for human natural languages?
A compiler cannot do these things. The code must follow the rules of the language perfectly, or the code won't compile. And further, it doesn't understand intention like a human can. It doesn't know that you intended to loop over all items in a collection - if you have an off-by-one error it'll happily compile that regardless.
Another area I wish Copilot understood: where my cursor will go next. Right now it can only feed me new lines, but I bet the underlying LLM is already powerful enough that with little fine-tuning, it could guess that after I edit a call to foobar() to add an argument the next thing I'll do is probably edit the definition of foobar.
It feels like there's a trove of UX improvements in there; we've barely scratched the surface.
My point is, I think Copilot-style tools could do more that insert text. They could recommend lines of code to remove, give you a quick keyboard shortcut to "place cursor in next file / line you're likely to want to edit", etc.
IDEs can do some rigid versions of that (eg "Jump to next diagnostic", "Refactor", "Jump to implementation", etc) but Copilot could be the more general, more powerful version of these features.
We're currently working on two forks off this... one, using the expected type information in conjunction with existing grammar constrained generation to enforce per-token type-correct generations, and two, using some program synthesis unevaluation techniques to also provided runtime trace data relevant to the current program hole we're trying to generate a completion for.
But yeah in general there is so much existing work on static and dynamic program analysis which can be applied here... I think a lot of the interesting challenges are going to be UI/UX ones... interactive processes to help more precisely specify intent and iteratively valid generations as more and more code is written autonomously.
One thought I've had but haven't experimented with yet is that we could leverage a lot of the existing tooling that was made for humans - e.g. tab-complete providers like Jedi, which do some type inference behind the scenes. It's able to provide suggestions of valid members for a given cursor position, and so logits could be warped to prefer tokens which match the suggestions (so if the output at a given time is `math.sq`, `math.sqrt` would be much preferred over `math.square_root` which doesn't exist). You'd have to be a bit smart around this though because in since situations such as when using an identifier for the first time it's not yet in scope and you don't want the LLM to never create variables. Maybe some beam search shenanigans and heuristics could be enough to get useful output, but at that point it no longer seems like a quick thing to just try out :)
Couldn't we say the same thing about almost all UNIX/Linux coreutils? There's no way to get a strongly-typed array/tree of folders, files, and metadata from `ls`; it has to be awk-ed and sed-ed into compliance. There's no way to get a strongly-typed dictionary of command-line options and what they do for pretty much every coreutil; you have to invoke `something --help` and then awk/sed that output again.
These coreutils are at their core, only stringly-typed, which makes life more difficult than it strictly needs to be.
Everything is a bag of bytes, or a stream to be manipulated. This philosophy simply leaked into LLMs.
It could still be text, or at least bytestreams: program traces, OpenTelemetry logs, objdump of the compiled binaries, LLVM IR dumps, compiler errors, syntax highlighting markers, etc.
Alone with copilot the amount of money flowing into this area is much more now than just a year ago.
And based on a Google research blog article, they use it internally with over 50% of suggestions being accepted.
The race is on and no one can afford not to play the game.
Are you sure 'traces' is the right word? Not something more like ASTs?
Btw, the predict-next-token approach has the benefit of also being able to deal with broken code.
You might also want to compare http://www.incompleteideas.net/IncIdeas/BitterLesson.html and https://gwern.net/scaling-hypothesis with your idea of adding more domain specific knowledge.
Just makes me more excited for future iterations!
Then maybe have a separate model trained on going from syntax tree to source code.
I don't see why you need a model for this. But yes, this is a very cool idea.
It must have gaming forums in its training data. Gamers know that the solution to all problems is to upgrade (game/drivers/OS) to the latest version.
I found this effect to be most pronounced when writing tests. I think Copilot shines in codebases with static typing, clear interfaces, doctstrings and unit tests. That's really about the densest, most richly annotated context you could give to an LLM. And that's before adding capabilities for more well-defined reasoning about static languages, types, etc. - there is potential for it to get even better at this.
For pure functions, it is 100% correct and complete essentially 100% of the time, allowing me to write descriptive test names and nothing more. Even when mocks or spies have to come into the picture, it is usually 95% accurate. The key is that you should always have the file you are testing open in an editor tab, as well as another test class that demonstrates the testing style you want it to emulate.
Forget everything else about Copilot, it's worth the cost for this alone. Time writing tests reduced ~80%. They could remove all other functionality and rebrand it as Test Copilot, and they would still get my money.
But even as a beginner--when I started coding some 10 years ago, I'd have killed for something like ChatGPT. Had some programming problem you needed to solve? You better hope that someone wrote something about it online, that it's on StackOverflow or been discussed on some other discussion board. Otherwise, it's up to you to splunk into StackOverflow, wait a week or two, hope you don't get ignored/downvoted, or try to find some IRC channel to post the question to. Comparatively having the ability nowadays to talk to a AI about the most niche programming concepts in your specific use case, in English, without being judged? It's straight up magical to me.
In fact the copilot plugin in PyCharm is awesome to also just write normal text, like for a article.
So I’m thinking someone should be building an automated product manager!
But also, I would never trust a script I threw together in 15 minutes to actually produce real code and solve the problem. All it does is generate the text I tell it to. It can't understand the system or how it works, it just procedurally spits out text.
In the same way that a dozen lines of Python cannot understand your program, an LLM is also fundamentally incapable of understanding it. That's the crucial part of programming. An LLM will give you text all day, but it can't write your program.
Sure an LLM can produce trivial scripts, but in my experience so far, it can only reliably generate trivial programs that I could write blindfolded in the same amount of time.
If we just treat these tools like the text generators they are instead of insisting they're code generators, we'll all be better off.
Me: <prompt 1: modify this function>
AI assistant (either ChatGPT or Copilot-X): <attempt 1>
Me: <feedback>
AI assistant: <attempt 2>, fixed with feedback, but deleted something crucial from attempt 1, for no reason at all
It keeps me employed, and even increases my rate quite a bit.
At lot of the confabulations we see today - non-existent variables, imports and APIs, out-of-scope variables, etc - would seem (to me) to be meaningfully addressable with these techniques.
Relatedly, I have gotten surprisingly great mileage out of treating confabulations, in the first instance, as a user (ie, my own) failure to adequately contextualise, rather than error.
In a sense, CFGs give you sharper tools to do that contextualisation. I wonder how far the “sharper tools” approach will carry us. It seems, to this interested layman, consistent with Shannon’s work in statistical language modelling.[4]
The term “prompt engineering” implies, to me, a focus on the instructive aspects of prompting, which is a bit unfortunate but also wholly consistent with the way I see colleagues and friends trying to interact with LLMs. Perhaps we should instead call it “context composition” or something, to emphasise the constructive nature.
[1] https://github.com/outlines-dev/outlines
[2] https://github.com/ggerganov/llama.cpp/pull/1773
[3] https://www.reddit.com/r/LocalLLaMA/comments/156gu8c/d_const...
[4] https://hedgehogreview.com/issues/markets-and-the-good/artic...
If you're already good at this stuff, you'll find the risk of coding assistants getting things wrong is pretty minimal for you.
If you have bad habits where you frequently write and commit code without first executing it and trying to poke holes in it, you'll find AI assisted coding full of traps.
... and you can use it to run a C compiler too, if you know what you are doing! https://simonwillison.net/2023/Oct/17/open-questions/#open-q...
And you can do fun things like point it at an openapi schema and asking it some questions about that API. I've given it screenshots of websites and asked it to criticize the design or document what is visible. It's amazingly good at supporting localization work.
I was working with some geospatial code for an algorithm that generates UTM coordinates from GPS coordinates a few weeks ago. I needed a Kotlin implementation that I could use on multiple of its platforms (js and jvm, i.e. no java dependencies allowed).
It was kind of useless writing the code for me for this (that's the first thing I tried obviously) but it unblocked me a couple of times and helped me figure out some details.
I ultimately found several old Java implementations that contained a lot of undocumented magical numbers and I just asked it the meaning of those numbers (about a dozen) and it came back with good explanations (basically things like wgs84 ellipsoid parameters, radius of the earth, etc.). The code wasn't quite working (it failed a test I wrote) so I asked it to identify possible causes for that and it came up with some uncovered edge cases that I was able to cross confirm with other implementations. In the end I was able to piece together a working implementation by combining different elements from several implementations. Each of them individually had issues. A lot of this code is ancient and there is apparently a lot of copy paste reuse in this space.
Over time, will less skilled programmers produce more critical code? I think so. At some point a jet will fall out of the sky because the coding assistant wasn't correct and the whole profession will have a black eye.
The programmers will be less skilled because the (up until recently) lack of coding assistants provides a more rigorous and thorough training environment. With coding assistants the profession will be less intellectually selective because a wider range of people can do it. I don't think this is good given how critical some code is.
There is another related issue. Studies have shown that use of google maps has resulted in a deterioration of people's ability to navigate. Human mapping and navigation ability needs training which it doesn't get when google maps are used. This sort of thing will be an issue with coding.
I think much the same will happen with regards to programming. Sure, most people will be able to bust out a simple script to do X. But if you want to do a "serious task", you're going to get a professional.
Like, I can't know if the code my pair writes has flaws, just like an AI coding assistant.
I've never learned so much about programming as when pairing. Having someone else to ask or suggest improvements is just invaluable. It's very rarely I learn something new from myself, but another human will always know things I don't, just like an AI.
Of course, you don't blindly accept the code your pair/assistant writes. You have to understand it, ask questions, and most of all write tests that confirms it actually works.
For local navigation, first and foremost. The goal is to teach you how to navigate your locale, so you use it less and less. You still will want to ask it for traffic updates, but you talk to it like you would between locals who know all the roads.
As a model for how to do AI in a way that enhances your thinking rather than softens/replaces it.
Roughly speaking, if you stick to using AI for writing code you could have written yourself, you're okay.
I hear it is also quite useful for doing things which you know extremely well but are tedious to do. Anything in between is certainly a danger zone.
Yeah, if you never knew what the code that got generated did in the first place, that's not gonna apply, but if you're using it as basically just a code expander for things you could do the pseudocode for in your sleep, you're probably gonna be ok.
The AI could do certain things much faster than I would be able to by hand. For instance, in a certain hash table containing structures, it turned out that the deletion algorithm which moves elements couldn't be used because other code retains (reference counted) pointers to the structures, so the addresses of the object have to be stable. AI easily and correctly rewrote the code from "array of struct" to "array of pointer to struct" representation: the declaration of the thing and all the code working on it was correctly adjusted. I can do that myself, but not in two seconds.
E.g. creating a new page in an app.
Feed an LLM the design, the language/framework/component library you're using, and get a page which is 90% of the way there. Some tweaks and you're good.
Far, far quicker than going by hand, and often better quality code than a junior developer.
Now I would never deploy that code without reading it, understanding it, and testing it, but I would always do that anyway. GPT4 is close enough to a good software engineer, that to resist is it to disadvantage your business.
Now if you're coding for pleasure and the fun of creation and creativity, then ditch the LLMs, they take some of that fun away for sure. But most of the time, I'm focused on achieving business outcomes quickly. That's way more productive with a modern LLM
I firmly believe that the worst kind of help is unreliable help. If your help never does its job, then you know you have to do everything yourself. If your help always does its job, you know you can trust it. If your help only sometimes does its job, you get the worst deal of all because you never know what you'll get and have to check every single time.
Unreliable help is still useful if you know that it's unreliable - you learn to keep a critical eye on what it's doing and correct when necessary. Still saves a ton of time, and I'm not guessing that, I'm saying that from my own experience.
Do you feel the same way about a PR from a coworker?
After all, whether a result is correct or not depends on whether it matches the user's desire, so verification criteria must come from that source too.
Sure there are certain types of relatively objective correctness that most of the time will line up with a user's desires, but this kind of verification can never be complete afaict.
The specific way we're training coding assistants for next-token-prediction would also be an incredibly difficult context for humans to produce code.
Suppose you were dropped off in an society of aliens whose perceptual, cultural and cognitive universe is meaningfully different from our own; you don't have a grounding in concepts of what they're trying to _do_ with their programs. You receive a giant dump of reams and reams of source code, in their unfamiliar script, where none of the names initially mean anything to you. In the pile of training material handed to you, you might find some documentation about their programming language, but it's written in their (foreign, weird to you) natural language, and is mixed with everythign else. You never get a teacher who can answer questions, never get access to a IDE/repl/interpreter/debugger/compiler, never get to _run_ a program on different inputs to see its outputs, never get to add a log line to peek at the program's internal state, etc. After a _lot_ of training, you can often predict the next symbol in a program text. But shouldn't we _expect_ you to be "unreliable"? You don't have the ability to run checks against the code you produce! You don't get a warning if you use a variable that doesn't exist! You just produce _tokens_, and get no feedback.
To the degree humans are reliable at coding, it's because we can simulate what program execution will do, with a level of abstraction which we vary in a task dependent way. You can mentally step through every line in a program carefully if you need to. But you can also mentally choose to trust some abstraction and skip steps which you infer cannot be related to some attribute or condition of interest if that abstraction is upheld. The most important parts of your attention are on _what the program does_. This is fully hidden in the next-token-prediction scenario, which is totally focused on _what tokens are used to write the program_.
This difference is even more stark when it comes to driving assistants. Video compilations of Teslas with FSD behaving erratically and most importantly, unpredictably, are all over the place. Experienced Tesla drivers seem to have some limited ability to predict the weaknesses of the FSD package, but the issue is that the driving assistant is so unlike a human. I've seen multiple examples of people saying "well, humans cause car crashes too," but the key difference is that I have to sit behind the wheel and deal with the fact that my driving assistant may or may not suddenly swerve into oncoming traffic. The reasons for it doing so are likely obscure to me, and this is a real problem.
However, there's a lot we can just throw humans at and trust that the thing will get complete and be correct. And even with the checks and balances, we can have the human perform those checks and balances. A human is a pretty autonomous unit on average.
So far, AI can't really say "let me double check that for you" for instance. You ask it a thing, it does a thing, and that's it. If it's wrong, you have to tell it to do the thing again, but differently.
In all the rush to paint these LLMs as "pretty much human", we've instead taken to severely downplaying just how adaptable and clever most sentient beings can be.
In any case, the point is that we have learned techniques to compensate for human fallibility. We will learn techniques to compensate for gen AI fallibility, as well. The objection that AIs can be wrong is far less a barrier to the rise of their utility than is often supposed.
The original argument put forth was that "an inherently unreliable tool cannot gauge its own reliability".
Someone responded that humans fit that description as well.
I said we don't. We can and do verify our own reliability. We can essentially test our assumptions against the real world and fix them.
You then claimed we couldn't do that "self-sufficiently".
I responded that while that is true for some tasks, for a lot of tasks, we can. That an AI can't check itself and won't even try.
And now you're telling me that they can check against each other.
But if you can't trust the originals, asking them if they trust each other is kind of pointless. You're not really doing anything more than adding another layer.
For example: If I put my pants on backwards, I correct that myself. Without the need to check and/or balance against any other person. I am self-correcting to a large degree. The AI would not even know to check its pants until someone told it it was wrong.
The objection isn't that "AIs can be wrong", the objection is that AIs can't really tell the difference between correct and incorrect. So everything has to be checked, often with as much effort as it would take to do the thing in the first place.
Your objections seem to rely on a restricted view that says "we can't do better" but with no evidence. Whereas we have plenty of evidence of massive, continual improvement in the very areas you are holding up as problematic.
One of the well-established techniques was always to give more worthwhile incentives that are especially meaningful to the hardworking assistants.