I expect the models will continue improving though, I feel like most of it comes down to the ephemeral nature of their context window / the ability to recall and attach relevant information to the working context when prompted.
I expect the models will continue improving though, I feel like most of it comes down to the ephemeral nature of their context window / the ability to recall and attach relevant information to the working context when prompted.
I don't think it's that simple.
From what I've found, there are "attractors" in the statistics. If a part of your problem is too similar to a very common problem, that the LLM saw a million times, the output will be attracted to those overwhelming statistical next-words, which is understandable. That is the problem I run into most often.
They're rather impressive when building common things in common ways, and a LOT of programming does fit that. But once you step outside that they feel like a pretty strong net negative - some occasional positive surprises, but lots of easy-to-miss mistakes.
I do a lot of interviews, and the poor performers usually end up running out of working memory and start behaving very similar to an LLM. Corrections/input from me will go into one ear and fall out the other, they'll start hallucinating aspects of the problem statement in an attractor sort of way, they'll get stuck in loops, etc. I write down when this happens in my notes, and it's very consistently 15 minutes. For all of them, it seems to be the lack of familiarity doesn't allow them to compress/compartmentalize the problem into something that fits in their head. I suspect it's similar for the LLM.
> I expect the models will continue improving though
I try to push back on this every time I see it as an excuse for current model behaviour, because what if they don't? Like, not appreciably enough to make a real difference? What if this is just a fundamental problem that remains with this class of AI?
Sure, we've seen incredible improvements over a short period of time in model capability, but those improvements have been visibly slowing down, and models have gotten much more expensive to train. Not to mention that a lot of the problem issues mentioned in this list are problems that these models have had for several generations now, and haven't gotten appreciably better, even while other model capabilities have.
I'm saying this not to criticize you, but more to draw attention to our tendency to handwave away LLM problems with a nebulous "but they'll get better so they won't be a problem." We don't actually know that, so we should factor that uncertainly into our analysis, not dismiss it as is commonly done.
If I ask Claude to do a basic operation on all files in my codebase it won't do it. Half way through it will get distracted and do something else or simply change the operation. No junior programmer will ever do this. And similar for the other examples in the blog.
I can do the rest myself because I'm not a dribbling moron.
Not sure exactly how you used Claude for this, but maybe try doing this in Cursor (which also uses Claude by default)?
I have had pretty good luck with it "reasoning" about the entire codebase of a small-ish webapp.
Write me a parser in R for nginx logs for kubernetes that loads a log file into a tibble.
Fucks sake not normal nginx logs. nginx-ingress.
Use tidyverse. Why are you using base R? No one does that any more.
Why the hell are you writing a regex? It doesn't handle square brackets and the format you're using is wrong. Use the function read_log instead.
No don't write a function called read_log. Use the one from readr you drunk ass piece of shit.
Ok now we're getting somewhere. Now label all the columns by the fields in original nginx format properly.
What the fuck? What have you done! Fuck you I'm going to just do it myself.
... 5 minutes later I did a better job ...
I expected the lack of breadth from the junior, actually.
To be fair the guys I get are pretty good and actually learn. The model doesn't. I have to have the same arguments over and over again with the model. Then I have to retain what arguments I had last time. Then when they update the model it comes up with new stupid things I have to argue with it on.
Net loss for me. I have no idea how people are finding these things productive unless they really don't know or care what garbage comes out.
Core issue. LLMs never ever leave their base level unless you actively modify the prompt. I suppose you _could_ use finetuning to whip it into a useful shape, but that's a lot of work. (https://arxiv.org/pdf/2308.09895 is a good read)
But the flip side of that core issue is that if the base level is high, they're good. Which means for Python & JS, they're pretty darn good. Making pandas garbage work? Just the task for an LLM.
But yeah, R & nginx is not a major part of their original training data, and so they're stuck at "no clue, whatever stackoverflow on similar keywords said".
Not sure if you’re being figurative, but if what you wrote in your first comment is indicative of the tone with which you prompt the LLM, then I’m not surprised you get terrible results. Swearing at the model doesn’t help it produce better code. The model isn’t going to be intimidated by you or worried about losing their job—which I bet your junior engineers are.
Ultimately, prompting LLMs is simply a matter of writing well. Some people seem to write prompts like flippant Slack messages, expecting the LLM to somehow have a dialogue with you to clarify your poorly-framed, half-assed requirement statements. That’s just not how they work. Specify what you actually want and they can execute on that. Why do you expect the LLM to read your mind and know the shape of nginx logs vs nginx-ingress logs? Why not provide an example in the prompt?
It’s odd—I go out of my way to “treat” the LLMs with respect, and find myself feeling an emotional reaction when others write to them with lots of negativity. Not sure what to make of that.
I expect I'd have to hand feed them steps, at which point I imagine the LLM will also do much better.
These problems are getting solved as LLMs improve in terms of context length and having the tools send the LLM all the information it needs.
We have to stop trying to compare them to a human, because they are alien. They make mistakes humans wouldn't, and they complete very difficult tasks that would be tedious and difficult for humans. All in the same output.
I'm net-positive from using AI, though. It can definitely remove a lot of tedium.
How? They've already been trained on all the code in the world at this point, so that's a dead end.
The only other option I see is increasing the context window, which has diminishing returns already (double the window for a 10% increase in accuracy, for example).
We're in a local maxima here.
I didn't say they weren't improving.
I said there's diminishing returns.
There's been more effort put into LLMs in the last two years than in the two years prior, but the gains in the last two years have been much much smaller than in the two years prior.
That's what I meant by diminishing returns: the gains we see are not proportional to the effort invested.