Performance of LLMs on Advent of Code 2024
jerpint.io
jerpint.io
When I ask them to code things that they never heard of (I am working on a online sport game), it fails catastrophically. The LLM should know the sport, and what I ask is pretty clear for anyone who understand the game (I tested against actual people and it was obvious what to expect), but the LLM failed miserably. Even worse when I ask them to write some designs in CSS for the game. It seems if you take them outside the 3-columsn layout or bootstrap or the overused landing page, LLMs fails miserably.
It works very well for the known cases, but as soon as you want them to do something original, they just can't.
We are doing a large EU healthcare project and things differ per country obviously; if I would assume people to have a modicum of world knowledge or interest in it and look it up, nothing would get done. It is easier to deliver excel sheets with the proper wording and formulas and leave out what it is even for.
Works fine with LLMs.
Disclaimer; the people in my company know everything about the subject matter; the people (and LLMs) that implement it are usually not ours: we just manage them and in my experience, programmers at big companies are borderline worthless, hence, in the past 35 years, I have taken to write tasks as concrete as I can; we found that otherwise we either get garbage commits or tons of meetings with questions and then garbage commits. And now that comes in handy as this works much better on LLMs too.
I'd expect more senior engineers to handle increasingly complex/ambiguous projects, while taking the business context into account as necessary.
Maybe we were mostly actually doing the exact same stuff and these things are trained on a whole lot of the same. Like with react components, these things are amazing at pumping out the most basic varieties of them.
But then there are those idiosyncrasies that every company has, a finicky build system, strange hacks to get certain things to work, extensive webs of tribal knowledge, that will be the demise for any LLM SWE application besides being a tool for skilled professionals.
There’s not much to ponder on here as the majority of human thought and action is unoriginal and irrelevant, so software development is not some special exception to this. This doesn’t mean it lacks meaning or value for those doing it, since that can be self-generated, extrinsically motivated, or both. But, by definition, most things done are within the boundaries of a normal distribution and not exceptional, transcendent, or sublime. You can strive for those latter categories, but it’s a path strewn with much failure, misery, and quite questionable rewards, even where it is achieved. Most prefer more achievable comforts or security to it, or, at most, dull-minded imitations thereof, such as amassing great wealth.
If you think about it, lots of professions are already divided in two. You often have profession of inventors / designers and a profession of creators. For example: Electrical engineers vs Electricians. Architects / civil engineers and construction workers. Chefs and line cooks. Composers and session musicians. And so on.
The two professions always need somewhat different skill sets. For example, electricians need business sense and to be personable - while electrical engineers need to understand how to design circuits from scratch. Session musicians need a fantastic memory for music, and to be able to perform in lots of different musical styles. Composers need a deeper understanding of how to write compelling music - but usually only in their preferred style.
I think programming should be divided in the same way. It sort of happens already: There's people who invent novel database storage engines, work on programming language design, write schedulers for operating systems and so on. And then there's people who do consulting and write almost all the software that businesses actually need to function. Eg, people who talk to actual clients and figure out their needs.
The first group invents React. The second group uses react to build online stores.
We sort of have these two professions already. We just don't have language to separate them.
The line is a bit fuzzy, but you can tell they're two professions because each needs a slightly different skill set. If you're writing a novel full-text search engine, you need to be really good at data structures & algorithms, but you don't really need to be good at working with clients. And if you work in frontend react all day, you don't really need to understand how a B-Tree works or understand the difference between LR and recursive descent parsing. (And it doesn't make much sense to ask these sort of programmers "leetcode problems" in job interviews).
ChatGPT is good at application programming. Or really, anything that its seen lots of times before. But the more novel a problem - the more its actual "software engineering" - the more it struggles to make working software. "Write a react component that does X" - easy. "Implement a B-tree with this weird quirk" - hard. It'll struggle through but the result will be buggy. "Implement this new algorithm in a paper I just read" - its more or less completely useless at that.
Meanwhile people who work on things that are even slightly niche or complex (ie. Things which aren't answered in 20 different stack overflow threads, or things which require of few field "experts" to meet around a table and think hard for a few hours) don't understand the enthusiasm at all
I'm working on fairly simple things but in a context/industry that isn't that common and I have to spoon feed LLMs to get anywhere, most of the time it doesn't even work. The next few years will be interesting, it might play out as it did for plane pilots over relying on autopilots to the point of losing key skills
I'm currently increasingly often confronted with overly bloated codebases that, with the added speed of llms / coding assistants, have reached levels of unmaintainability much faster than traditionally possible.
Now, fixing those repos with LLMs fails miserably in my experience, so you have to invest hundreds of hours of senior engineer time to make the projects workable again, regardless of the novelty of the problem solved by that codebase.
It also frequently makes the engineers pulled in to help miserable.
An LLM is, famously, a "stochastic parrot". It has no reasoning. It has no creativity. It has no basis or mechanism for true originality: it is regurgitating what it has seen before, based on probabilities, nothing more. It looks more impressive, but that's all is behind the curtain.
It surprises me that so many people expect an LLM to do something "clever" or "original". They're compression machines for knowledge, a tool for summarisation, rephrasing, not for creation.
Why is this not that widely known and accepted across the industry yet, and why do so many people have such high expectations?
To me this proves that llms learn concepts and multilayeres representations within its network and not just some dumb statistical inference. Even famous llm skeptic like Francois Chollet doesn't invoke this stochastic parrots anymore and has moved on to arguing that they don't generalize well and are just memorizing.
With GPT-2 and 3 I was of the same opinion that they seemed like just sophisticated stochastic parrots, but current llms are a different class from early gpts. Now that o3 has beaten memory resisting benchmark like ARC-AGI I think we can confidently move on from this stochastic parrots notion.
(And before you argue that o3 is not an llm, here is an openai researcher stating that it is an llm https://x.com/__nmca__/status/1870170101091008860?s=46&t=eTe... )
Biological networks of more intelligent species contain a few billion neurons and upwards from that while even the big LLMs are somewhere in the millions of "equivalent" at best. So, bad topology + much less "neurons" and the resulting capabilities shouldn't be too surprising. Plus it is clear that AGI is nowhere close, because one result of AGI is a proper understanding of "I". Crows have an understanding of "I", for example.
And that is where these "meta abstraction levels" come in: There are many needed to eventually reach the stage of "I". This can also be used to test how well neural networks perform, how far they can abstract things for real, how many levels of generalization are handled by it. But therein lies a problem: Let 2 persons specify abstraction levels and the results will be all across the board. This is also why ARC-AGI, while dives into that, cannot really solve the problem of testing "AI", let alone "AGI": We as humans are currently unable to properly test intelligence in any meaningful way. All the tests we have are mere glimpses into it and often (complex) multivariable + multi(abstraction)layered tests and dealing with the results, consequently, a total mess, even if we throw big formulas at it.
However, out of distribution generation and generalization are a better more useful metric. In these Yann Lecun has argued that interpolation is meaningless in the high dimensional spaces these "curves" are embedded in.(https://arxiv.org/abs/2110.09485)
ARC has proven they can generalize, and Alpha Go (not an llm but a deep network) has proven it can generate novel/creative solutions. We don't need AI to have a sense of self "I" for them to beat us at every human skill and activity. Infact it might be detrimental and not useful for us if AI developed a sense of self.
Here's a much better analysis from someone who got 45 stars using LLMs. https://www.reddit.com/r/adventofcode/comments/1hnk1c5/resul...
All the top 5 players on the final leaderboard https://adventofcode.com/2024/leaderboard used LLMs for most of their solutions.
LLMs can solve all days except 12, 15, 17, 21, and 24
"I used gpt-4o with zero shot prompts and it failed terribly!"
"I used Claude/o1/o3, I fed various bits of information into the context and carefully guided the LLM"
Those two approaches (there are many more) would lead to very different results, yet all we read here are comments giving their opinions on "LLMs", as if there is only one LLM and one way of working with them.
stone->your ingredients->soup
LLM->your prompts->solution
“Some other models tested that just didn't work: gpt-4o, gpt-o1, qwen qwq.”
Notably gpt-4o was used in the post linked here.
If the top 5 people on leader board, 'used LLM's'. Meaning, they used an LLM as a helping tool.
But article I think is, what if the LLM played by itself. Just paste the questions in, and see if it can do it all on its own?
Perhaps that is the difference, different goals.
Note that the leaderboard points are given out on the time taken to produce a correct answer. 100 points for the first to submit an answer, 99 for the second, 98 for third and so on. No points if you're not in the first 100.
So if an LLM fails on 5 problems, but for the other 20 it can take a 600-word problem statement and solve it in 12 seconds? It'll rack up loads of points.
Whereas the best human programmer you know might be able to solve all 25 problems taking around 15 minutes per problem. On most days they would have zero points, as all 100 points finishes go to LLMs.
That said, if I cared only to produce the best results I would 100% pair up with an LLM
Source for this?
The author could make a model on huggingface routing requests to him. Latency might vary.
Once we use tools to easily iterate on code (e.g. generate, compile, test, use outcome to refine prompt) we will turbocharge LLMs coding abilities.
It's a reasonable test.
When it couldn't solve it within the constraints, was it 'broadly correct' with some bugs? Was it on the right track or completely off?
Glancing through the code for some of the 'easier' problems Claude (his best performer), it seems to have broadly correct (if a little strange and overwrought) code.
But my phone is not a good platform for me to do this kind of reading.
A better implication would be that proper use of LLMs implies more than OP did here (proper prompting, looping the answer w/ a code check, etc.)
Anyway, I’m not at all surprised that they can handle AoC. If anything I would worry that AoC will still be a fun project to author when many people solve it with AI. It sure won’t be fun to curate any form of leaderboard.
But I’m not surprised because I’ve seen them fail on many problems even with lots of prompt engineering and test cases.
And so it is with success of LLMs in one-shot challenges and any job that depends on such challenges: cooked a long time ago.
For most coding problems, people won't magically know if the code is "correct" or not, so any one-shot answer that is wrong could be a real hindrance.
I don't have time to prompt engineer for every bit of code. I need tools that accelerate my work, not tools that absorb my time.
o1 got 20 out of 25 (or 19 out of 24, depending on how you want to count). Unclear experimental setup (it's not obvious how much it was prompted), but it seems to check out with leaderboard times, where problems solvable with LLMs had clear times flat out impossible for humans.
An agent-type setup using Claude got 14 out of 25 (or, again, 13/24)
Anyhow this is like night and day compared to last year, and it's impressive that Sonnet is now apparently 50% as good as a professional human at this sort of thing.
My gut tells me you can get much better results from the models with better prompting. The whole "You are solving the 2024 advent of code challenge." form of prompting is just adding noise with no real value. Based on my empirical experience, that likely hurts performance instead of helping.
The time limit feels arbitrary and adds nothing to the benchmark. I don't understand why you wouldn't include o1 in the list of models.
There's just a lot here that doesn't feel very scientific about this analysis...
https://huggingface.co/spaces/jerpint/advent24-llm/tree/main
There's zero mention of how long it took the LLM to write the code vs the human. You have a 300 second runtime limit, but what was your coding time limit? The machine spat out code in, what, a few seconds? And how long did your solutions take to write?
Advent of code problems take me longer to just read than it takes an LLM to have a proposed solution ready for evaluation.
> they didn’t perform nearly as well as I’d expect
Is this a joke, though? A machine takes a problem description written as floridly hyperventilated as advent problems are, and, without any opportunity for automated reanalysis, it understands the exact problem domain, it understands exactly what's being asked, correctly models the solution, and spits out a correct single-shot solution on 20 of them in no time flat, often with substantially better running time than your own solutions, and that's disappointing?
> a lot of the submissions had timeout errors, which means that their solutions might work if asked more explicitly for efficient solutions. However the models should know very well what AoC solutions entail
You made up an arbitrary runtime limit and then kept that limit a secret, and you were surprised when the solutions didn't adhere to the secret limit?
> Finally, some of the submissions raised some Exceptions, which would likely be fixed with a human reviewing this code and asking for changes.
How many of your solutions got the correct answer on the first try without going back and fixing something?
In the same way that I don't consider the time an author spends writing a book when saying how long it took me to read the book. OP lost zero time training ChatGPT.
Or do you mean how much time OP spent training themself? Because that's a whole new can of worms. How many years should we add to OP's development times because probably they learned how to write code before this exercise?
Would you also count the number of hours it took for someone to write the OS, the editor, etc? The programmer wouldn't be effective without those things.
Not the OP, but I was able to zero shot 27 solutions correctly without any back and forth, and 5 more with a little bit back and forth. Using local models.
It does not "understand" or "model" shit. Good grief you AI shills need to take a breath.
1. You could have pointed out the grammatical errors and explained why they matter.
2. You could have pointed out the structural errors, and explained what structure should have been used.
3. You could have written a new prompt.
4. You could have re-run the experiment with the new prompt.
Otherwise what you say is just unsubstantiated criticism.