ChatGPT does Advent of Code 2023
themotte.org
themotte.org
Depends what you are benchmarking for... If you are benchmarking the ability of the solution to solve LEETCODE challenges, that is different to the ability of GPT4 to assist everyday programmers knock out business logic or diagnose bugs.
My experience of GPT4 is that it's significantly better at the latter than GPT3.5.
Additionally, the real test is for me is "Can an average programmer using GPT4 as a tool solve Advent of Code faster than an equally-skilled programmer without an LLM?".
Literally says in the article that GPT's main drawback is that it can't debug in the same way a human can.
When stuck I often paste code into ChatGPT and ask why it doesn't work, and it will often help me quickly identify the error and propose a fix.
One example are bugs caused by precondition violations, which ChatGPT can’t diagnose without also being given the code to all of the incoming call-sites, which means you end-up solving the problem yourself before you’ve even explained the issue to ChatGPT - so (to me, at least) my use of ChatGPT is more akin to rubber-duck-debugging[1] than anything else.
yeah, we run the random number generator again and hopefully this time less buggy code pops out
I mean, yes, obviously? Any tool, even if it had near-zero marginal utility must necessarily improve performance, as you can simply elect not to use it.
The disagreement is not in the direction, it's in the magnitude. Therefore the real test is: "Can an average programmer solve AoC faster with GPT4 and without syntax coloring or with syntax coloring and without GPT4?", or "GPT4 in C versus no-GPT4 in Python", or "GPT4 on a crappy laptop vs no-GPT4 on a high-end workstation" and so on.
That might be true for any tool with consistent, predictable output that you thoroughly understand. Someone could easily lose a few hours dickering around with chatgpt only to realize they’re better off just starting it from scratch.
No opinion on the specific tradeoff here, but in general it's not obvious to me that your statement is true. Using tools involves opportunity cost, so sometimes "not electing to us it" can on average be a net win, no? There is a cognitive and time load to using it at all, right?
A simpler test would be keeping the development environment the same and then adding GPT4 to see if there is a statistically significant and meaningful speed increase (of a decent magnitude).
I'm not looking for a 1% speed improvement, I'm looking for a >50% speed improvement. Maybe I should have stated 'significant' speed improvement in the initial post.
It seems like you're accepting that (at least in its current state, barring unknowable future discoveries) this is a technology that complements developers, making them more productive.
If this is the case, then the question is how relevant this improvement is in quantitative terms. The cumulative improvements in developer productivity since the days of punching cards have easily been several hundred percent.
A thought experiment to figure out how relatively impactful this is would be to compare it to other technologies and see which gives the greatest boost.
My prior belief: somewhere around Intellisense level of useful, but not significantly more than that.
Although my personal belief is this is more likely going to be an improvement similar to Assembly -> C (even if the LLM component doesn't improve, but assuming that the tooling does improve).
Personally I think there is going to be a new 'higher, higher level' of programming paradigms that are about to be invented that are supported by the ability of LLM's to write code - still augmenting humans, but making them many times more productive.
So the "LLM language" already exists - its Javascript, Java, C#, etc.
I'm talking about a higher-level language that outputs Javascript, Java, C# etc (but presumably some sort of lower level byte-code in the future).
A language where intent and clarity is more important than syntax or implementation details.
Where
menu_options = [process['name'] for process in processes if process['active']]
is more clearly expressed as
menu_options = list of the active process names
(or something similar)
There are a lot of anecdotes about how "it's the programmer, not the language" or "any language can be productive", etc...
The actual results were that many popular languages are more than an order of magnitude slower to achieve time-to-correct-solution than others. From memory, the fastest was F# with about 20 minutes, then Python and C# with about 40 minutes for both, and then C/C++/Java were hours to even days!
Well, it's going to need to rewrite functions to add debug due to not having edit capabilities, but I tried this and it absolutely added debug info which it then used to debug issues:
In an inner part it added:
debug_info.append({
'Hand': hand,
'Bid': bid,
'Type': hand_info[0],
'Sorted Hand': hand_info[1],
'Rank': rank,
'Score': score
})
I don't know quite what's happening but I feel like people constantly say it can't do something and the very first thing I try (just asking it to do the thing) usually works.I gave it the hands in the problem statement and the expected result, and the explanation as to why (copy pasted). It ran the code, looked at the debug output, identified the problem and rewrote the function. I'm not saying it immediately solved the problem, but it easily added debug information, ran the code, looked at the output and interpreted it.
I didn’t have the solutions generated by ChatGPT to be clear, I used a prompt along the lines of “take this text and extract it’s requirements and generate bullet points, also make the example inputs and outputs clear”.
I did the same for the previous year too, when ChatGPT hadn’t been out for long.
I find that a lot more enjoyable and less tedious.
Start at the end and work my way backwards until the conclusion makes sense.
If you subscribe to ChatGPT+ you can just ask GPT-4 to verify the approach by doing a web search, that often works quite well.
It wouldn't be notable if someone specifically asked ChatGPT, knowing its limitation, but using it to automatically populate Quora and Google with it is pretty bad. People are using LLM to fill the web with BS.
[1] https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3073417/
[2] https://www.quora.com/Why-does-drinking-olive-oil-burn-the-b...
I have no idea why in 2024 people are still lumping all LLMs together, as if they're all identical clones of the same thing.
It's like saying vehicular travel will never work because you didn't like the handling of the cheapest car you can buy.
One is dumb as a brick, then other isn't. If you don't specify, then your comment should be dismissed out of hand.
Also, it's a well-known limitation of all current LLMs that they're terrible at basic algebra. Instead of trying to replace BLAS or Maple with it, ask it to write the Python code or produce the Mathematica expression.
This creator has some excellent videos of chat GPT attempting advent of code.
He uses it in a more generous format where he is often giving it multiple attempts and trying to coax the correct answer out of it. He definitely has more success with it than the article, but it is hard to tell how much of that success is due generous assistance and prompting, and so it is hard to know how much the model has actually improved year over year.
My main criticism of learning about algorithms using a framework is that they handhold so much with optimisation techniques and accessible APIs that you can know very little about what's actually happening and build something that works.
When writing AI algorithms just knowing how to implement in theory is rarely ever enough. The difficulty is always in the debugging and optimisation techniques.
I notice that for "novel" code that CoPilot hasn't seen before, it's mostly useless, but when writing a hobby project which is a path tracer (of which there are probably 1000 implementations on GitHub) it's excellent. Which isn't surprising. It has seen the exact same function I'm writing, written 100 times in every language imaginable. There are books on the topic etc.
// convert color from linear to sRGB
let srgb = [some formula]
// calculate wavelength dependent index of refraction
let nw = [something]
But yes, obviously somewhere between this and just saying "Create all the code I need" All the fun would disappear too. For the most part, that wouldn't work either. It would shit out 200 lines of code that doesn't work, and now it instead of taking the boring job of googling a formula from an optics reference replacing the variable names, it's now taking the role of shitty developer and made me the code reviewer. But see that's my day job and it's not what anyone wants.
It is also quite possible that there are fundamental limits to scaling LLMs that we have not discovered yet. You won't find a way to let airplanes fly to the moon no matter how many resources you pour into it. To make it there, you need better methods (ie rockets in this case) that look nothing like the first few steps to get off the ground.
By that I mean try to solve the Day 1 problem at the point of release, then try to solve it in a fresh ChatGPT session the next day, and then get it to solve it again the day after that, and so for the next couple of months.
Do that for each day that it runs for and then see what patterns emerge. Part of me expects it to get better at solving each problem as the rest of us write our solutions and then make them available on the web in some form for it harvest - but it would be interesting to see if that is the case.
Kind of like saying "let's watch the same youtube video every day and see if it changes" - it won't, that's not how youtube works, there is no point trying.
On the other hand, even move-fast-and-break-things mentality knows that some of December should be in production freeze, emergency-only updates.
Otoh, I can easily imagine helper/censorship models and chat’s system prompt being updated. System prompt doesn’t change capabilities much though, and the censor model behaviour doesn’t change the output, just cuts it off when it discovers copyright or other violations (the case of chat not being able to recite the Dune poem, DALLe not being able to produce certain images etc)
Of course, an LLM will use the solutions that are in its training set. I asked an LLM to explain a pretty generic solution to one of last years AoC questions, and even with no (or very little) reference in the code to the fact it was an AoC solution the LLM used terminology from the question.
It wasn't ChatGPT, i think it was Claude.ai. ChatGPT gave an explanation you would expect from just seeing the code without the context it was related to AoC.
> The model can unquestionably code. [https://news.ycombinator.com/item?id=38205052]
Seems pretty clear most of the fantastic results from before was overfitting. Like every other case. It's amazing to me how these models can be caught red handed being overfitted again and again and again, and people don't get the memo.
The data is out there to objectively assess these tools, but it’s not good for business.
Generative AI is only as good as the dataset that you give it, so for problems that exist in a heavily contrived and parameterized space (like leetcode-style problems) it works really well. But give it a novel problem with intertwining libraries and external dependencies along with custom type and structure definitions, it's going to fall flat.
Humans do something AI can't, and that's draw from experience to apply a solution on a novel problem. This is why I'm not terribly worried about AI coming for engineer jobs anywhere in the near future. I use ChatGPT all the time to write me small functions, generate regular expressions, etc. Basically all the drudgery.
I'd argue we've hit peak-AI at this point because from this point on, all datasets are going to be colored by AI-generated results. Generative-AI is now on a trajectory were it'll simply regress to the mean of knowledge.
I don't know about the rest of your comment, albeit I consider myself more in the non-Chomsky camp. This emergence thing seems a little more than an elaborate hoax to me.
Well, maybe they've just trained GPT-4 to wiggle my balls, when I ask it to analyse a poem that I wrote 15 years ago.
The local llm scene, regardless of this debate, is nevertheless the hottest topic IMHO atm.
Several finetunes have repeatedly blown my mind even in the past 2 weeks. On a Raspberry Pi.
AI can interface. Natural Language Processing. This is the first time in history humaity has such knowledge and technology.
It's basically C-3PO.
Maybe I shouldn't, though I justify it because it allows me to focus on the bespoke bits of whatever I'm doing...
Feels a bit double edged to me still.
Even when I do need some, I use swagger or copier templates or something similar.
If you’re using ChatGPT for scaffolding, I feel you’ve fucked up?
And the above applies to a lot of prod code as well.
Now keep in mind that it had never seen this problem, and it was written in a very terse "competition solving" style with zero comments, one-letter identifiers, and generic function names like "day1" and "parse".
It figured it all out. It worked out that it was for solving a maze -- even though the word "maze" was not used anywhere in the code!
It worked out that a constant "(-1,0)" in one isolated bit of code was "orienting the character west", even though to figure that out it had to trace the logic through 4 or 5 layers of indirection. It connected a single 'w' character in the parser to a vector somewhere else!
Etc...
PS: It wrote a more useful, more accurate, more coherent, and more grammatically correct comment than I had seen in any codebase I had worked on in something like two years.
Allegedly.
I think this wildly overestimates the programming skills of the average CS graduate. My estimate of the fraction of CS graduates able to do that is closer to 1%.
Earning a silver star took less than 1 hour each day for almost everybody on average. The gold star was tougher, between 2-5 on average.
Even people who didn’t complete any of the challenges have told me they plan to work them over the course of the year.
We’re going to make this an annual thing and announce it at the next conference in August, so hopefully we will have more participants next year.
You don't need a CS degree for that.
~5% of those who solved the first 2 problems solved all 49 of them. Given that it is a significant time commitment to spend several hours per day every day for 25 days straight, more people could have done it, given the time. These are not some elite self-selected geniuses. 1/4 of those who solved 1st problem, hasn't solved the 2nd one despite its simplicity https://adventofcode.com/2023/stats
But personal opinion aside nobody provided any real evidence that this indeed the case and only that is what I wanted to point out in my comment.
Everybody who participated seemed to enjoy solving the problems. Many of them have said they plan to solve the rest over the next couple of months, just not on the contest timeline.
I think most CS grad students can solve Advent of Code. Some people, probably don't finish it not because it is hard, but probably because they lose interest.
But this is much different, and likely much worse, than paid ChatGPT. Code Interpreter does its own debugging-and-revising loop.
It's bizarre to me that this author wouldn't pay $20 one time to evaluate the higher quality product, the one most people would use if they cared about code quality, am I missing something?
ChatGPT is a product built on top of gpt-4.
It includes many features that were built on top of the api (not as a a part of it)
So it’s the same model, but the chatgpt product has more features than just text generation
And depending on the subject and/or what you're trying to do, it will have a larger impact the just talking to the model
- ChatGPT pros:
Has some bells and whistles like code interpreter (which I can easily get via Open Interpreter).
Has plugins (although I found web browsing to be inferior compared to Poe/Perplexity).
- Pure GPT-4 API pros:
Is not dumbed down or forced to "forget" things or be lazy in coding.
I use the API either programmatically or through Poe.
They’ll spend weeks, years even, coding a solution in an effort to not pay for it.
Penny wise but pound foolish.
An extreme blindspot.
Every introverted dev should go outside and meet other people, check out the irl marketplaces with 1 coffee for 10$, thats the best investment they can make.
And I yet I still get baited everytime
Edit: apparently not, the author is just really good at coming up with ai adverse puzzles. When testing ChatGPT did much better on last year’s puzzles.
> Here are things LLMs didn't influence:
> The story.
> The puzzles.
> The inputs.
> I don't have a ChatGPT or Bard or whatever account, and I've never even used an LLM to write code or solve a puzzle, so I'm not sure what kinds of puzzles would be good or bad if that were my goal. Fortunately, it's not my goal - my goal is to help people become better programmers, not to create some kind of wacky LLM obstacle course. I'd rather have puzzles that are good for humans than puzzles that are both bad for humans and also somehow make the speed contest LLM-resistant.
> I did the same thing this year that I do every year: I picked 25 puzzle ideas that sounded interesting to me, wrote them up, and then calibrated them based on betatester feedback. If you found a given puzzle easier or harder than you expected, please remember that difficulty is subjective and writing puzzles is tricky.
At least one betatester might very well have been using chatgpt.
Some people have even speculated that the problems this year were deliberately formulated to foil ChatGPT, but Eric actually denied that this is the case.
Citing the author of AoC: Here are things LLMs didn't influence: The story. The puzzles. The inputs.
I did the same thing this year that I do every year: I picked 25 puzzle ideas that sounded interesting to me, wrote them up, and then calibrated them based on betatester feedback.
https://old.reddit.com/r/adventofcode/comments/18bp8id/why_d...It's interesting it turns out this year was not written with gpt in mind.
aka new unique problems that aren't in its training set
A bit among the lines of "if my grandmother had wheels, she would be a bicycle". Emulating neurons doesn't make intelligence.
Sentience isn't even well-defined and I'm not sure we can even point to any unique quality of any individual human as indicative of sentience. We certainly can't agree on what point an individual human becomes sentient, as evidenced by the abortion debate. At best we seem to have some statistical evidence based on what we've achieved as a species being superior to what oak trees have achieved, but at an individual level, I don't know how to prove to anyone that I'm not just a very advanced LLM.
There is another sense where sentience is just being used to mean "that which is human", and by definition, nothing aside from humans will ever qualify as that.
What about animals makes them sentient? Until we can answer that, people are just going to be talking past each other. Even if you forget about AI, whether animals are sentient, and, if so, which animals are sentient, is a big argument in biology/ethics/law that's been going on for at least half a century. The UK recently passed a law that declares animals to be sentient, but not all invertebrates are considered sentient in that law. Their justification is that they don't contain a central nervous system, but what special property about a CNS confers sentience?
There aren't any similarities in the way animal species and LLMs develop.
There have been rumours that the current ChatGPT loses money even at $20/month, and that to economise on running costs they've changed to a less capable model.
And if the modern LLMs have already mined everything there is to get out of pirated ebooks and the common crawl dataset - who knows how long it'll take for them to make the next big step forward?
But synthetic training data has its problems. Oh, it's great if you want a limitless supply of templated high school math problems. In other fields, though? If GPT-whatever is a bit unclear on whether Magnus Hirschfeld was a Nazi or a victim of the Nazis, and you use it to generate synthetic training data - you can't expect the student model to know better than the teacher model does.
Why not? People outgrow their teachers all the time, even in fields where there is no new external data, like maths. Often improved understanding comes just from reflecting on an issue from new perspectives, which synthetic data can provide. Perhaps a teacher/student model helps AI develop those new insights, just as it does for us.
I don't believe anybody is suggesting that we exclusively use synthetic data, but rather that synthetic data can augment other types of training. The other thing to consider is that less sophisticated models can be prone to hallucinating nonsense, but the hallucinations are usually inconsistent, whereas truthful responses tend towards consistency in various directions: between each response, internal consistency, and consistency with reality.
It's conceivable that a more sophisticated model would be able to learn a sense of certainty based on the consistency of its training data, much as we do. If you consider your education, I'm guessing you probably had lots of people tell you incompatible nonsense over the years. In my case, I've had teachers give me a ton of inconsistent explanations about how electrons "know" which path to take in a circuit - probably one of my biggest questions since a child. Only one explanation turned out to be internally consistent and demonstrable with experiments I've seen on YouTube. The result is that I now have, I think, a pretty decent understanding of how electric current forms a path within a circuit, or at least one that can be used to make valid predictions, despite being told a vast amount of inconsistent and wrong information over my life.
ChatGPT discussion here has been totally dominated by this problem from the beginning, but usually it's people raising the bar on what you have to do to get good results from it.
It shows that people fundamentally do not understand the tools they are using.
Given a book of numbers, here are two tasks:
1) copy out the entire book, but replace every prime number with 7.
2) write down the list of prime numbers in the book.
Which one is easier?
LLMs have to generate tokens one at a time, and it’s very very difficult to perfectly generate a set of input tokens except for some tokens.
Since you are almost certainly randomising the probabilities to some degree (that’s what temperature does), you’re also asking for both deterministic and random outputs.
TLDR: ask LLMs what is wrong with the code.
Ask for a diff.
Don’t ask for an LLM to refactor, bug fix or annotate code…
That’s extremely naive usage.
Back to my stupid analogy: “please copy out this book, but fix the numbers which are ‘wrong’”
I can hardly complain when I get terrible results can I?
If you don't understand that an LLM generates output token-by-token, and that as a result of randomizing the token output probabilities that you cannot generate an error free copy of the input you've invested so little time in understanding as to be farcical.
There's 'wow, these are complicated and I don't fully understand them'
...and there's, 'What this. It shiny. Not worky. Make some random change to prompt and pray to LLM gods'
Come on, make an effort.
I've seen an IT professional type this into Google: "Why did my PC crash?"
I couldn't believe that after two decades of using web search technology, he still hadn't figured out how to extract value from a text index using specific and relevant key words.
Similarly, very few developers know what a database index does or whether they need one or not. My pet theory is that NoSQL databases become popular because many of them automatically index every column, making it feel like a magic black box instead of an evil black box.
LLMs are not only black boxes, but they're soooo fundamentally different and new that a lot of people are really struggling to wrap their brains around it.
Just in this thread, today, there are people that are complaining about the direct equivalent of "my poorly thought out Google search didn't work, so Google is bad."
Here's a re-wording of our conversation without that; you decide how you want to take it.
me: I am sad because people are clearly using these tools without understanding them; here is a specific example and reason of why what they're trying to do doesn't work.
you: these tools cannot be understood.
me: not only is that is literally false, it's obviously and self evidently false, and I just gave you a specific example of how; I can't take anything you say as being in good faith when you believe that anything to do with AI is literally unfathomable, and you can't even be bothered responding to what I actually wrote.
Maybe, looking into it, you would find that it's not nearly as complicated or difficult as you imagine.
If not, I guess we have nothing to talk about.
This kind of attitude is why I'm sad.
/shrug
I also felt there were more problems than usual this year that could not easily be solved without looking at the input for special cases not alluded to in the problem descriptions. (As someone who has solved all 25 for the past 3 years).
An extreme example was this year's day 20 circuit-simulating problem, which was made far easier by having the given circuit split up into a few independent chunks that are only connected at the start + end. (I suspect it might be NP-complete without this feature)
It's a slightly different kind of problem solving to think "what makes this particular input easier than the general version of this problem", and one that I'd naively assume LLMs are less skilled at.
Wastl himself denied that this is the case[1]. This is a lie.
> Other tasks required studying the input data
Always been the case for AoC, how is that different from other years?
We have a clear example of a set of tasks that state-of-the-art LLMs cannot perform. We are doing science, for once. Why do we need to get into full conspiracy mode?
[1] https://old.reddit.com/r/adventofcode/comments/18bp8id/why_d...