ChatGPT is better at generating code for problems written before 2021
ieeexplore.ieee.org
ieeexplore.ieee.org
(1) ChatGPT is better at generating functionally correct code for problems before 2021 in different languages than problems after 2021 with 48% advantage in Accepted rate on judgment platform, but ChatGPT 's ability to directly fix erroneous code with multi-round fixing process to achieve correct functionality is relatively weak;
(2) the distribution of cyclomatic and cognitive complexity levels for code snippets in different languages varies. Furthermore, the multi-round fixing process with ChatGPT generally preserves or increases the complexity levels of code snippets;
(3) in algorithm scenarios with languages of C, C++, and Java, and CWE scenarios with languages of C and Python3, the code generated by ChatGPT has relevant vulnerabilities. However, the multi-round fixing process for vulnerable code snippets demonstrates promising results, with more than 89% of vulnerabilities successfully addressed; and
(4) code generation may be affected by ChatGPT 's non-determinism factor, resulting in variations of code snippets in functional correctness, complexity, and security. Overall, our findings uncover potential issues and limitations that arise in the ChatGPT -based code generation and lay the groundwork for improving AI and LLM-based code generation techniques.
Aside from the very important fact that GPT-3.5 is still far and away the most frequently used LLM model, it's not like GPT-4 has a completely different architecture with completely different characteristics. It's clearly better, but much of what they describe should generalize to LLMs as a whole (for example, knowledge cutoff dates matter a lot and these things likely have memorized a lot more than we thought they did).
While currently this is a relatively weak strength of genAI - assuming technological improvement of this technique over time, isn't it just as possible that the data quality will converge positively rather than negatively over time, in the future? That is to say, the web would be consistently "refined" as time goes on, by predominant VLLM?
Assuming that the internet is even "filled" as you say in the first place (personally I don't think organically-generated content is ever going to be pushed out of the internet, but that's my opinion, and I'll entertain the opposite case for the sake of the discussion). It also assumes that people are using models trained on the current state of internet "slurry" in the first place - that we are continually ingesting more internet YoY into these models. If we come up with a better model that needs less data to produce high-quality content, neither my nor your assertion is even relevant. Same case if the internet just decides to use small, low-quality models trained on only a portion of the internet.
But if the internet is continually recycling the entirety of itself through a model that has tens of millions of dollars of funding and research focused on directly improving the quality of it's answer metrics, it's not necessarily 100% locked into a downward quality convergence slide. Especially if we assert that humans /will/ continue to be consistently putting more organic data into the internet over time. It's a pessimistic take.
We never used to have such an elaborate or accessible thing before. Kids and senior citizens likely engage constantly without even realizing it - and adopt one another's, idk the words, but I'm going for something akin to dialects/accents/language/opinions/etc
We're all a bunch of parrots
The result shows that the LLM has a hard time generalizing to new (basic!) problems that were not in its training set; this suggests it is memorizing solutions rather than understanding the mechanisms... It's the undergrad who has crammed Leetcode solutions for two weeks before their interview, rather than a student with a deeper understanding who can successfully address new problems.
Having a friend who is objectively pretty dumb but has memorized the entire existing literature of your field of study is still pretty damned useful, however. You just need to understand what kinds of questions you can trust them with.
Along these lines, my own best use of LLMs is looking up jargon and old results. I see some complicated problem, propose a method for attacking it, and ask my friend who has read every stats paper ever digitized whether they have seen this kind of analysis before, what it's called, why it's a bad idea, and what people do instead. The answers point me to relevant literature by giving me the right jargon and method names to search for using Good-Old-Fashioned-Google.
They just have a particular set of strengths and weaknesses, same as any tool. Figuring out their strengths and limitations is how you use them well. And dismissing them for not fully solving AGI is short-sighted.
I am once again reminded of Searle's "Chinese room" thought experiment [0] in which it is argued that computers cannot understand Chinese, merely memorize and regurgitate a canned set of responses to given prompts, and execute procedures by rote; that their computations cannot be "about" a subject matter (such as Chinese, or programming) the same way the operation and contents of our minds are "about" something.
Are they really publishing a paper based on GPT3.5 in July 2024? I am not sure these results are relevant in any way today.
Edit: Just for reference. The best model for coding today (according to most benchmarks) is Claude-3.5-Sonnet which is freely accessible. Also GPT-4o is freely accessible and is still vastly better than GPT-3.5.
The lm sys arena coding leaderboard (https://chat.lmsys.org/?leaderboard) lists sonnet-3.5 and gpt-4o jointly on #1 and GPT-3.5-Turbo on #35. You can freely download and run LLMs locally on your machine that are significantly better than GPT-3.5, for example Mistral Codestral.
There is really no reason to accept any results on GPT3.5 for relevant today. This is as if you were complaining that a computer from the 00ies is not running <recent operating system> well.
A publication based on an LLM that was state-of-the art only until March of 2023 cannot be justified by long review times.
Edit: To be fair, it seems their preprint was first submitted in August 2023 and the IEEE article that is based on the paper was a bit slow...
> Thus, in this study, we take the state- of-the-art ChatGPT (the default version of GPT-3.5), the recent popular product, as the representative of LLMs for evaluation.
it rubs the lotion on its skin, ... now it places the lotion in the basket
More discussion on blog post: https://news.ycombinator.com/item?id=40897958
HN is not a curated map of links to discussions, it's a forum, a set of link-discussion pairs. If a discussion fails to take off and receive "significant attention" [0] there's nothing wrong with resubmitting it, much as that may offend your sensibilities.