How Good Is ChatGPT at Coding, Really?
spectrum.ieee.org
spectrum.ieee.org
From my personal experience, Claude 3.5 or GPT-4o work best. They're more coding assistants, not really capable of writing anything more than very simple programs on their own. They make a lot of mistakes, and you need to know how to debug code they produce.
Claude is my favorite but it just randomly added a division by 100,000 to a line of code for no discernible reason. According to Claude it was "an oversight on [Claude's] part".
It seems relevant to mention that the "ChatGPT" in the article isn't the one most of us are using for coding.
We can discuss percentages in this or that task but the article points to the root cause which I think is inescapable.
It seems like we're talking about two different things.
If you understand what ChatGPT is, and what it can and can't do, then there's no problem. It's a useful tool. If you think ChatGPT is AGI or it's going to replace actual programmers soon, then there's a problem.
To me, it sounds like the issue is mostly unrealistic expectations.
In any case, these models are good for simple stuff as far as I can tell, but can't do anything hard or off the beaten path. They can give me ideas for solving hard issues, but any code they generate is killed by hallucinations and general lack of anything resembling contextual knowledge. It can be quite obstinate too, even if I explain I'm trying to solve an unusual problem, and so need out of the box "thinking", 4o will try to force me back to some mainstream solution that I can't use.
I also use Amazon's Q and that is good for simple/repetitive automation but often generates extremely wacky stuff otherwise.
They are all apparently better at Python than anything else. I don't use that so maybe that's an issue.
We knew that already, but it is good to have an academic publication to link to.
Blame OpenAI.
"We’ve trained a model called ChatGPT which interacts in a conversational way."
I don't need help coding. That's my core skill!
But I do need a lot of help with how packages/libraries/languages work, and ChatGPT can usually give a decent answer in a minute that could have taken me hours of deeply frustrating searches.
I won't blindly trust either partner, and at least ChatGPT isn't insulted when I check if it is right :)
However, I am not a great writer, and not a native English speaker, and I find that ChatGPT is better than I am at this. It is after all, a large language model, it is really good with words, it is what it is designed to do, more so than problem solving. I usually feed it my code and let it document it, if it gets it wrong, I correct it, add some context,... Essentially, ChatGPT is my editor (the job, not the software).
I look at hard to understand code pretty often. Maybe ChatGPT can be helpful there too.