ChatGPT's odds of getting code questions correct are worse than a coin flip
theregister.com
theregister.com
> For each of the 517 SO questions, the first two authors manually used the SO question's title, body, and tags to form one question prompt and fed that to the Chat Interface of ChatGPT. The generated answers by ChatGPT are then stored in CSV files. For the additional 2000 SO questions, ChatGPT 3.5 Turbo API is used.
I would be interested in seeing the results for GPT-4. For a paper released in August 2023, it’s strange that GPT-4 isn’t even mentioned.
Except for when you look at the results. ;)
It outputs some boilerplate code, and inserts a method with comments saying "insert logic here" - entirely sidestepping the request.
You can say that's remarkably human-like, but if it was a human cow-orker, you would not keep them on the payroll. It is a cheat, like ELIZA, like SHRDLU, like a useless employee.
I asked ChatGPT-3.5 for the usual rejoinder to complainers:
"If you're looking for an enhanced experience, you might want to consider trying out the paid version called ChatGPT-4. It offers even more capabilities and improvements. Feel free to explore the options available to find the best fit for your needs!""
But oddly that doesn't change my opinion of ChatGPT.
The only reason I know that ChatGPT was right is because I blindly followed the first Googled suggestion.
If I started with ChatGPT, then I would have to verify the answer somehow. If I use Google, I should verify the answer. But ChatGPT cannot do that job.
Wikipedia is starting to feel as useless to me as ChatGPT. I don't know if anyone is actually replacing the articles with AI generated fiction at scale yet, or whether it's still an artisan, human-centered craft but the more I click on references, the more I find they are uncheckable, even if it's ambiguous whether they are fraudulent. People write rants on a topic and cite a single book without page numbers five times. Or, of course, dead links to websites without archive.org.