Testing ability of GPT4 and open-source LLMs to detect C++ bugs
catid.io
catid.io
Edit: the post says open-source models.
> Can open-source LLMs detect bugs in C++ code?
> No, they cannot.
GPT-4 is not open source.
Then you could fix them
I don't get those 'Open source not good enough' posts. However good openAI can get, their secret sauce is known and open source will get there even with a delay. The don't have an exclusivity to the data on the internet
You can provide additional examples of these specific issues, sure... but that doesn't mean it'll catch ones you haven't provided. It's not reasoning about what makes the patterns dangerous.
"GPT 4: 0 false alarms in 15 good examples. Detects 13 of 13 bugs."
I've found Vicuna 13B to be quite capable at _generating_ code. But not as good as GPT-4. Model size probably also plays a large role: 13B parameters runs nicely on normal-ish home user hardware, GPT-4 not quite.
I'm thinking something like a linter with access to the AST, that can produce warnings like "you forgot this corner case", "you are not freeing resources. Is this intended?" and so on.
I do wonder how to reliably tell it to ignore noisy, incorrect warnings though. They're potentially sensitive to any new input / weight / random-seed changes, so it seems like literally every LLM upgrade will run the risk of ignoring existing suppressions (or you say "ignore this whole line" and miss useful warnings) due to small perturbations...
Would be interesting to see how they do.
May. This is by far two most hilarious bytes I ever heard! :D :D :D