Someone in the thread below says their bug was closed because of mentioning that they use AI to confirm the bug they had encountered. Is he going to go back and reopen all of those now that he learned what pretty much everyone else already knew?
Someone in the thread below says their bug was closed because of mentioning that they use AI to confirm the bug they had encountered. Is he going to go back and reopen all of those now that he learned what pretty much everyone else already knew?
I don't know about this particular case, but if I saw someone report a bug and as evidence claim they had X Y and Z LLMs verify it I would be pretty upset. If you're going to use an LLM to make a replication, just do that and give me the replication, don't point to your notoriously error-prone tools as though they lend your report credence.
It's in a similar vein to people who reply to questions with "well Claude says: <chat transcript dump>"
or substantially worse: "<chat transcript dump>"
If a report is improved and becomes actionable, that's great.
That you used an LLM to "confirm" the bug is an example of such non-information. The reproduction testcase and/or the reproduction steps and description of the expected outcome confirm the bug.
In the state of the tagged video he says still not accepting AI submissions until a certain set of preconditions is met. So... No?
- using the LLM to find (possible) bugs and a human confirms it by testing, reviewing, etc.
- using the LLM to find and confirm the bug without the human confirming it
> that I had every LLM check it to confirm it's a bug
I'm primarily objecting the way he phrased it, as opposed to just saying "I've tested and confirmed and reproduced the bugs". Instead, it sounds uncertain and detached, like he asked the LLMs if it's indeed a bug without further verification.
Bug fixes should get the same treatment. A patch is either correct or it isn't. Projects that ban AI-written fixes outright are asking "who wrote this?" instead of "is this right?", and users live with the bug in the meantime.
I get why maintainers are fed up. Review time is scarce, and they're drowning in plausible-looking garbage. But that's a problem with low-quality submissions, not with AI as such. Require tests, require a human who vouches for the patch and will answer for it, and ban repeat offenders. Then hold every patch to that same bar, whoever or whatever wrote it.