That's not only an uncharitable take, it's also wrong.
If 999 out of every 1000 "reports" from a specific source is wrong, then it is not irrational to disregard all 1000, especially when they can be generated faster than you can read.
I mean, it's just probabilities, right? If you're okay trusting output from an LLM, you should be okay with using statistics in general as a source for informing decision-making.
Right, but that assertion does not contradict what I said: there's a difference between the articles premise (AI reports are mostly valid) and what I said (3rd-party submitted AI-reports are mostly invalid).
See my reply to a sibling poster who also implies I did not read the article.
Out of, say, 40 bugs that SOTA models can find, you're still going to have to sift through all the hopeful wannabes who each submit that same list of 40, but differently worded, differently explained and with different PoC code.
The problem still remains when welcoming AI-reports from the world: you could potentially spend all your time on examining and then discarding these reports without even getting to any new bugs in those reports.
No one did though, because the friction involved in finding and submitting bug reports meant that a single individual could not overwhelm a project in spurious reports.
> Have you heard that most AI bug reports are “slop?” Not so in 2026. That was true for most of 2025, but the quality of AI-generated vulnerability reports has drastically improved. That is not to say that we no longer have problems with bad vulnerability reports, but in general, nowadays most of them are pretty good. (Daniel Stenberg reports the same pattern for curl.)
I read the article very carefully, including the bit that you quoted. Here's what I read:
> Red Hat’s scan of GLib found 118 vulnerabilities. Or at least, it claimed to. However, due to the way we ran the scans, several of these are actually unnecessary duplicates of each other, which we have not fully deduplicated yet, so the number I report is not entirely trustworthy. Moreover, 46 of these “vulnerabilities” are bugs in gobject-introspection, mostly in the typelib support, which is evidently not very robust. A typelib controls how your program calls libraries; it is effectively calling convention, so it must inherently be fully trusted: a malicious typelib would be able to induce vulnerabilities even without any bugs! I would expect an AI ought to have been able to figure that out, but apparently not. These bugs are still real problems that we ought to fix, but all maintainers agree they are not security vulnerabilities, so let’s count all of them as false positives. That alone creates a 40% false positive rate. Ouch.
And that's with them running the scanner, not with submitted reports by 3rd-parties! When you welcome AI reports, everybody is going to submit the same report, just differently ordered and differently worded.
I mean, he even said:
> I requested that the bug bounty program end because I was overwhelmed with incoming AI-generated issue reports. The final issue was reported on February 23, 2026. Here are the results:
Sure, he attributes it to a financial incentive, but it's clear that submitted AI reports will overwhelm, and the only way they got to a measly 40% real-bugs was by do the scanning themselves.
(Also, I wish all these sibling posters implying that I did not very carefully and thoroughly read the article would, themselves, read the article!)
I hope you understand that software engineering is now going to require the use of AI. In the same way that software engineering requires the use of compilers. Sure, some people refuse... but...
> I hope you understand that software engineering is now going to require the use of AI. In the same way that software engineering requires the use of compilers.
When LLMs actually can reliably do their jobs (which compilers do), then they might be an essential tool. Not before. For now, they are slop machines used by people who care more about going fast than getting things correct.
LLM agents do a fantastic job of finding exploitable bugs in code. MUCH better than humans.
They would also do a fantastic job isolating the duplicate reports as described above.
So what's your issue?
Then look. If you can't judge, then trust the experts. This article is written by experts.
The people writing this article are experts. They cite other experts.
If you can cite real data from the last few months that still claims there are no wolves - and it's not just insane anti-wolf propaganda - I'd love for you to show me.
"Economists have predicted 18 of the last 2 recessions".
I mean, c'mon! You have never read that?
Besides, when "expert in $FOO" means "familiar with $FOO that's only 6 months old", then it's not unreasonable to be skeptical.
In other fields, an expert is someone who's studied the specific field $FOO for decades. Here we're talking about a skill level that is not distinguishable between "1 weeks experience" and "two years experience".
I've found that software engineers are good at analyzing the public reports they receive for their own projects.
Let's not over generalize, shall we?
Cite data from the last few months.
Let's return to actual data: the article we are discussing is experts claiming that AI should be used to report errors in software.
Do you have data to support your side?
EDIT: I did a little research. These projects accept AI contributions: Python, NumPy, SciPy, pandas, scikit-learn, Django, Kubernetes, the Linux kernel, Firefox, Flutter, Homebrew, curl and PyTorch.
Why do you think they do, if it's such a bad idea? Are they all idiots? They have no idea what they're doing? Linus? Really?
You can read how skeptical I was if you dig through my comments if you want. But at some point it's not skepticism, but pigheadedness.