The era of software quality, or the era of ostriches?
blogs.gnome.org
blogs.gnome.org
in this case it would be organizations refusing to adapt more rigorous quality tooling alongside the rate of release, similar to the footgun allusion
https://www.mcgill.ca/oss/article/did-you-know/ostriches-do-...
This may actually be the time where it becomes feasible to create high quality software with reasonable (human) effort.
What we are seeing with AI generated incident reports causing issues is maybe mostly an impedance mismatch between people who have fully adopted AI and people /projects who haven't.
Reporters that submit useful, quality reports, should not only get their money back, but also a dividend on the money taken from low-quality reporters.
> Currently the best solution is to not use programming language package managers, but GNOME’s Rust code depends heavily on Cargo. Accordingly, I recommend against using Rust for writing GNOME software.
It didn't seem like the article was really about Rust, honestly.
It's not even false, you do know that KDE exists
I was pretty sure that was false. I did a quick check and found
https://develop.kde.org/docs/plasma/scripting/
The shell got a lot of hate for being written in gjs, but you won’t get that kind of functionality from a shell written in a compiled language.
In fact, for many complex projects, it's essentially impossible, even with AI.
Eventually, you get to a point when you have every feature you could possibly want but the list of tradeoffs is long and yet not worth trimming.
Care to explain the circumstances/context that you mean?
I found that quite humbling. I thought I was a lot better at writing secure code than that.
When I worked with blockchain before, that was another level because each node had to talk with thousands of other nodes; each running different versions of the code and they had to propagate messages throughout the network in a secure and scalable way without missing any nodes or reaching the same node twice (spam vulnerability) so your code is basically talking to different versions of itself which feels a bit like recursion, but harder and with a network between which adds latency and errors; and you can encounter tricky issues like message storms among others where the nodes keep bouncing and replicating a message between themselves. Or if any kind of propagated processing doesn't deduct tokens, then that's a massive spam vulnerability. Also, if the peer-discovery algorithm is structured, then the network is vulnerable to eclipse attacks. Also, most parts of the code had to be fully deterministic and idempotent or else some nodes could fork off the network. Stamping out non-deterministic, non-idempotent logic is hard work, especially when you have to do it across multiple distinct nodes running on different machines operated by different people! Paradoxically, some parts of the code actually benefited from being non-deterministic like the P2P networking and peer discovery logic since that made it harder to game/carry out eclipse attacks. So you need deep, nuanced knowledge to get it right.
For the kinds of systems I work on, loose coupling + high cohesion is mandatory, but building reliable systems is still difficult and time consuming regardless.
That blockchain project I mentioned wasn't architected well in the beginning and it was a nightmare to work with; it had to be rewritten because it was literally impossible to add new features on top (in a reliable way); nobody could do it, even the people who were there since the beginning couldn't explain how it was working. Making it modular allowed us to make it reliable and extensible but each module was still difficult to work on in the sense that every feature required careful discussion and debate with team members because so many tradeoffs had to be considered.
Also you made a general statement about writing correct code being insanely hard. I mean, writing hard code is hard, but many if not most of the things we do in this profession are not in that category. And most of the correctness we need to achieve isn’t insanely hard. Agreed that there are of course areas where the confluence of challenges can be extremely challenging.
Sometimes even something which sounds really simple, isn't simple at all once you consider all the possible formats, modes of interaction and all the things which could go wrong.
1) person's name 2) their address 3) their email address
This form must
1) meet or exceed WCAG AA 2.1 2) handle people from all over the world 3) only allow valid inputs
If you've never done this before you will think it easy. If you have you will know that it's nearly impossible. This is web development btw.
I've also had registered domains turn the box red with "Please enter a valid email address" as if it needed to end in .(com|org|edu|gov)
Any of those little junctions between features could very well have <10 users even on multi-million users pieces of software. Now every one of those little possible cross-sections in the venn diagram is another possible state the software can be in and that needs to be tested.
In the software I work currently we have one of such case, a combination of 3 different features that almost no users use in conjunction, but if they do it breaks the system. I am aware of it, my manager is aware, the product owner is aware, the fix is not trivial (in fact it would require rewriting large portions of the system).
It just stays there and we have monitoring and customer-support manually handling it. The product is fundamentally not correct (in fact the correct behavior is not even defined), yet it still works for 99.99% of the users.
Practical example? I have ran multiple times in multiple different products (even ones I worked) where if you login with Google/Apple/Whatever that makes your account completely borked to be used with SAML/OIDC if you get invited into one later. You basically need to call support to ask them to remove your user and be invited again. All those software were, by definition, not correct.
Once I saw that I knew that this article was a ridiculous stack of bullshit. Gnome might have some bugs and shortcomings but now millions desktops are using Gnome based window manager for decade and the world didn't collapse yet.
And let's not forget the ridiculous assumption that you can't handle a software quality without using AI...
What was good enough in the past may not be good enough now. In the past these overflow defects were as hard for attackers to find as they were for developers, because they had the same tools available.
Now we have LLMs, and they can apparently find all sorts of problems that, for whatever reason, weren’t being found with manual review, static analysis and fuzzing. We have to assume that hackers will use the technology to find vulnerabilities. If maintainers don’t do the same, then they are ceding an advantage and leaving their users unnecessarily exposed.
I don’t know that I completely agree with the above. (For example, I don’t know how scrupulous GNOME has been in the past about non-AI tools for automated defect discovery, or how well they compare to AI.) But it at least feels like a much more charitable interpretation of the article’s main thrust.
And even then there is the risk that the frontier model providers collapse. There’s no indication yet that these companies are going to become profitable. And open source models simply piggy back on frontier ones and are generally not powerful enough for this level of adversarial attacks as far as I know.
And finally are LLMs the only method to hardening software? There is still a lot left on the table that could still allow a project to resist attacks from a frontier model.
Sure, humans are bad at catching this stuff. But in order to drive an LLM you have to be able to catch this stuff. Otherwise your only option is to trust the model and give up.
For example: https://link.springer.com/article/10.1186/s42400-020-00058-2
Perhaps using static analysis generates false positives in cases where the code can't be proven safe? But when I was working on a project where we used Sonarqube, we ended up deciding as a team that we'd prefer changing the code to eliminate false positives over "wontfix"ing them, and I was happy with that decision. It led to more regular coding practices that ultimately made the codebase easier to read and understand. For largely the same reasons as Dijkstra was getting at in "Go To Statement Considered Harmful."
There is no excuse for not having an LLM look for flaws in your codebase. Your attackers are using them already.
I think writing tests is a great use of ai, but only if the tests themselves are throughly reviewed or are themselves validated by statements in a formal system.
That's more-or-less how I define QA work. The goal is not proving overall correctness, it's surfacing individual issues.
I think people higher up are warning against a naive mistake, which is asking one agent to write both the product and the test suite. Here, making things up and cheating becomes a problem of quality theater...
The QA role (regardless of whether it's performed by an AI or a human) typically involves breaking tests, not writing them. (This applies more to unit tests, integration tests blur this line.)
this doesn't sound like your model, and I'm unclear what it means to break a test. maybe test automation, in which case, sure that seems fair game for AI, but that not where the real meat is.
This pretty much. My last job didn't really have QA so I did the next best thing which is writing a bunch of integration tests for the use cases that matters for the product. They were not an indication for correctness, but more like a canary to warn me if I break something while developing. Bugs reported by consumers usually warn me of area not well covered.
QA would play the same whole. They shouldn't need to check for code correctness, their most useful task is to surface bugs that breaks the product requirements (performance, security, business logic,...). And for that, having knowledge of the implementation is unnecessary. In the above example of writing integration tests, I took care of only using the public interface of the modules.
In my mind, tests are everyone's responsibility, but the goal differs. A regular developer adds features and should be writing tests to prove their code functions as intended. QA does not add features, so their goal is finding holes in code added by others.
It's the adversarial relationship mentioned by a parent:
> Have [QA] build tests to break the code. Don't give [QA] the job of making a test suite that passes.
While other projects rewrite things in memory-safe ways, GNOME's response is to ask a chatbox if there are any memory vulnerabilities.
That's not only an uncharitable take, it's also wrong.
If 999 out of every 1000 "reports" from a specific source is wrong, then it is not irrational to disregard all 1000, especially when they can be generated faster than you can read.
I mean, it's just probabilities, right? If you're okay trusting output from an LLM, you should be okay with using statistics in general as a source for informing decision-making.
Right, but that assertion does not contradict what I said: there's a difference between the articles premise (AI reports are mostly valid) and what I said (3rd-party submitted AI-reports are mostly invalid).
See my reply to a sibling poster who also implies I did not read the article.
Then look. If you can't judge, then trust the experts. This article is written by experts.
The people writing this article are experts. They cite other experts.
If you can cite real data from the last few months that still claims there are no wolves - and it's not just insane anti-wolf propaganda - I'd love for you to show me.
Cite data from the last few months.
Let's return to actual data: the article we are discussing is experts claiming that AI should be used to report errors in software.
Do you have data to support your side?
EDIT: I did a little research. These projects accept AI contributions: Python, NumPy, SciPy, pandas, scikit-learn, Django, Kubernetes, the Linux kernel, Firefox, Flutter, Homebrew, curl and PyTorch.
Why do you think they do, if it's such a bad idea? Are they all idiots? They have no idea what they're doing? Linus? Really?
"Economists have predicted 18 of the last 2 recessions".
I mean, c'mon! You have never read that?
Besides, when "expert in $FOO" means "familiar with $FOO that's only 6 months old", then it's not unreasonable to be skeptical.
In other fields, an expert is someone who's studied the specific field $FOO for decades. Here we're talking about a skill level that is not distinguishable between "1 weeks experience" and "two years experience".
I've found that software engineers are good at analyzing the public reports they receive for their own projects.
Let's not over generalize, shall we?
You can read how skeptical I was if you dig through my comments if you want. But at some point it's not skepticism, but pigheadedness.
> Have you heard that most AI bug reports are “slop?” Not so in 2026. That was true for most of 2025, but the quality of AI-generated vulnerability reports has drastically improved. That is not to say that we no longer have problems with bad vulnerability reports, but in general, nowadays most of them are pretty good. (Daniel Stenberg reports the same pattern for curl.)
I read the article very carefully, including the bit that you quoted. Here's what I read:
> Red Hat’s scan of GLib found 118 vulnerabilities. Or at least, it claimed to. However, due to the way we ran the scans, several of these are actually unnecessary duplicates of each other, which we have not fully deduplicated yet, so the number I report is not entirely trustworthy. Moreover, 46 of these “vulnerabilities” are bugs in gobject-introspection, mostly in the typelib support, which is evidently not very robust. A typelib controls how your program calls libraries; it is effectively calling convention, so it must inherently be fully trusted: a malicious typelib would be able to induce vulnerabilities even without any bugs! I would expect an AI ought to have been able to figure that out, but apparently not. These bugs are still real problems that we ought to fix, but all maintainers agree they are not security vulnerabilities, so let’s count all of them as false positives. That alone creates a 40% false positive rate. Ouch.
And that's with them running the scanner, not with submitted reports by 3rd-parties! When you welcome AI reports, everybody is going to submit the same report, just differently ordered and differently worded.
I mean, he even said:
> I requested that the bug bounty program end because I was overwhelmed with incoming AI-generated issue reports. The final issue was reported on February 23, 2026. Here are the results:
Sure, he attributes it to a financial incentive, but it's clear that submitted AI reports will overwhelm, and the only way they got to a measly 40% real-bugs was by do the scanning themselves.
(Also, I wish all these sibling posters implying that I did not very carefully and thoroughly read the article would, themselves, read the article!)
I hope you understand that software engineering is now going to require the use of AI. In the same way that software engineering requires the use of compilers. Sure, some people refuse... but...
> I hope you understand that software engineering is now going to require the use of AI. In the same way that software engineering requires the use of compilers.
When LLMs actually can reliably do their jobs (which compilers do), then they might be an essential tool. Not before. For now, they are slop machines used by people who care more about going fast than getting things correct.
LLM agents do a fantastic job of finding exploitable bugs in code. MUCH better than humans.
They would also do a fantastic job isolating the duplicate reports as described above.
So what's your issue?
Out of, say, 40 bugs that SOTA models can find, you're still going to have to sift through all the hopeful wannabes who each submit that same list of 40, but differently worded, differently explained and with different PoC code.
The problem still remains when welcoming AI-reports from the world: you could potentially spend all your time on examining and then discarding these reports without even getting to any new bugs in those reports.
No one did though, because the friction involved in finding and submitting bug reports meant that a single individual could not overwhelm a project in spurious reports.
Or you end up with an AI version of Chinese whispers.