Pop!_OS bans AI-generated code from much of its codebase
neowin.net
neowin.net
Are you in Python by chance? Python has a lot of crazy hidden/inexplicit/spooky action at a distance stuff (especially in the frameworks) that can make LLMs gunk up code by defensively programming or just burn context chasing data provenance
A model can reproduce large swaths of its training data exactly. It’s a different algorithm that powers its learning process (it’s why it needs trillions of tokens to even learn basics of language).
If there was a spectrum from copying on one end to creative production inspired from something else on the other end, the human generally lies heavily on the right end, while the model is much more on the left, that gap is large enough, that yes the model is in some sense “copying”.
Just because a model can reproduce parts of its training set doesn't mean that's what it's doing when it solves a programming problem. It can also reason about the problem, draw on its knowledge of algorithms and data structures, write tests targeting APIs it's never seen before, generate synthetic data and run experiments, etc. etc. Also, the claim that it can reproduce large swaths of its training data verbatim is an empirical one. I would be surprised if it could even reproduce 0.1% of the books it's ingested, for example.
Saying that AI can't do anything but copy or steal from humans seems to be a rhetorical technique used by people who are still unaware or in denial about the capabilities of the agentic systems released in the past few months. They can now one-shot theorems and programming problems in a few minutes that would be difficult and time-intensive for even the 99.9th percentile human expert.
No human on earth can copy at this scale. Yes AI is beyond a database lookup, it does have reasoning on top of this knowledge, but my original claim that if there is a spectrum between copying with minimal changes and creative inspiration with minimal copying, AI is one the copying end while humans are on the creative end. Humans are very bad at reproducing anything verbatim, even if they wrote it like a week ago.
I'm skeptical that there is a single spectrum like you're describing. It's not well defined. Say a human and an LLM prove a new theorem independently (without external help, i.e. from their own neural weights and reasoning). How do we measure how much each of them copied from previous work, as opposed to having learned from or been influenced by it?
A good workman shuts up and finds better tools without complaining.
Save your "you're holding it wrong" if you're not going to suggest how to hold it.
Cult speak escape hatches are intellectually lazy.
Edit: My latest project is all GPT-6 Astra High. It takes a lot of steering to keep it from adding a bunch of, while useful, features that are not strictly enough to the point. That main issue is it’ll use a lot of extra tokens in the process!
What was your process?
In case it is unclear, I am genuinely curious. I have great success with chatbots, but vibing coding has never gotten me further than a proof-of-concept.
https://williamcotton.github.io/datafarm-studio
Some demos of the above charting language:
https://williamcotton.github.io/algraf/demos
WASM, in browser editor, LSP, and more.
https://github.com/NousResearch/hermes-agent is 99% (just a guess) LLM generated. 1140 closed pull requests this week. 1.5k closed issues. The github insights page for commits doesn't load for me presumably because it can't handle this scale of commits. But I estimate ~1K commits per day on average.
There's a blog entry https://nousresearch.com/refactoring-hermes-with-1393-agents that details some work that was done by LLMs to refactor and improve the code.
I guess they know how to hold it?
I had a look at the kind of issues that are reported at that project (there's 15k of them, so I can at best assess a couple). It looks like a complete mess: A lot of concurrency and resource mismanagement issues and edge cases that in a better-managed project would have been avoided by construction. They will now will likely be solved by more defensive programming, driving overall complexity ever upwards.
If you really want to check some quantity metrics to try to reason about code quality, look at whether "fix" PRs are overall LOC neutral or negative (not counting tests). In this project, almost every "fix" is an addition. Worse, almost every fix is more branching.
If almost every PR is some sort of fix, and most of them add branching, and there's thousands of them weekly... That leads to only one place and I want to be nowhere near it.
Show me an AI that adds features by deleting code (https://www.folklore.org/Negative_2000_Lines_Of_Code.html) and I'll pay attention.
Your response was: "Well you're not doing it right, but these hermes devs know what they're doing".
But the blog post you linked to shows their prompt, which is:
> I want god files broken up. I want simplification across the board. I want unification of helpers and methods that can be reused. I want less if-if-if-if-if-if-else routing. I want code legibility up. I want interpretability of the codebase and how things connect to each other up.
So it sounds like AI made their code a mess too. They then tried to make the point of how much money they saved cleaning up the code with AI, that AI made a mess of to begin with.
And if you look at the merged PRs on that project, a ton of them are bug fixes... to the code the AI wrote. And that's been my personal experience too: AI creates a huge amount of churn in a codebase. Just vast amounts of PRs fixing code that the AI itself wrote.
I’m saying it’s probably multiple factors and both you and GP are right.
Seriously, I find I need to slow down the rate of change. I don't move forward until I understand the change proposed and have updated the docs. At the same time, I find that keeping up with the LLM/agent is exhausting. 3 hours with an LLM leaves me as tired as 6 hours with a keyboard had previously. I find that coding when tired or fuzzy yields code that shouldn't have been written in the first place. Sadly, once it's been written and debugged, the temporary fix becomes permanent.
I'm using local models, and they go slow enough that I have no trouble following along with what they're doing; but visually, both go and typescript, along with react native, make me puke. So I wouldn't be able to do this without AI.
I describe how to do it in my comment history, but it's basically a Super-TDD along with some custom engineering harness.
I don't want to say skill issue, but the same way you can give a chain saw to a teenager and one to a skill craftsman, well, AI can obviously create whatever you want it to do.
I think some of the variety is simply how fast SOTA models pump out garbage that you simply have to close your eyes because it's not sensible to just watch characters flow across the screen.
Almost all the coding I'm doing via AI is just faster than readable. But I can see the thinking traces and I stop to model when it's obvious it doesn't understand my intent, etc.
So I'm not doubting you created garbage. I'm just doubting that it's a product of soley AI use.
`total = dev + review`
If dev approaches zero, but you review at the same pace as you always have, are you in a better position? Yes.
Will you potentially have a backlog of code waiting for review? Also yes.
Would you prefer to be waiting for the dev team for all of the time instead, then still have the same amount of reviewing to do at the end of it? Absolutely not.
Late 2025 also had a step change when agents could largely code autonomously without handholding like previously, and to be honest it's not worth hearing opinions about AI from before that time, that's how significant the change was.
The models aren't good at architecture and design. But they take direction on architecture and design and design well and can refactor code quite effectively. AI agents can absolutely be used to clean up vibe coded code bases once you figure out if the investment is worth it. The mess can be avoided if you give them sufficient guidance on architecture and design upfront.
That said, doing so purely in text form doesn't feel great right now. I've been thinking about UML lately. The problem with that was the roundtrip after the code was generated and then the implenetation happened. I don't necessarily think UML is the solution, but neither is walls of dense text.
Can you elaborate to back up this claim? WHat exactly is your yardstick for "being good at SW design and architecture"?
Because I found the current SOTA AI models being great at architecture and design, much better in fact than most average real-world devs. Is your yardstick just the John Carmacks of the world by any chance? Because most devs are not John Carmack. They are also not Linus Torvalds, they are not Stallmann, etc.
Maybe your LLM experience is still stuck in the 2023 era of ChatGPT?
And do you consider yourself to be representative of the average developer, above them, or below them?
LLMs don't even need to be better than the average dev, let alone the top performing ones, like you. If they can just be better than the bottom 20% of devs and white collar workers in general(easily achievable when you've been around the block and saw how many useless people just keep warm chairs for high wages in large companies), that's already a huge win for those products.
What I mean, at a previous job I had ran into a memory leak issue in our backend and discovered a colleague pushed a library into prod which came with comments in the source code saying "DO NOT USE IN PROD, IT CAUSES A MEMORY LEAK!". There's cases where human stupidity and carelessness far surpasses whatever issues LLMs cause so maybe the average dev isn't really that much better than the SOTA LLMs.
1) duplication - LLMs are great at generating lots of text, so its faster and easier for them to generate entirely new facilities that overlap heavily with existing ones then it is for them find existing facilities that should be expanded and refactored (note I just said 'find'; actually editing raises the time and difficulty even more). This is fine for a while as the duplicate facilities usually work just fine, up until something needs to be changed across all of them and they miss changing one or more of them, things break, and a bunch of tokens have to be burned tracking down the issue.
2) ever increasing surface area - even when making changes that do expand a facility without much duplication they frequently only add without removing much of anything or changing the overall design of the facility to reduce the amount of state its tracking and the number of branches it has based on that state. I've never seen one decide to split up something large or with too many responsibilities on their own. They will happily create a god class or function and just keep making it bigger.
Unlike a compiler it won't give up at the first sign of trouble but that just means it left alone it will dig bigger and bigger holes.
Treat prompt engineering as a discipline and refine your technique. When it produces garbage throw out the work and start over until you figure it out.
You say you used AI and your projects turned into unmaintainable messes, so your conclusion is that it means AI is not living up to its promises.
I guess if the argument is “AI makes it so you always get a great result no matter how you use it”, then your argument is sound. Your projects not working out proves that AI doesn’t always work no matter what.
However, that doesn’t mean you can’t use AI to create sustainable and well organized code. Failing to do something doesn’t mean it’s impossible and anyone who thinks they can is not paying attention.
I can’t run a marathon. If I went out and tried to run one, I would get a few miles and collapse, failing completely.
I don’t think it would be reasonable, though, at that point to say “running a marathon is impossible, anyone who says they can do it clearly lying. I tried and didn’t even make it 5 miles!”
I wish people would stop assuming their experience with something is the only possible truth.
As ai becomes better these people will begin changing their story because it’s utterly obvious what’s happening.
On the other hand asking these clankers "review the feature branch I wrote" and "review my entire codebase for bugs" or "help me debug this" has saved me months of prospective work.
And more recently most major models have been getting _really_ good at RE, for example you can have OAI models (and maybe A/'s if they don't refuse) use idalib MCP and reverse-engineer stuff from start to finish, then follow up with GLM 5.3 for vuln assessment and exploit PoC.
Stuff that used to take weeks or months now just takes a few hours, or less.
Where AI did make a lot of impact is triage. I can throw a messy bug report of an intermittent issue at claude, and tell I need issue reproduced and fix developed, and there is a well above 50% chance it'll deliver. Never commit that fix to the codebase as-is, of course.
Unless you believe there is substance in social interactions over time, like trust. But then it would be a socially weird move to dismiss promises of responsibility out of hand instead of picking up the invitation to build trust if that's what's perceived to be lacking.
I think a PR "in flight" shouldn't have been closed like that.
All this will do is push out developers like you that honestly disclose, and instead people will now just lie.
You build your OS atop thousands of open source packages, many of which contain AI generated code. Are you going to audit them one by one and remove offending packages? What about the ones you won't remove because the OS would be irreparably broken?
This is just really silly.
The more I think about this the more I think it's like self-driving cars. We have this expectation that self-driving cars MUST be 101% safe and never get into any accidents, ever, before the technology is worth adopting. LLMs are the same -- it's like we think if you can't one-shot a prompt and get perfect software out of it, it's failed. You can choose to spend time getting the LLM to refine the code it's written, review the architecture, come up with an actual engineering process around the LLM. Yes that means you'll be producing less code per time spent -- which is a good thing.
Being hand written is no guarantee of high quality, just like using LLMs is no guarantee of low quality.
I’ve read a lot of code hand-written by programmers. By smart and hard working people. And I know from experience that the code the average programmer writes is not great either.
I have yet to understand how maintainers can't distinguish beyond (1) PRs that literally include Claude co-author notes or (2) low quality code contribution regardless of the creator.
PopOS is Ubuntu with extra problems. Ubuntu itself is fine, but then PopOS adds weirdness.
Cosmic has been in beta for how long ?
I daily drove the alphas before the betas, and of course there were a couple rough edges. But I had a minimal working desktop instead of sway or KDE/Gnome (too heavy).
It has been releasing non-betas for a good year now, and it's been a smooth sailing.
They aren't saying that AI produces bad code or is terrible for the world in some way.
It's mainly just resulting in a lot of PRs that they don't have enough time to review or features they don't plan to add.
I think even amongst the HN and Linux userbase, pop_os is still niche, let alone amongst normies who never heard about Linux. So they can afford take the high road and treat it like their personal sandbox, accepting only human written code.
But larger and more important projects like Fedora and Debian are more pragmatic with the fact that they'll have to accept AI written(but human reviewed) code, if they wish to keep up with the real world development and threats, as expressed by Linus Torvalds himself.
The thing is, the cat's out of the bag on this one now, especially in the field of pen-testing and reverse-engineering. AI can brute-force its way into projects in ways that beat even experienced researchers, so your only choice to keep up is to accept the use of AI generated fixes as a counter defense.
Do you think Netanyahu, Trump or Xi-Jinping are somehow secretly using Cosmic DE at home, to be worthy targets?
Bad actors have limited time, lives of their own and mouths to feed as well, so they concentrate their efforts where "the fish are" if they want to PWN someone for profit.
That's why Windows was the biggest target in the past for so long and why MacOS and Linux were ignored. Because most of the fish were on Windows.
Previously, time was the most precious resource. Now its tokens, and more cheaply at at.
If you assume Mossad and NSA are Token-maxxing every single niche FOSS project out there to cast as large as possible fishnet on hacking all Average Joes on the planet just in case, then maybe using Mozilla and MacOS gets you hacked too, maybe even visiting HN and commenting here gets you hacked by some zero days you don't yet know.
Where does this open-ended paranoia argument end?
(I wasn't able to make it work on a scrap Dell I tried it on because the GPU was too old. Booted the USB key and COSMIC greeter failed to start)
1) Firstly, your source plase? My research according to Google Gemini 3.8 Pro shows top 5 DEs are as follows:
+---+---------------+------------+-------------------------------------+
| # | Environment | Est. Share | Primary Ecosystem / Defaults |
+---+---------------+------------+-------------------------------------+
| 1 | GNOME | 45% - 50% | Ubuntu, Fedora, Debian, RHEL |
| 2 | KDE Plasma | 25% - 30% | SteamOS, openSUSE, Kubuntu, Manjaro |
| 3 | Cinnamon | 8% - 12% | Linux Mint flagship |
| 4 | Xfce | 6% - 9% | MX Linux, Xubuntu, low-spec PCs |
| 5 | MATE | 2% - 4% | Ubuntu MATE, Mint MATE (GNOME 2) |
| - | Others / WMs | 3% - 5% | Hyprland, i3, Sway, LXQt, Budgie |
+---+---------------+------------+-------------------------------------+
So then COSMIC isn't even TOP 5, for this to be a major target by market share as originally claimed.2) Secondly, keep in mind that distrowatch is not representative of linux userbase. MX-Linux kept showing up at the top spot for many years despite being niche.
3) And thirdly, TOP 5 DE isn't really an achievement when Linux DE market share is overwhelmingly dominated by KDE Plasma and Gnome as the majority shareholders, with XFCE and Cinnamon trailing. So Cosmic DE if it somehow made the no. 5 spot, would still be ignorable sub <1% market share, as I initially claimed, basically invisible to bad actors.
"Way too defensive" how? By asking and bringing data for my PoV?
>GP replying to you was clearly trying to have a conversation
As am I, except I ask for, and also bring data to back up my PoV, instead of vague opinions.
>Chill.
Where am I not being chill?
For funsies, I threw two different DE-by-marketshare inquiries at Gemini 3.8 Pro in separate sessions and it ranked Cosmic 12th in the first and 7th in the second. Very trustworthy stuff.
Please show me what part of what I said before was not "a normal discussions" according to you?
>it ranked Cosmic 12th in the first and 7th in the second. Very trustworthy stuff.
And that disproves me how exactly? I originally showed "Cosmic is NOT a TOP5 DE", and the LLM data you posted also shows that IT IS indeed NOT a top 5 DE.
Do you have a better source than GP?
>if needed post the actual source
Define "actual source"? In good faith, I mean.
Where else do you get this information that's, quote, "actual source"?
What do you expect exactly? Do you want me to now manually parse through terabytes of information at your whim for your own convenience? Sorry, but I'm not your personal unpaid servant.
If you wish to disprove me in the comment section, then you need to do the manual work and show us that the my quoted LLMs statistics are wrong. I'm not your personal errand boy to do your bidding, 'massa'.
Did you just blindly trust the opinions of others on this topic, and yet I'm the one who has to provide peer-reviewable data for you to back-up mine? Sorry, I'm not your unpaid lackey. Try to formulate a better (counter)argument for why their baseless opinion is right, but my LLM backed up opinion is wrong, if you wish for a even-footed good-faith argument.
I also saw it now. That blog is not a representative ground truth, but just another opinion piece, which I can respect as an opinion of the blog's user base, but I can't take as an accurate real world statistic, same how aggregate opinions you read on HN are not representative of the actual real world.
I hope you can understand my PoV. You can also disagree if you want, but you'll need to bring something more than "that's wrong because LLMs sometimes hallucinate" as proof that Gemini's data is wrong in this case.
> you'll need to bring something more than "that's wrong because LLMs sometimes hallucinate" as proof that Gemini's data is wrong in this case.
actually I can say it's wrong or likely to be wrong especially if it doesn't cite the sources it uses. And if it does, then just paste the sources here instead of the LLM output. It is also unknown where it got the info and as someone else said, you ask it two different times and it gave two different answers, thus it is unreliable.
How did you verify that those people from the blog are "real"?
I also talked to real people for my own data, case in point, I asked my mom and dad which linux DE is most used, and the results came out different. Which "real people" are the ones representative for the ground truth of Linux DE sahre?
> and yours is just a hallucinated list from one of the me-too LLM vendors
How do you know it's hallucinated? Ask the LLM the population of your country? Is the answer mostly accurate or is it hallucinated in an inaccurate way?
Aren't LLMs just outputting the highest statistical probability from the aggregate of their scraped data, which in this case would be including opinions on Reddit, and every website and blog on the entire internet (including that random one posted by badc0ffee) on the Linux DE uusage topic, making it a more accurate real-world representation than just a single random blog?
You can call it "hallucinated" if you want, but that doesn't mean it's not accurate. I asked for proof that my answer was inaccurate, not that it was "hallucinated", those are two different things, and your argument didn't prove it was inaccurate nor did it prove it was hallucinated. Would you like to try again?
Edit: I see you posted elsewhere.
That is such a weird statement that I am not entirely certain where to begin. PopOS is hardly niche. Its base are all fairly common components by linux standards. And, more importantly, attackers and bad actors may other considerations in mind than sheer population size -- just to point out the glaringly obvious.
<< I think even amongst the HN and Linux userbase, pop_os is still niche
I think rather than trying to disprove it, I think I should ask why you think that? If anything, PopOS annoys me because it is just a step before ubuntu ( and ubuntu is just windows at this point ). Maybe I am defensive, because my first real distribution ( that did not share disk with windows was popos )?
If that's your threat model then you shouldn't use any SW in exitance, FOSS or otherwise. In fact you shouldn't even go online, or even outside you own house, since one single bad actors exist everywhere. You can walk down the street and suddenly someone in a car runs you over(witnessed myself). And yet live goes on.
Because your comment didn't disprove that Cosmic DE isn't too niche for bad actors to get involved.
Then they shouldn't reject AI aids to help them patch vulns found by bad actors faster, no?
Consider that it's packaged for many well-known distros, so pop_os install base alone doesn't tell the whole story: https://system76.com/cosmic/download
Reminder that AI is quite stupid.
I'll trust the words of groups like curl (https://daniel.haxx.se/blog/2026/06/10/a-human-in-control/), Linux, and even the infamously anti-AI Gnome (https://blogs.gnome.org/mcatanzaro/2026/10/02/the-era-of-sof...) that AI is finding real vulnerabilities and you're your project a disservice by ignoring them.
Edit: Though Greg did recently have a talk (that I skimmed) where he was a little reserved on LLMs: https://www.youtube.com/watch?v=NnV_cWeoo5Q
For example, I was helping work on an open-source game engine earlier this year with a longstanding text rendering bug dating back to around 2021 that prevents the engine from being production ready, which the community and myself have developed extensive workaround for. So, one day I've finally said enough and got Claude to debug it. It took Claude 10 minutes to find the bug, it was 3 lines of code change in the renderer (yes, three).
So, I wrote up the regression tests, documented the bug and opened up a PR for the fix, thinking it'll get merged in like less than a week and then we can all move on. The maintainers received it fairly well on the PR, but the PR sat there for nearly 6 months, unmerged, until it finally closed from a bad squash upstream. I'm pretty sure the bug is still there too.
And as a side note, I would be ecstatic if someone wants to contribute to my Github projects with their AI.
Other people's open-sourced project is different though, so I tend to check/test the code myself much more rigorously when contributing to other people's repos than if it is my own projects.
In the meantime, others have experienced this bug and are waiting for the maintainers to notice and fix it. I'll just continue to run my fork (so far it has been no burden at all) while I (slowly) decide what direction to take.
Sounds reasonable, even to avid LLM users, I suppose. You have to draw a line. This line is too simplistic, but it'll work, for now.
"Apparently, AI doesn't lead to positive results in all software development teams. Customers are asking me whether they should slow down the adoption of AI in their teams.
How do you respond to that? Yes? No? "
https://www.linkedin.com/feed/update/urn:li:activity:7510679...
Feels like part of it is a learning habit, getting proficient in use of the tools (AI agents) themselves, and better workflow around it, but it remains AI is not perfect yet?
In fact I've start doing that myself. Sending PR and convincing the maintainer why the fix is necessary is just too much effort.
Faith is the problem. Extraordinary claims must stand up to scrutiny. They do not.
PopOS is not a product being sold by techbros trying to pump their stock like AI is, it's free open source software, it doesn't need to "outcompete" anything. If you don't like it, don't use it.
Shipping them full of bugs which LLMs already fixed upstream.
Fortunately, as it is an open source project, if anyone actually cares about the distribution, they'll fork it and continue advancing it using modern tools.
Which is why LLM PRs should just be issues (if there isn’t one already). Make the issue author a co-author on the PR. But let the maintainer actually oversee the LLM generated solution.
So basically you're saying you reject drive-by PRs.
My repos probably don't see as much traffic as SQLAlchemy though.
if i did have to think about tokens I use something like together.ai with an open weight model like glm 5.3 (which ill sometimes use to review a claude change for something intricate). glm 5.x is just very chatty though
I have a hard time knowing if anti AI is a mental illness or propaganda coming out of China.
Unironically.
Separately, why would anyone use a Debian based desktop OS? Your $11 Amazon mouse won't work. An Nvidia card won't work. Just use Fedora.