Protecting our FLOSS commons from LLMs
blog.codeberg.org
blog.codeberg.org
I disagree with basically all of the logic they're using here, but that's the point: not every site needs to be for every person. Clearly there are quite a few software developers who agree vehemently with the stance they're taking. They should have a community hub, and now they do.
I also don't see your points of contention - it may help to be more specific.
Last but not least, I am also not sure they are a community hub. This depends on how many people use codeberg. IMO it is still a niche right now; perhaps that will change, but right now I am not entirely certain it is a community hub per se when compared to github. But competition is good and if there are alternatives to fatso Microsoft controlling github, I am all up for that. The upcoming mandatory age sniffing will be a waterfall moment.
I didn’t interpret GP that way. Rather, they were saying that like-minded projects are appropriately homed.
It's a real impasse where the pro-AI or indifferent opinions seem to simply stop posting when they realize there is going to be no real debate, and the anti-AI individuals swarm on any new input.
My personal opinion is I don't care either way -- AI is a tool and if contributors use that tool to improve code or the product, I could care less. I just find this interesting.
> 7. You must not share projects that mostly consist of code written by "generative AI"-tools (including services such as Claude, OpenAI Codex). Such projects having an unclear copyright status (see requirements § 2 (1) 1 and § 2 (1) 3) and furthermore have little safeguards to ensure that they do not include harmful code (c.f. § 2 (1) 5).
https://codeberg.org/Codeberg/org/commit/96fac426a32d1ba91ff...
For me that would be the single biggest reason. People who vibe code don't legally own the copyright to the output (to the best of my hobbyist-non-lawyer knowledge), so they cannot themselves provide permission for others to be able to legally copy.
If something is not covered by copyright law then others do not need permission to copy it. However, you cannot impose conditions on copying it either (not even minimal conditions like the 2 clause BSD license).
However you can enforce licenses in jurisdictions where even pure vibe-coded LLM output is covered by copyright. So you might find yourself in a position where, for example, you can enforce the GPL on your vibe-coded software on someone in the UK, but not someone in the US.
If a photographer can own a point and click snap of a painting then I expect we'll land somewhere where LLM outputs are or can become owned by the tool weilder with at most some minimum effort.
In the US at least, the PTO has not said this. They have said that artwork is not copyrightable if it was substantially "generated" with little human input. They have not said anything about software or other creative works. And certainly nothing about software that has been carefully "vibe engineered" with lots of human input, direction, and review.
It seems unlikely to me that the PTO would declare that a software project that heavily utilized an AI-based advanced autocomplete would make it not copyrightable.
Now, if you just one shot the most advanced Tetris algorithm of all time and post it on the web without any scrutiny or review, and no input, then it's probably not going to be copyrightable.
Why would they have different rules for one potentially copyrightable material versus another?
What the PTO (courts, actually) have not said is how much and what kinds of work a human must perform to transform a machine-generated work into a copyrightable work. It could be that this would be domain specific (i.e. different rules for software and art). But totally different rules entirely for software and art? Seems unlikely to me.
There is a longstanding distinction between expression (copyrightable) and facts and ideas (not copyrightable). Nor is this merely used as a binary category: the degree to which a work rests on facts and ideas, versus expression, is very much taken into account when considering the question of infringement.
Classic cases involve works like maps and biographies, versus works of pure fiction. You can surely see how the expression of the former is constrained in ways in which the latter is not.
The algorithm to reverse a binary tree is not subject to copyright. Nor is a court likely to accept the use of "left" and "right" as variable names as evidence of infringement.
They would treat the cases differently because the cases are different.
My opinion has always been that copyright law is an awkward fit for software, and remains so. That said, the iteration and edit cycle which is normal to agent-driven software development, has a rather different character than the more normal 'workflow' for art, which is: "drawing of Asuka eating an elote in front of the Chinese Theater" then picking the one the user likes the most.
There's also the realpolitik to be considered: courts are simply unlikely to shut down the use of LLMs in software development, that horse has left the barn. There's always some margin between 'can' and 'will'.
I agree that the copyright office has not said this either -- they have just said that you need to be a human to register a copyright.
There's no reasonable belief that or similar scenarios cannot happen and does not happen.
That's my problem with Linus's announcement regarding AI in the kernel -since the announcement I'm counting the days until we see SCO II Electric Boogaloo.
There's a lot of people with deep pockets and a deeper hatred of FOSS, and AI gives them the backdoor to take FOSS out -legally, on the grounds of province.
(I'm not a lawyer, this is not legal advice, and this is only my speculation from a layperson's point of view)
That's not how copyright works. If you use a pirated copy of photoshop the resulting image is not illegal.
It's not as simple as your reductive analogy. It's for a court to decide.
How many youtube videos have gone dark over bogus copyright claims over the years? And you think people will overlook what's in AI -when there's over a trillion dollars on the table?
You seem to think one needs to have copyright to publish something but that's not true. If you don't infringe on anyone else's copyright there's no problem. And others can of course copy themselves.
The only problem is if someone else had copyright to the output.
The blog post is more consistent with the actual change to the TOS. It makes a lot more sense now. If the TOS are a problem then you can always use another host or self-host. Codeberg is intended for a specific type of project (open source license, public repo, encouraging collaboration etc.)
Beyond the reasons stated I think there is probably a more general (albeit less tangible) benefit to having a forge with a reputation for being human-focused. Like if you see a project on Codeberg you can reasonably expect it's probably not completely machine-generated slop, which is no longer an expectation I have for GitHub. People are getting offended about it but why not just use a different forge for your vibe-coded projects?
The main risk is that it leads to loads of distracting discussions about whether projects are "too vibe-coded" or whether they fall on the right side of the line. I would assume Codeberg would be relatively light touch about enforcement in marginal cases, which equally might blunt the effectiveness of the rule but ultimately, it doesn't have to be perfect to be useful.
See also: 100MB per-user quota for repos that are not meaningful software.
This is all reasonable.
I will have to move my projects off Codeberg, but that's okay. We have incompatible views. No grudges.
I couldn't care less about if AI was used to develop software, if it does what it is supposed to do, and if it is secure. Bad software can be produced with or without AI, so ultimately it is up to _you_ to vet the software you use if you don't pay anything for it.
Any voices that said anything on the PR or on threads about it on Mastodon even ask for clarifications on where the limits etc were intended to be were immediately attacked and shut down in a black and white way... It was very clear where it was going to go and there was zero nuance to the discussion. The "community" shouted out anybody who would have voted against this.
I got ready to pull my project the moment I saw all that going down, and just waited to pull the trigger til the day it passed.
It's their right to narrow their definition of community on this basis. I also think the way they did it is crude, lacks nuance, and was executed poorly.
I'm sure my project probably would not have fallen afoul of their judgements. I hosted my own CI in part because I didn't want to suck up their resources. I was by no means producing "slop" -- I'm a software engineer with a 30 year career and capable of engineering judgement. I'm also long time open source user and contributor, an advocate of copyleft licensing, have my own critical opinions about aspects of LLM usage, and wanted off GH because I saw a lot of ... evil.. there and wanted to not be reliant on American infrastructure among other things.
But... I like probably the majority of authors of large packages these days, use AI tooling as do my contributors. And parts of my project have hooks to use open weight models, and provide MCP etc. hooks, and so on.
Having seen the type of discussion both on the PR and on Mastodon, I felt no desire to have these particular people holding open judgement over me and my project. [Aside: I'm fairly certain someone has gone through the PR that proposed the ban and removed egregious comments ... there are some comments missing that were formerly there, wish I had archived them]
Take the comments on this Mastodon thread: https://wandering.shop/@silvermoon82/116964029691087138 "I'm sorry you caught abuse off the slop fondlers. Consent and boundary issues are inbuilt in their community."
Or here, when someone tried to ask for similar clarification: https://tldr.nettime.org/@tante/116880003584050912 "You're coming off similar to someone trying to probe the limits of consent instead of trying to honor the request in good faith. It's a bad look." "Opening your conversation with a group by saying "well I'll get to violate your group's values a _little_, right?" does not project the mentality of someone who wants to collaborate in good faith."
Comparing AI usage or questions about it to rape culture or violating consent is absolutely repulsive and trivializes the latter and exaggerates the former.
Or making analogies to Claude Code users being NAZIs? "@piggo @tante I see such a rule as being similar to the "no Nazis" rule. It works by Nazi sympathizers self-identifying by arguing about the rule."
Pre-defining an in-group or out-group based on how contaminated they are by the tools they use... I'm sorry, it's ridiculous. And not in the slightest "democratic" even if it holds a formal vote.
It's possible that nuance would prevail when the people sitting in judgement over things happened to look at my project. The actual public blog post seems going out of its way to provide a bit of "don't worry" ... but none of that was provided before. And I felt no need to be under that microscope.
... and so then the flood of "good riddance, so glad we got rid of people like you" comments... and people ritually going through my other comments to downvote me, blah blah. It's all the hallmarks of an angry echo chamber.
I posted this elsewhere about the kind of un-dialectical purity-of-essence thinking I saw on display:
[I fear] Blunt purity tests administered by people with reductionist and categorical thinking.
The reality is that even tools produced by concentrated capital and mass appropriation can also increase the productive capacity of a small free-software project whose code remains available under copyleft. They're not "infected" and contagious. They're tools.
They can deskill and dispossess workers; they can extend what a particular worker is able to make. Which tendency dominates, under what conditions, and to whose benefit are material questions.
"AI bad" is not an answer to them, any more than tech-bro "innovation good" is.
Their comparison was to a no-Nazis rule, not to Nazis. That comparison works. Only vibecoders argue about the intricacies of a no-vibecoding rule, just like only Nazis argue about the intricacies of a no-Nazis rule. They aren't saying vibecoders are Nazis, they're saying both vibecoders and Nazis are dishonest and try to find loopholes in the rules and accidentally identify themselves by doing that. The people who argue about age of consent also do that, which is why the person mentioned rape.
More worrisome is I saw this same set of comparisons made over and over in multiple threads.
Like, it's clear many people in this community see themselves as violated in some sense by the actions of American AI companies. The emotional response is legitimate. Transitioning from that emotion to indiscriminate policy is not, and making comparisons like that in a public forum is distasteful.
As for "if someone doesn't know whether their project is mostly AI-written, they're an idiot or lying. Nobody doesn't know that" -- that is not the question. The question is the precise definitions of what "AI-written" means, because people legitimately wanted to know if their projects were going to be voted off the island. Because like it or not based on your moral principles -- vast amounts of people are using these tools now. I can't say a majority because I don't have the numbers but it would absolutely not surprise me. And I can guarantee a large number of projects on codeberg use(d) them. So clarify in advance of the policy passing what the policy will be. Instead the language was left ambiguous so they could choose to exercise personal judgement.
Calling them stupid or dishonest is a dick move. People legitimately wanted to know.
Are they correct? That should be the most important thing.
> More worrisome is I saw this same set of comparisons made over and over in multiple threads.
Maybe because they're correct
> because people legitimately wanted to know if their projects were going to be voted off the island.
Thus outing themselves as AI projects. Nobody with a non-AI project doesn't know that it's non-AI.
I'll move on now.
I suspect the reason you're not saying this is because when you write it out, it's obviously a policy rooted in simplistic thinking, similar to "We won't host code if it was written in Emacs" or "We won't host code if it uses spaces instead of tabs."
Your second paragraph is a strawman.
In world where we cannot copyright facts, but we can copyright compendiums, we've demonstrated that we acknowledge the effort and value of curation. The way I use AI, I'm carefully curating what makes it into the released version, either by writing it myself or iterating with an AI to get it written.
You're just being dismissive.
I can understand why. The cost to the host is entirely in volume of code, and creating volume is cheap now. Without some constraining measure they will have lots of AI code. Even just by itself that would make it impossible to have.
It’s also a cultural thing. This is for that niche of developers. That’s all right. Communities have their own (often-hyperspecific) rules. HN certainly does and we’re all here.
I don't get the anti-AI sentiment.
From the maintainer standpoint, it basically eliminated entire classes of frustrating, tedious jobs that existed before - especially with regards to packaging and CI/CD
No longer do I have to do the back and forth dance of getting slightly further into a 20-minute build over the course of an evening until GitHub Actions finally produces a correct build. I just log into `gh`, give Claude the specs ("Please make it as statically-linked wherever sensible, support these platforms xyz back to versions abc") and it loops 2 or 3 times itself until the build is done.
If a user posts an issue, Claude can track it down in 15 minutes flat. Our issue backlog was cut by around 70% in a couple of months after Opus 4.5 was released - I could fix 5 issues in a coding session rather than 1, and new users who would vibecode a fix would often upstream the solution via PR.
As a user, deployment of many FOSS apps has gone from a nightmare to 15 minutes of prompting. I have a Claude skill that has the user accounts, IP addresses and configs of the servers. You just say "Install Immich on server xyz" and it does it. My wife says she wants the Immich photos on the Chromecast, I sign into adb and say "Make the Immich photos the screensaver on the Chromecast", and it installs ImmichFrame.
If I run into issues, instead of just giving up I'll just prompt a fix. I'll often clean things up, draft a PR and upstream it - and to be honest results have been mixed. Some maintainers (myself included) are appreciative that someone has upstreamed a fix that others can find useful, others are actively hostile to it and plenty more are just too busy to review the PRs.
I do empathise with them about the wasted CI/CD resources for small projects with a handful of users, but the rest just sounds academic and out-of-touch. From my own experience as an early adopter I've found that AI has massively increased my participation in the FOSS sphere - as a maintainer, contributor and consumer.
There are valid social, ecological and maybe philosophical concerns.
But on the technical aspects... I don't get it either. It's a wonderful tool. even if it didn't write a single line, the fact that it can analyse a source and explain things to me in context is a game changer.
My issue however, is that this is changing the rules on people. Codeberg initially was a FLOSS alternative to big company Github. Now it's more narrow, an alternative to Github and only for people who don't want to use AI. What's the next step? This is the problem, you either a platform and be, within legal reasons, agnostic, or you are trying to be a specific community. It is unclear to me if the specific community that Codeberg wants to be today will be the one it wants to be in a year from now, because clearly it's changing.
I cannot rely on a service provider that does change that quickly.
Codeberg never was a full alternative to Github. From the start it's been an open platform, run by European non-profit, exclusively for Free and Open Source projects. I see the recent changes to be inline with their initial values and community they want to foster.
As the other commenter said, they didn't change their values and target.
Also, these changes are enacted by members of the e.V. Nobody is forcing them to do something. It's all public and transparent.
In the blog post they complain about AI-assisted projects being too polished, using CI resources even though it's a one-man project with zero users. (Reasons they don't give for banning it in the TOS)
So, what's the rule exactly? Can I set up release builds on my one-man project or not? What about when the clearly vibe-coded project does it, do I report them? Does everyone just remember to not check in their AGENTS.md when pushing to codeberg?
I think things should just cost money. Far better than a nonsensical honor system around LLMs which has the obvious outcome of hiding LLM use.
That said, they can foster what ever environment they want. I think it's just a regressive self-limiting view that tries to pretend like the world isn't changing with rules that have bad downstream outcomes, as opposed to a platform that accepts reality and figures out what to do from there. And I find that disappointing.
Codeberg has some of these features but the networking effect there is pretty weak because there are way fewer people. Almost Every programmer is on Github.
Codeberg bans vibe coded projects
Regarding discussion here o democracy - pure democracy can lead to suppression of minority views. Democracy can mean 6 / 10 people voting to beat the crap out of the other 4. The US, for example, is not a democracy, it is a constitutional republic, with checks and balances to prevent such abuse. Not perfect by any means, but doing pretty well compared to the alternatives available today.
The US is both. Or, to phrase it another way, it is neither a pure democracy nor a pure republic. That's a good thing, it helps to reduce the downsides of both.
Fair enough if Codeberg isn't the correct place for stupid little personal hobby projects, but that wasn't the sense I got when I signed up.
I get the feeling I should move the repos to my own Forgejo instance out of respect soon enough, but I'm certain they will not be deleted any time soon, at least not without me being notified of it. It's a community, not a corporation.
I would hope so, that split would be a natural parting of the communities.
My pessimism tells me that the more any community rejects LLM and polices against generated output, the heavier they would be crawled. Which means that proLLM community by nature puts pressure on all antiLLM infrastructure – an asymmetrical problem, where other people pay for your choice.
My optimism, on the other hand, tells me that after bubble pop, hopefully, far less people would have money to feed aimless crawl results into machine learning algorithm in hopes of slight improvements.
I don’t currently use LLMs for any significant part of my work but I’m not willing to put myself in a position where I might someday have to prove that something WASN'T written with an LLM because someone is making spurious accusations to troll or "DoS" a project.
It's unlikely to happen, but there are plenty of alternatives where this isn't a concern at all. Why open yourself or your projects to this?
This seems fully orthogonal to whether LLMs are used or not
That's not healthy democratic culture.
Since when is Codeberg responsible of the Internet culture?