> To the others: I apologize to the world at large for my inadvertent, naive if minor role in enabling this assault.
this is my position too, I regret every single piece of open source software I ever produced
and I will produce no more
> To the others: I apologize to the world at large for my inadvertent, naive if minor role in enabling this assault.
this is my position too, I regret every single piece of open source software I ever produced
and I will produce no more
The Open Source movement has been a gigantic boon on the whole of computing, and it would be a terrible shame to lose that ad a knee jerk reaction to genAI
it's not
the parasites can't train their shitty "AI" if they don't have anything to train it on
this is precisely the idea
add into that the rise of vibe-coding, and that should help accelerate model collapse
everyone that cares about quality of software should immediately stop contributing to open source
I see this as doing so at scale and thus giving up on its inherent value is most definitely throwing the baby out with the bathwater.
My point was that the hypothetical of "not contributing to any open source code" to the extent that LLMs had no code to train on, would not have made as big of an impact as that person thought, since a very large majority of the internet is text, not code.
Unless you're in the camp that believes ChatGPT can extrapolate outside of its training data and do computer programming without having ever trained on any computer programming material?
It will however reduce the positive impact your open source contributions have on the world to 0.
I don't understand the ethical framework for this decision at all.
I'm not surprised that you don't understand ethics.
I couldn't care less if their code was used to train AI - in fact I'd rather it wasn't since they don't want it to be used for that.
which is the exact opposite of improving the world
you can extrapolate to what I think of YOUR actions
my comments on the internet are now almost exclusively anti-"AI", and anti-bigtech
My position on all of this is that the technology isn't going to uninvented and I very much doubt it will be legislated away, which means the best thing we can do is promote the positive uses and disincentivize the negative uses as much as possible.
they're using your exceptional reputation as a open-source developer to push their proprietary parasitic products and business models, with you thinking you're doing good
I don't mean to be rude, but I suspect "useful idiot" is probably the term they use to describe open source influencers in meetings discussing early access
IMHO their are going to be consequences of these negative effects, regardless of the positives.
Looking at it in this light, you might want to get out now, while you still can. Im sure its going to continue, its not going to be legislated away, but it's still wrong to be using this technology in the way it's being used right now, and I will not be associated with the harmful effects this technology is being used for because a few corporations feel justified in pushing evil on to the world wrapped positives.
I think of LLMs as more like the invention of cars or railways: enormous negative externalities, but provided enough benefit to humanity that we tend to think they were worthwhile.
Are the negatives of LLMs really that bad? Most of them look more like annoyances to me.
The ones that upset me the most are the ChatGPT psychosis episodes which have lead to loss of life. I'm reassured by the fact that the AI labs are taking genuine steps to reduce the risk of that happening, which seems analogous to me to the development of car safety features.
If bringing fire to a species lights and warms them, but also gives the means and incentives to some members of this species to burn everything for good, you have every ethical freedom to ponder whether you contribute to this fire or not.
For your fire example, there's a difference between being Prometheus teaching humans to use fire compared to being a random villager who adds a twig to an existing campfire. I'd say the open source contributions example here is more the latter than the former.
"It barely changes the model" is an engineering claim. It does not imply "therefore it may be taken without consent or compensation" (an ethical claim) nor "there it has no meaningful impact on the contributor or their community" (moral claim).
if true, then the parasites can remove ALL code where the license requires attribution
oh, they won't? I wonder why
There's also plenty of other open source contributors in the world.
> It will however reduce the positive impact your open source contributions have on the world to 0.
And it will reduce your negative impact through helping to train AI models to 0.
The value of your open source contributions to the ecosystem is roughly proportional to the value they provide to LLM makers as training data. Any argument you could make that one is negligible would also apply to the other, and vice versa.
Not if most of it is machine generated. The machine would start eating its own shit. The nutrition it gets is from human-generated content.
> I don't understand the ethical framework for this decision at all.
The question is not one of ethics but that of incentives. People producing open source are incentivized in a certain way and it is abhorrent to them when that framework is violated. There needs to be a new license that explicitly forbids use for AI training. That may encourage folks to continue to contribute.
In both cases I get the frustration - it feels horrible to see something you created be used in a way you think is harmful and wrong! - but the world would be a worse place without art or open source.
Well maybe the AI parasites should have thought of that.
I would never have imagined things turning out this way, and yet, here we are.
Rather, I think this is, again, a textbook example of what governments and taxation is for — tax the people taking advantage of the externalities, to pay the people producing them.
The open source movement has been exploited.
The exploited are in the wrong for not recognising they're going to be exploited?
A pretty twisted point of view, in my opinion.
For instance, does Richard Stallman fit this mold?
If you disagree, please explain how RMS and/or Perens do not fit the mold.
Stallman cared / cares mostly about user freedom, but was canny enough to understand that businesses would also been to be able to engage with this freedom too.
Copyleft licensing was put in place _because_ commercial exploitation was expected. It was designed to preserve user freedoms.
Compromise was built in the model.
But the unexpected twist was cloud; which broke the safety mechanism.
This is the reason I feel it's unfair to say that proponents of the movement are naive. Exploitation was predicted from the outset. A complex turn of events drew the shape of the current landscape.
And if blame is to be apportioned anywhere, it should be firmly at the feet of the corporations profiting.
But, anyway, I do not think your analysis captures the situation in the right way. The free-rider problem, where users contribute no/negligible code or money back to FOSS, is the heart of the exploitation; Tivoization or cloud SaaS-ification are merely forms of free-riding. Other forms of the free-rider problem would have eventually become a thorn in the side of FOSS even if those two things had never happened and there is no way to plug that hole in the concept of copyleft.
And I maintain that was entirely predictable (and it was predicted by many a few decades ago!): there is no reason for a business owner to contribute back to FOSS when not contractually obligated to do so. Like the fable of the scorpion and the frog, even if it's valid to do so it's kind of pointless to blame capitalists for doing what everyone knew what they were going to do all along.
These corporation are not run by people who have no choice, they're run by people who choose to run the system to the absolute limit for absolute material gain.
All the FAANGs have the ability to build all the open source tools they consume internally. Why give it to them for free and not have the expectation that they'll contribute something back?
Are there any proposals to nail down an open source license which would explicitly exclude use with AI systems and companies?
But for most open source licenses, that example would be within bounds. The grandparent comment objected to not respecting the license.
Most companies trying to sell open-source software probably lose more business if the software ends up in the Debian/Ubuntu repository (and the packaging/system integration is not completely abysmal) than when some cloud provider starts offering it as a service.
Because it is "transformative" and therefore "fair" use.
As an analogy, you can’t enforce a “license” that anyone that opens your GitHub repo and looks at any .cpp file owes you $1,000,000.
The fact that they could litigate you into oblivion doesn't make it acceptable.
Even if you could construct such a license, it wouldn't be OSI open source because it would discriminate based on field of endeavor.
And it would inevitably catch benevolent behavior that is AI-related in its net. That's because these terms are ill-defined and people use them very sloppily. There is no agreed-upon definition for something like gen AI or even AI.
"The only thing that matters is the end result, it's no different than a compiler!", they say as someone with no experience dumps giant PRs of horrific vibe code for those of us that still know what we're doing to review.
Anyone can use your software! Some of them are very likely bad people who will misuse it to do bad things, but you don't have any control over it. Giving up control is how it works. It's how it's always worked, but often people don't understand the consequences.
As they say, "reduce, reuse, recycle." Your words are getting composted.
no, it hasn't. Open source software, like any open and cooperative culture, existed on a bedrock, what we used to call norms when we still had some in our societies and people acted not always but at least most of the time in good faith. Hacker culture (word's in the name of this website) which underpinned so much of it, had many unwritten rules that people respected even in companies when there were still enough people in charge who shared at least some of the values.
Now it isn't just an exception but the rule that people will use what you write in the most abhorrent, greedy and stupid ways and it does look like the only way out is some Neal Stephenson Anathem-esque digital version of a monastery.
If you care about what people do with your code, you should put it in the license. To the extent that unwritten norms exist, it's unfair to expect strangers in different parts of the world to know what they are, and it's likely unenforceable.
This recently came up for the GPLv2 license, where Linus Torvalds and the Software Freedom Conservancy disagree about how it should be interpreted, and there's apparently a judge that agrees with Linus:
https://mastodon.social/@torvalds@social.kernel.org/11577678...
But you can be sure that even the risk-adverse companies are going to go by what the license says, rather than "community norms."
Other companies are more careless.
But AI is also the ultimate meat grinder, there's no yours or theirs in the final dish, it's just meat.
And open source licenses are practically unenforceable for an AI system, unless you can maybe get it to cough up verbatim code from its training data.
At the same time, we all know they're not going anywhere, they're here to stay.
I'm personally not against them, they're very useful obviously, but I do have mixed or mostly negative feelings on how they got their training data.
Might be because most of us got/gets payed well enough that this philosophy works well or because our industry is so young or because people writing code share good values.
It never worried me that a corp would make money out of some code i wrote and it still doesn't. AFter all, i'm able to write code because i get paid well writing code, which i do well because of open source. Companies always benefited from open source code attributed or not.
Now i use it to write more code.
I would argue though, I'm fine with that, to push for laws forcing models to be opened up after x years, but i would just prefer the open source / open community coming together and creating just better open models overall.
Some Shareware used to be individually licensed with the name of the licensee prominently visible, so if you had got an illegal copy you'd be able to see whose licensed copy it was that had been copied.
I wonder if something based on that idea of personal responsibility for your copy could be adopted to source code. If you wanted to contribute to a piece of software, you could ask a contributor and then get a personally licensed copy of the source code with your name in every source file... but I don't know where to take it from there. Has there ever been some system similar to something like that that one could take inspiration from?
Most objections like yours are couched in language about principles, but ultimately seem to be about ego. That's not always bad, but I'm not sure why it should be compelling compared to the public good that these systems might ultimately enable.
Nah, don't do that. Produce shitloads of it using the very same LLM tools that ripped you off, but license it under the GPL.
If they're going to thief GPL software, least we can do is thief it back.
which they don't
and no self-serving sophistry about "it's transformative fair use" counts as respecting the license
Characterizing the discussion behind this as "sophistry" is a fundamentally unserious take.
For a serious take, I recommend reading the copyright office's 100 plus page document that they released in May. It makes it clear that there are a bunch of cases that are non-transformative, particularly when they affect the market for the original work and compete with it. But there's also clearly cases that are transformative when no such competition exists, and the training material was obtained legally.
https://www.copyright.gov/ai/Copyright-and-Artificial-Intell...
I'm not particularly sympathetic to voices on HN that attempt to remove all nuance from this discussion. It's challenging enough topic as is.
thankfully, I don't live under the US regime
there is no concept of fair use in my country
There's a huge political aspect here: copyright hasn't worked for decades (I've written about this at length), and this is the latest iteration in that erosion. Countries that enforce IP as a natural right are going to have trouble navigating the change: they either need to avoid AI entirely (this will have higher costs than many anticipate), or they need revise how they think about copyright. Or they can just ignore it. There are no good options.
My instinct is that countries that embrace change will do better.
What a joke. Sorry, but no. I don't think is unserious at all. What's unserious is saying this.
> and the training material was obtained legally
And assuming everyone should take it at face value. I hope you understand that going on a tech forum and telling people they aren't being nuanced because a Judge in Alabama that can barely unlock their phone weighed in on a massively novel technology with global implications, yes, reads deeply unserious. We're aware the U.S. legal system is a failure and the rest of the world suffers for it. Even your President routinely steals music for campaign events, and stole code for Truth Social. Your copyright is a joke that's only there to serve the fattest wallets.
These judges are not elected, they are appointed by people whose pockets are lined by these very corporations. They don't serve us, they are here to retrofit the law to make illegal things corporations do, legal. What you wrote is thought terminating.
Thanks for your contributions so far but this won't change anything.
If you'd want to have a positive on this matter, it's better to pressure the government(s) to prevent GenAI companies from using content they don't have a license for, so they behave like any other business that came before them.
I fixed it... Sorry, I had to, the quote template was simply too good.
Yes.
I don't see how "We couldn't do this cool thing if we didn't throw away ethics!" is a reasonable argument. That is a hell of a thing to write out.