IP law (like a lot of other things) has been skewed toward the interests of business, even when that conflicts with fairness or societal good. For all its flaws (IMO), the free software movement tends to be principled. Just because something is legal doesn’t mean it’s right.
News back then were about intentionally prompting to output known copyrighted material.
The parent comment still stands in my opinion:
When, despite millions of developers using agentic AI already, are these lawsuits supposed to manifest?
I would be surprised if a frontier model generated unexpected copyright headers during typical usage.
> News back then were about intentionally prompting to output known copyrighted material.
First, there are other cases if you take the time to dig. This is quite an old example (GPT-2) as i haven't kept up to date on this field recently, but it does show that this problem has been known about since before these systems were widely adopted: https://arxiv.org/abs/2012.07805 [0]
Second, GP said nothing about the type of effort required to make it happen, just that it can be done and that the copyright owner could come along and cause legal problems later. It's absolutely possible to have a fly-by contributor who purposefully asks for code that reproduces X/Y/Z without a maintainer knowing about it.
But then the maintainer is the one in legal trouble.
> When, despite millions of developers using agentic AI already, are these lawsuits supposed to manifest?
Legal / copyright / etc. cases often take a lot longer than a couple of years to come to fruition.
---
[0]: edit -- to clarify this is an example of the reproduction problem, not an example copyright infringement case.
The concern discussed here is copyrighted material being generated unintentionally and the original author asserting their rights.
This has, to my knowledge, not happened once.
If we are not talking about unintentional violations, I don't understand the point of the discussion.
I can also intentionally copy paste the copyrighted material into my merge request without the use of AI in an attempt to get the maintainer into trouble.
both intentional (malicious contributor) or unintentional (Large-Laundering-Model) are copyright issues -- which is the point of GCC's policy.
> I can also intentionally copy paste the copyrighted material into my merge request without the use of AI in an attempt to get the maintainer into trouble.
You can. You can also do it significantly faster with significantly less effort while being harder to detect using agents etc.
That is obviously not what anyone was referring to, nor does it make sense, when there is a much more reasonable basis to prohibit the same contribution.
Namely inserting vulnerabilities. This one actually happened before afaik, and provides a clear benefit to the attacker.
If a contributor doesn't care about submitting copyrighted code, they can do it without an LLM as well.
Plenty of github accounts now are agent-driven monstrosities just trying to inflate someone's contribution stats etc.
Someone tried to contribute “vibe-coded” device support to a project I’m involved with, they said they did it all based on the device documentation, the code their agents spit out was copied verbatim out of a (GPL’d) project with which I’m familiar which supports that device.
LLMs are not learning things and then using that learning to construct new things. They are essentially a form of lossy compression of their training set. And you don’t need to be explicit about trying to reproduce a portion of that training set for an LLM to output one.
I am not aware of any study attempting to measure unintentional reproduction.
With your example, I question whether you have seen this happen first hand. For all I know, the contributor could have explicitly prompted the model to reference the GPL project and had the agent clone the code from the web.
the "entire history of" is circa 3-4 years, which is very much a tiny period of time compared to normal legal system / copyright law stuff (IANAL).
alternative perspective: it's just taking time for the lawyers to figure out what they can sue them for.
It’s a hard balancing act to do. Give in too much randomness and you get non-sensical outputs that are difficult to align. Fit too closely to the training data and the model regurgitates the training data.
And oh, what’s that copyrighted material we never made any agreement to use doing in there?
> The complaint argued that "the basis of the Gaye defendants' claims is that "Blurred Lines" and "Got To Give It Up" "feel" or "sound" the same. Being reminiscent of a "sound" is not copyright infringement. The intent in producing "Blurred Lines" was to evoke an era. In reality, the Gaye defendants are claiming ownership of an entire genre, as opposed to a specific work"
they lost (eventually) https://en.wikipedia.org/wiki/Pharrell_Williams_v._Bridgepor...
wider point -- whether or not a copy is a copy and whether it is is infringing on copyright or not ultimately has to be decided by a court case when it's not an obvious and clear cut violation. especially in the USA with the utterly mental fair use law.