It is truly amazing how many people will shill for these massive corporations that claim they love open source or that their AI is open while they profit off of the violation of licenses and contribute very little back.
It is truly amazing how many people will shill for these massive corporations that claim they love open source or that their AI is open while they profit off of the violation of licenses and contribute very little back.
Even the license itself states it's optional, and you don't have to agree it (if you don't, you get the copyright law's default).
Author of the article is a former member of the Pirate Party and EU parliament, so they have expertise in the copyright law.
So the same persons that supported Napster and the Pirate Bay now want to circumvent copyright for open source software.
An unholy alliance, but the recent comments from some Microsoft brass about everything on the Web being freeware seems to indicate that these are the talking points that Microsoft and its new allies will put out.
I expect that people professionally dedicated to a copyright reform are very familiar with it, regardless of which way they want to reform it.
The copyright laws were written before generative AI existed, so they may not be adequate or fair in the new reality, but that's the current state anyway. As Reda notes, the law is not specific enough to draw the difference between collecting and processing data for search engines (that may be using ML for retrieval) and using the same data with LLMs.
Frequency signal data over an image are not the image, but no one argues a JPEG encoded copy of a PNG isn't the same image. I think the weights vs code are similar in that regard.
As for releasing weights, probably more if we're talking about AGPL code.
There's existing analogies like encyclopedias and dictionaries.
One interesting aspect to those sorts of consolidation works is that they may contain errors and other artifacts, specifically to identify duplications of their work vs new from-scratch work.
It also feels similar to the recent article posted about photography and how during its early days pictures were used for advertising without the consent of those photographed. [0]
[0] https://www.truthdig.com/articles/the-troubled-development-o...
However, the broader issue that Microsoft has infiltrated OSS and its organizations successfully by hiring and donating remains. It would not surprise me at all if they now hire people with an ostensibly "freedom fighter" background for credibility.
Look at how many people here cite his (former?) membership in the Pirate Party for credibility! Party membership means nothing. Politicians (in general!) change their minds, can be bought, etc. The Green Party in Germany started out as a peace party and has been used repeatedly to lend credibility to the Kossovo and other wars.
https://blog.okfn.org/2022/03/03/microsoft-to-support-open-d...
Today, we are pleased to announce that Microsoft will once again be supporting Open Data Day by providing mini-grants to organisations to help them run events, the call will launch on Open Data Day 2022.
They also supported "Open Data Day 2021". Sounds like a nice trojan horse to influence EU legislation through purported activists.
What about BSL, SSPL, or other source available (for your eyes only) licenses? Copilot harvests all public repos, regardless of its license.
To simplify:
- imagine all code Copilot trained on is GPL licensed. - we have a universal function `isInfringing(code)` that has access to all GPL code, and returns `true` if it is infringing some GPL code.
for a given prompt; if `isInfringing(copilot(prompt))==false` we cannot claim copilot infringing on GPL code, even it is trained on GPLed code.
so the problem starts here; does the piece of code copilot emits, if written by yourself also would be infringing ?
why everyone on discussions tries to bring "if a human made it"? a generative AI operates way faster than anyone ever existed and ever will and probably a person aware of the license & acting respectful towards it, will create something more sensible/plausible to avoid plagiarism
now having dozen/hundreds/thousands of humans substituted by a machine that makes money for some for-profit company is really fair? even if they were a non-profit, as someone pointed up, people who create the content that feeds the weights aren't recieving a penny! they already made money with it, they will make more & that is/will upgrading/e the state of gen. AI
for sure legal battles on people copying code from permissive licenses should exist but it's feels a different discussion
At what measure is one sufficiently inspired for it to be a derivate work? That is up to courts to decide.
Technically when you get a piece from each, there is no infringement legally. ( as they have all different copyright holders )
If CoPilot indeed derives a function (or a functional block) from a single source, it might plainly violate the license of the repository where it derives the code from.
There are many questions, and nothing is clear cut. The only thing I know is, I will never use that thing.
EDIT: I remembered that people were able to make CoPilot emit their code almost as-is with the correct prompts: https://x.com/docsparse/status/1581461734665367554
So it's not we're taking a bit from n different sources, and generate something with that.
False in ex-Commonwealth countries and Japan.
And for which jurisdictions has this been established? What is the legal argument that "weights" are derivative but output is not?
I'm surprised it's so clear cut as you say, but I haven't really been following the whole kerfuffle.
If you can get a decent facsimile of licensed code out the other end, how is it really any different from lossy compression? I doubt the courts would consider a lossy re-encode of a disney movie as free from copyright.
That's not really Microsoft's problem as long as people aren't afraid of using Copilot to generate (potentially GPL'd) code. And from what I've generally seen from genAI discussions at work, people think very little about any legal implications.
However, copyright isn’t cooties. If the output is not similar, then it is not infringing regardless of how much GPL’d training data was used to generate it.
>"How much change is enough" has always been a gray area for courts and humans to decide.
But copilot has been shown to generate chunks of sufficient size and specificity that as a layman it very much feels like "copied GPL code". And my boss agrees too - we have a blanket ban on generative AI tools in our work because it's not considered worth the risk.
Only when given chinks of copyrighted code as input. I don't think anyone has demonstrated big chunks of copyrighted code in the output when copyrighted code isn't present in the query/context.
In fact, I suspect microsoft specifically filters the output for that.
No, of course not. We should probably revisit copyright law, given that it was written at a time when no-one foresaw modern AI tools, its capabilities, and its effects on creators and societies.
Free Software already views all proprietary software as inherently immoral. So there is no need to take a detour of what went into making the software to reach that conclusion from that angle.
Transformers predict the most likely next token, the most likely next token is usually related to the surrounding context.
So yes it can create code similar to GPL code but it can only do that consistently when the GPL code is included in the context. So don’t do that.
This is not immediately obvious to me.
A small though experiment: the Harry Potter books are clearly copyrighted works. If I generate a frequency list of all words in these books, i.e. a list of all words and how often they appear, that frequency list is derived from the original work, in the normal way we would use the word "derived". But is it a "derivative work", under the strict legal definition of this term?
There are other problems with releasing model weights under the GPL. It just doesn't fit, in the same way as releasing non-software under the GPL doesn't make sense.
Calling the output of generative AI copyrightable violates the spirit of copyright, as it is neither creative nor labor-intensive. We could quibble about that, but I think we can at least agree that the point is that this generative AI stuff requires very little skill to use in most cases and can't operate without prior art to train on. Other lame stuff has been copyrighted before, like paint splatters and stuff, but even that type of art appears to involve more skill than entering a few words into a generative AI.
EU courts disagree:
> Under European copyright law, scraping GPL-licensed code, or any other copyrighted work, is legal, regardless of the licence used.
edit wording about the shill