https://docs.github.com/en/site-policy/github-terms/github-t...
In addition, one has to consider whether the act of studying a project whose sources you have legally obtained to gain experiences applied in other contexts where the original license may not apply or be upheld is an issue. This applies to humans as much as to ML.
A couple things - you agree github can do it, ostensbly taking archives for backups sure, but in this case it's to distribute out to some 3rd party and explicitly to be locked in a vault with no business purpose.
My personal uploading of say - a GPL licensed code, or hell, microsofts leaked code DOES NOT magically grant github to do whatever they want with it. There's already a license, no?
from the github archive page: > Archiving software across multiple organizations and forms of storage helps to ensure its long-term preservation. "The Arctic Code Vault"
> On 02/02/2020 GitHub captured a snapshot of every active public repository. Those millions of repos were then archived to hardened film designed to last for 1,000 years, and stored in the GitHub Arctic Code Vault in a decommissioned coal mine deep beneath an Arctic mountain in Svalbard, Norway. Our partners include the Long Now Foundation, Software Heritage, the Internet Archive, Microsoft Research’s Project Silica, the Arctic World Archive, GHTorrent, and GHArchive. Our advisors include both technological visionaries and world-renowned experts in the humanities.
so, like, I am sure it was vetted by lawyers, but my plain english understanding of that is github may need to archive it, whats not clear is that github may send it to any number of non profits to do with it what it pleases.
Then you've got software heritage who literally just scrapes the web for cgit repos and copies them without any agreement or license between you as the host and software heritage.
Anyway, it's all very Uber-model. Fuck the legality, we're just gonna do it and fuck it because the worst that can happen is a slap on the wrist.
Granted, the machine has to learn to not directly plagiarize in the same manner humans are not allowed to when they use acquired knowledge - unless the license allows it of course - but the act of studying material that is allowed to be read cannot be considered harmful.
1) The LLM owners really can't guarantee that it won't directly plagiarize without attribution or licensing. Your code may contain a unique algorithm or method for solving something, and when someone asks the right question, your code may simply be the only answer it knows to give.
2) While the code being used as training input was open source and visible to the public to learn from, the models being built often aren't. It seems unethical to train from public data yet keep the resulting weights private and charge for access to use the trained weights.
1) These tools are generally used in a pair-programming fashion, and in that function, the output can be considered similar to when you ask a coworker on Slack and they paste you a snippet, or if you browse github and read someone else's implementation (without having the LICENSE text within your field of view at all times). A possible violation would then only occurs once the snippets in question are included into your code-base and distributed in ways that violate the original license.
2) One could argue that sharing the snippet with you was a form of redistribution, but I would not consider this to apply if a human did it and would therefore not apply it to machines either, and I do not think that is what people generally consider redistribution of an open source project. GPL technically has a clause second-hand violations, but I do not think that one holds.
It should also be noted that licenses like MIT only require the copyright and permissions notice included in substantial portions of the program, and so smaller snippets are always fine. Humans also do not bother attributing smaller copy-paste blocks - we'd run out of storage linking to all the stackoverflow answers!
3) The issue gets a bit hairier when the machine reproduces large/important portions of projects with no hint as to its source, license or ways to do proper attribution, but even then I'd consider the violation to occur only if included verbatim into a project which is then redistributed under incompatible terms.
4) Even when code is largely identical, it generally only an issue if the code is a unique invention, not if the code trivially follows for a skilled individual of the trade. That's a principle in the practice of many laws, including patent law.
For the second aspect, I do not see any importance to the fact that the trained model is not public. A person studying open-source projects do not upload a brain dump afterwards, and others only directly benefit from their experience (their "weights") if they decide to teach the subject. Nor is every project they write afterwards with their knowledge necessarily open-source, only being public if they want to make them public. Licenses generally do not restrict private or internal usage, including modification and derivative works. It is redistribution they trigger on (with some catches for things like AGPL).
(I would of course like the model to be public for the betterment of mankind, but that's different from the legal aspect of it.)
It's like Common Crawl, another non-commerical mass scraping project just benevolently stealing creations for "AI" companies.
> To request the removal of a content from the Software Heritage archive, you must file a formal request containing all of the following informations:
> (...)
> Please send your request by e-mail to takedown@softwareheritage.org
https://www.softwareheritage.org/legal/content-policy/
The HN thread you posted even shows that they contacted the forge's admin week in advance to check for possible concerns.