Microsoft has been super litigious in the past when it came to copyright violation starting all the way back with Bill Gates' letter in Byte magazine about those pesky pirates. To see them do this makes pirating MS software fair game from here on. They could have asked nicely, instead they just took.
It is. Laws are adapted based on widespread technological capabilities and progress.
As an example, if it is easy to create real voice or signature using AI models - they should no longer be considered effective evidence for contractual reason instead of enforcing that it is illegal to forge it. That is not going to work.
Past shouldn't dictate what we allow tomorrow.
All laws are made in interest of someone.
Does the copyright apply to AI models since they are out of scope and weren't widespread when it came into force?
Does the proposed benefit in the original law apply in practice?
Are they more beneficial than the progress allowed by AI models who use them as training data?
Is the copyright law practically enforceable on output generated by AI models?
https://en.wikipedia.org/wiki/Berne_Convention
Copyright is what FOSS depends on. For Microsoft to shit all over GitHub contributors rights is despicable.
Does the same apply to a human? Do we now define copyright violation differently for computers? I don‘t know the perfect answer here. But I‘m not so sure we should have standards that change depending on if a program is doing it or a human is doing it. Perhaps a bad standard to begin with.
I do tend to learn towards thinking „company uses publicly available, open source code in product“ is somewhat of a nothing-burger though.
EDIT: After more reading on the subject, I'm willing to accept that copyright infringement is unlikely here. This link [1] was the one I found most convincing.
However, I would still shift the goalposts and look at this ethically, and I still think it's wrong that Microsoft is profiting from code with licences like GPLv3. This is a whole other topic, though.
[1] https://www.technollama.co.uk/is-githubs-copilot-potentially...
Copilot's API is surfacing snippets of work without licensing information attached alongside. It can be shown in discovery that Copilot does access the origin work.
The sooner this is slapped down, the sooner we can avoid addressing the even more troubling question that exists today: is someone who used Copilot to throw together a bunch of code infringing copyright of works where those portions originate?
This is a complex problem with no satisfying conclusions... how could one be violating copyright if they never accessed the 'copied' work to copy? Copyrights aren't patents. Infringement requires copying.
Using Copilot launders the user's awareness of the origin works, yet making the Copilot users liable for widespread "accidental" copying would be troubling.
That you got the code from and entity that stole it somewhere else doesn't really matter. Generative models should respect copyright for their sources, and using a generative model to create new works that you intend to claim copyright on is stupid: someone may well show up one day with ironclad proof that you used their code without permission.
Also a reminder that outside the copilot debate, the online rights movement has largely been pushing for scraping, deep linking and transforming scrapped data to not be considered copyright infringement, regardless of any TOS on the site being scraped.
To me, co pilot is a exactly that, a scraper that has scraped public websites and is now presenting me the scraped data in an alternative and often transformed form. It’s my responsibility as a developer to ensure that my released product complies with applicable copyright law, but copilot and the use thereof is not in and of itself copyright infringement.
That a tool can be used to create infringing work or infringe on copyright in general is no more a valid argument against co pilot than it is against CD burners, de-drm tools, vcrs, kodi or plex, scanners or any number of day to day items that have the ability to infringe copyright if the user uses it for that purpose.
https://decoded.legal/blog/2021/06/github-copilot-initial-th...
https://fossa.com/blog/analyzing-legal-implications-github-c...
https://felixreda.eu/2021/07/github-copilot-is-not-infringin...
Until then all you have is opinions, mine is pretty straightforward: if the generative model can be made to work without first training it on other people's code then it isn't copyright infringement, if not then it is transforming one set of works into another.
The only thing that might let GitHub off the hook is their terms of service, but that might mean mass exodus from GitHub because if they interpret you using GitHub to host your code as a blanket permission to do with that code whatever they want then that's clearly not the original intent of the service.
If Microsoft buying GitHub claims that gave them a blanket license to do as they please with the contributions of millions of FOSS contributors then they are still just as bad as they were in the past.
Almost every GitHub repository comes with a license file, even GitHub should have to abide by that license, otherwise the whole thing is pointless.
If Microsoft/GitHub want to field the argument that they own the rights to all of the code uploaded to GitHub then I'm perfectly fine with that, the only problem I see with that defense is that it will likely kill GitHub overnight.
As for the jury argument: that's fine, but juries aren't lawyers either. I'm not sure if that should weigh as a positive or a negative for Microsoft.
Finally, regardless of the legality: there is such a thing as ethics and in my book you don't appropriate a large body of work from a whole community without so much as a by-your-leave. There have been other threads on HN regarding this and it is interesting to see the various opinions, even so if Copilot is challenged legally than I'll be cheering on the party bringing the suit.
Time shifting: recording TV on VHS to play later.
> GitHub Copilot is trained on billions of lines of public code.
> In one instance, GitHub Copilot suggested starting an empty file with something it had even seen more than a whopping 700,000 different times during training–that was the GNU General Public License.
https://github.blog/2021-06-30-github-copilot-research-recit...
This indicates that they are training it on github's public repositories and at the very least including 700,000 GPL licensed projects or code files. Since the GPL is one of the most "restrictive" open source licenses one can assume they are not caring about the licenses much.
There is no precedent on if a computer reading your code or looking at the image is fair-use or not.
The code being covered by copyright and the code being publicly accessible are two different things.
It was easier, at least - but probably gives some nice guarantees about the statistics of "public" code, with the norms and conventions you're "used to", because you're used to the internet's coding norms.
Company code can be pretty bad, even at Microsoft, often riddled with hyper-verbose variable names and strange design patterns.
I've heard people saying this since the mid-70's.
I lump it into the same trash bin with flying cars and orbiting space hotels, and "90 minutes from New York to Paris — undersea by rail." Things envisioned by artists that will never happen in my lifetime, or yours.