Copilot under fire as dev claims it emits 'large chunks of my copyrighted code'
devclass.com
devclass.com
I remember using Napster for the first time back in the day. Amazing! All this free music at my fingertips, sharing P2P, so cool. And I thought Aaron Swartz was a hero. More recently, SciHub is one of the greatest things to happen to knowledge sharing ever. Software patents are stupid and cause more harm than they're worth, and everyone hates patent trolls, right? Copyright is just a tool to be abused by Disney's lobbyists.
I think the latest tools like Copilot and the stuff Stability AI is releasing fits right in here as amazing tools heralding a new future of exploration and leveling up the power of creators to create. Remixing existing data, maybe copying it, whatever.
Obviously, I'm an adult now, and when working for my corporation I have to be diligent about respecting software licenses and getting lawyer clearance for using tools like Copilot and such, but at some level I've got the Mr Anderson vs Neo thing going on in my head and it feels a little bit like we've lost something from the freewheeling days of early tech.
I tried co-pilot and understand why statistical auto-completion based on the efforts of large groups of open source volunteers is useful but if Microsoft wants to charge a fee for the service then they need to respect the licensing terms of the code they are using because without the data set provided by all those volunteers co-pilot (and statistical auto-completion in general) would be impossible.
You assume that is even possible. I do not believe that would be possible, at all, at least completely impractical. Copilot would not exist. Also, remember Copilot is from the United States, where we have a thing called "fair use."
If a court finds that Copilot is "fair use," copyright and license are completely irrelevant. Cry me a river, it doesn't matter. We already know many things in tech are fair use:
- Google v Oracle showed copying API titles was fair use. Even if you copied tens of thousands of them, for the near entirety of Java. Also see ReactOS, which clones the Windows API; and Linux, which cloned the UNIX API. Imagine how Microsoft and the UNIX vendors felt - but it did not matter.
- Just because John Goodall had `AddTwoNumbers(int x, int y)` implemented as `return x + y;" back in 1981 doesn't mean you can't too. Even if you read his work. There is a degree of copying that is legal, despite copyright or potential infringement, that is decided in courts on a case-by-case basis when someone complains.
- And finally... even if it regurgitates proprietary code once in a while, that does not make a service de-facto illegal. See YouTube. Just because someone gets clearly copyrighted material (such as an entire movie) on there doesn't make YouTube illegal if YouTube takes it down and tries to prevent it from occurring. If GitHub argues the same, that they try to prevent large regurgitations in good faith, and remove complained-about regurgitations on request, the odds of the court shutting Copilot down approach 0%.
The equivalent argument here is homomorphically encrypting an entire codebase and searching through it for matching snippets and not attributing credit to the original code/source. This is essentially what co-pilot is doing but using AI hype as a cover.
1. What do our human brains do? How is that different from how a human brain may function? Our brains take large amounts of information, from all sorts of places, and boil it down just as Copilot tries to do.
2. According to courts, website scraping is completely legal and "fair use." Even if you scrape 200,000 profiles. How is this not different than website scraping, but with code? I expect that argument to also come up. And if it isn't quite the same, consider the legal arguments that allowed website scraping to pass as "fair use" - they will likely apply.
Edit: 3. Currently, ingesting large amounts of material, for the purposes of creating an AI-generated result, is considered legal, even if there aren't many cases. This is also how Stability AI, DALL-E, and others exist. Courts always loathe to shut down fledging industries because it makes them look political. Odds of success overturning that principle outright? I'd give it less than 10%.
In a nutshell, legally, it is not obvious.
Edit @belorn (stupid posting too fast limit):
There is actually a difference. YouTube uses a "rolling key" scheme, which actually qualifies as a DMCA "technical protection measure" under some courts. In which case, pardon a narrow exception granted by the Library of Congress every 3 years (say, for news reporting), your dataset would be illegally obtained - for bypassing DRM. A really, really weak DRM, but courts don't care about strength when determining these things. Which is why it isn't OK. However, if YouTube didn't do that, you actually would, under current law, be OK with training a model against YouTube whether YouTube liked it or not. Code on GitHub does not have any TPMs on it, so it is legal to obtain, and thus legal for training.
2. Scraping by itself is probably mostly legal. What you do with the scraped information is still subject to copyright and other laws.
Edit: Also, the degree to which it must be transformative is, I would argue, lesser for code. There's only so many ways to implement an algorithm - meaning that if it is slightly different, it may be transformative enough, whereas a more creative and less utility-based outlet like painting may require greater changes.
I've posted this elsewhere, but it doesn't matter what our brains do and how similar that is to how AI operates. Humans have rights that machines do not. For example I can watch a movie and not be sued for infringement because I made a copy of the movie in my head.
No AI model works that way either though. Think Stability AI: It doesn't have a copy of every image that it was generated with, but has "distilled" the patterns out of images. It no longer has a copy of any specific image inside it, nor is there an ability to extract training data from it.
In which case, it does not have a copy of the movie in its head either - but it does, for example, recognize the Disney look.
Right now, GitHub Copilot's argument is that training AI models, on copyrighted material, is currently legal. This is also the position almost every AI startup also takes, and it is rooted in "fair use" taking the "transformative" qualities of a work into consideration. There is no doubt that AI-generated suggestions generally are "transformative," it is only about whether they are transformative enough.
My point was that it doesn't matter whether AI works exactly like a human brain, humans have additional rights that an AI does not. This comment made my point better than I could: https://news.ycombinator.com/item?id=33273621
> There is no doubt that AI-generated suggestions generally are "transformative," it is only about whether they are transformative enough.
I don't think even Microsoft/Github actually believes this to be the case because they choose to train on public Github repos and did not include their own proprietary codebase.
There are good reasons they would not include proprietary codebases aside from this, so I don't see this as a smoking gun. Large codebases often involve elements that don't make sense outside of their immediate environment. Projects the size of Linux have this same issue, but large open source projects tend to have significant cross-pollination with the broader community so it's less of an issue.
For example, one large corporate codebase I worked in had a library of ~20, strangely named, short utility functions that were called from many thousands of places. People in the broader community would not find such ubiquitous use of these functions idiomatic at all, but it formed a "dialect" within the company that was often useful given that everyone knew it. These functions caused a lot of signatures/code to be structured in ways that assumed their existence - they were painful to extract when we open sourced internal things. There also tends to be a lot of business logic in corporate code (e.g., `max_space = 100 if has_excel_license else 10`) that would make zero sense in other codebases.
A real example in this case... I'd bet that many millions of lines of Microsoft's internal code still use Hungarian notation. Few users would want Copilot generating such names. I could see interest in a version of Copilot that augments the standard Copilot model with your own codebase if you have millions of lines of code in your GitHub Enterprise account.
It depend on the task. If we are to edit a movie then we cut and paste different parts from large amount of raw material until something new is created. If we do this too work we don't own we call it a remix.
Law is always about the context. Web scraping is OK. Scraping youtube is not. There is no technical explanation to distinguish the two.
I would very much like to see an AI that uses youtube videos as the training set. It could co-sing with people, co-edit movies and music videos, and even generate video game assets from just scraping videos of people playing video games. All very unlikely to ever occur.
If Copilot emanates chunks of copyrighted code verbatim, that makes it harder to argue that it's a truly transformative use. An apposite ruling can be found in Warner Bros. v. RDR Books, 575 F.Supp.2d 513 (SDNY 2008), which found against the defendant due to the reproduction of extensive verbatim passages from the Harry Potter books in a reference book about the series; in this case, one could argue that Copilot fixes in a tangible medium (thus evading Galoob) verbatim copyrighted code without annotation or commentary; indeed, the entirety of the model is nothing more than copyrighted material. (The counterargument is that the text embeddings are fundamentally transformative and thus protected in the same way a third-party index or concordance would be.)
Google v. Oracle, meanwhile, distinguishes between an API signature, which "provides a way ... to access prewritten computer code" from "the code that actually instructs the computer to execute a task," noting that "Google did not copy the task-implementing programs ... from the Sun Java API." In addition, the decision in that case rested in part on the fact that Google only copied a small portion of the Java codebase, whereas in this case Microsoft used the entirety of the repos.
As for YouTube, that's a different matter -- DCMA doesn't apply here, because Copilot isn't a service provider, it's the alleged infringement itself.
To summarize, one argument against Microsoft is that:
1. Training an AI model on copyrighted work does not substantially "transform" the work, but merely fixes it in a different format and generates additional derivative works from the copyrighted material;
2. The code being used is not merely "descriptive" (as in API signatures) but is creative in nature;
3. The entirety of the work is being indiscriminately consumed in the process; and
4. By generating derivative works, Copilot reduces the value of software developers and undermines the market for further software creation.
While there are plenty of arguments in favor of Microsoft, I wouldn't assume that fair use is a slam dunk defense.
FWIW I completely agree with regards to copyrighted code, I think there's definitely an argument to be made against intellectual property rights and laws as they currently are, but I'm not qualified to make those arguments.
> [Microsoft] doesn't attach a GPL license if it draws of GPL code.
It's been a routine practice by developers in industry, for at least a decade, to go read an implementation of a GPL thing, then "rewrite" it in "C#" or "Swift" or "Go", and diligently not tell anyone, and that's it, the GPL code was whitewashed.
GPL has been effective at preventing a specific kind of commercial productization of a specific codebase. Given how influential GPL protected code really is, it failed to prevent Amazon, Apple, Google, and Microsoft from copying GPL code unattributed in all the ways that actually matter.
Copilot is distilling big corporate programming culture. Copying ("rewriting") code found online, let alone in complete implementations, isn't something I do anymore, maybe for years, in my daily software engineering practice. But copying is essential to big corporate programming as trackpants, oily hair and working 1 day a week.
[^1]: Probably, I only know about the French court case a few years back and I recall there being a bit of nuance in what the actual implications were.
That doesn't mean an author that exactly rips off other authors isn't in violation of copyright.
Normal people and small businesses get completely demolished by DMCA takedowns etc which often times are completely abusive, but yet go entirely unreviewed. Because fuck those dumb, money thieving criminals.
But if it's Google or Github doing the thing, it's suddenly an act of brilliance and they're above these plebian rules. Disgusting
Seems like above a certain size, stealing becomes a great business strategy.
Actually it is kind of both. The full phrase was "Information Wants To Be Free. Information also wants to be expensive. ...That tension will not go away", from Stewart Brand's "The Media Lab: Inventing the Future at MIT".
That was a more succinct expression of something he said earlier to Steve Wozniak at a conference:
> On the one hand you have — the point you’re making Woz — is that information sort of wants to be expensive because it is so valuable — the right information in the right place just changes your life. On the other hand, information almost wants to be free because the costs of getting it out is getting lower and lower all of the time. So you have these two things fighting against each other.
It surely must be stopped and punished, or no license will ever be respected again.
MS is taking OTHER PEOPLES CODE and using it ways those people explicitly do not allow and putting it in a walled-garden tool for paid access: license violation as a service.
Copyright today is grossly abused by corps, but it was created for good reason, as plenty of authors and artists went bankrupt because of piracy before copyright law was created (example: Alessandro Manzoni, father of the modern Italian language).
Copyright will be eliminated after money will be eliminated and everyone will get equal access to a decent basic living.
I agree with this, but information seems only to be free for very large players, while it's still very expensive for small/medium-sized players and individuals. It's an absolute inversion of the nominal purpose and justification of copyright if it only protects the strong from the weak.
How curious that copilot doesn't ripoff Microsoft Windows or Office code...
Spare us the astroturfed ideals. You are not one of "us", Microsoft PR person.
You sound like a Russian diplomat defending the indefensible.
I could be wrong but it's double standards unless something like this applies other way around as well. I should be able to sue and bring down a company I don't want using my code under shitty licenses by using say GPL, so that they have to make their own code OSS as well if they use my work.
I want to be able to sue them to the ground, if they can sue me and my friends to the ground on a bullshit whim.
Always get it in writing if you decide to do something like this.
It’s common for academics working in this field to release their software open source instead of licensing it, both as a service to the public and in order to secure support. Davis is an incredibly successful academic whose software is used absolutely ubiquitously. This approach has worked for him: he has had many industrial partnerships and support from the government.
Attribution here means supporting science. If you think that it isn’t a big deal that Microsoft isn’t properly attributing him, what you’re saying is that you think it’s OK for Microsoft to be subsidized by tax payer (and private) dollars, since that’s what’s happening here. The code is clearly identical, with the exception of a few tiny transformations which leave the functionality unchanged. Arguing otherwise is being disingenuous.
It’s almost like Microsoft is capable of not only trawling through huge amounts of code, but accurately identifying Tim Davis’s code. Would be pretty crazy if they used these tools to identify instances of copyright/left infringement and help its users and the open source and scientific communities instead of making it clear that they do not want to do those things.
The internet would explode i think.. lol.
True. And those are license violations as well. But we aren't discussing those.
We are discussing an online service created by a large corporation that automatically hands millions of developers code directly in their IDE's which was entirely stripped from its license. There's no informing of consent or terms of use whatsoever.
> If you want to steal his code you can just copypaste it
When you copy the code from his GitHub repo, a reference to the applicable license is directly mentioned in the files you're copying from.
"stealing" in that regard would mean: willfully ignoring the license and its terms. You can't deny you weren't informed about the existence of the license. It's an integral part of the code.
As opposed to o-pilot which presents you code, giving you the impression that you can just plug it into your own projects, without informing you about any applicable terms of use e.g. attribution.
Put differently, if you're a professional and you copy/paste entire files from some random GitHub repo that doesn't contain any licensing or copyright information, well, you're creating a legal liability for yourself, your employer and anyone using that code through your work.
> No need for a convoluted prompt baiting Copilot into reproducing it.
In an utopian world, everyone would just be able to freely use whatever code is available on the Internet without having to ask for anyone's consent. We don't live in such a world. The reality is one where intellectual property rights and court cases are a thing.
This is really much ado about nothing. Nobody baits copliot into reproducing existing chunks of code, and it's very rare to use a big chunk of code from copilot without changes. I've been using it for months and I can't recall a single instance of doing that. Little bits of code should not be copyrightable in the first place. I can't copyright a single sentence fragment. All this fuss is way overblown.
[1] https://www.trademarkandcopyrightlawblog.com/2013/12/innocen...
There it is. If you have to ask that question, it follows that co-pilot isn't the right tool for developers to use in the first place.
Would you - or your employer - want to risk a lawsuit because you inadvertently copied code that's actually released under a license with particular terms of use?
"Oh, my due diligence went as far as trusting Microsoft and their product for making sure any and all rights where cleared before their tool provided me code." is a meager excuse in a court case when the original author of said code decides to sue you.
> Little bits of code should not be copyrightable in the first place. I can't copyright a single sentence fragment.
True. You can't copyright single fact statements. But clearly, there are cases where co-pilot suggests way more then that.
It's also true that most people don't really consider how stringent copyright laws really are until they are affected by the consequences of their actions (fines, suits, cease & desist,...).
> it's very rare to use a big chunk of code from copilot without changes
That's a tentative claim.
Copilot doesn't do this unless you bait it with very obvious prompts. But to answer your question: yes, never worked at a company that actually cares about licenses. It's a mirage. No-one ever gets sued. If your code is out there it will be used and abused and everyone gets away with it.
... Yet.
Yes. In actuality none of the code I've seen out of copilot is likely to be an issue. This is a mostly theoretical problem that is even less likely to be litigated. If you can't use copliot at work because of this, your lawyers have too much say inside the company, or management is too risk averse.
Also fuck copyright law. It's messed up. I'm not saying get rid of copyrights, but it needs a serious overhaul that is not getting.
But almost: https://en.wikipedia.org/wiki/Ancillary_copyright_for_press_...
Copilot is a large-scale public distribution service which pumps out copyright violations which directly degrade incentives of competent software authors which are known to have brought about the code that Copilot wants to pump out.
It is magnifying the problem of the volume in play of a structural solvent that is dissolving the means to produce quality original software.
> It, or any human doing the same, is not doing anything morally wrong.
A person has placed an object into society, asking for certain treatment of it, upon which they base their livelihood. Knowingly disrespecting and ignoring the requests placed with it – which are in part embedded in a formal social structure of international law and the accompanying expectations of behavior – BY USING A GIGANTIC MACHINE TO DO SO… to support your livelihood… is not a moral failure? Can you confirm?
It’s the same if you don’t use the giant machine for it.
It doesn’t matter if you don’t see it or choose not to accept it as such. The original author may still stop working, and there is a completely unnecessary moral failure there.
>True. And those are license violations as well. But we aren't discussing those.
It feels like we are now talking about it since they choose to bring it up.
What do you mean? Open source is not just source available.
So the funding bodies can support you financially over the code you released, and other users can support you with improvements to make code more valuable and usable.
This is beyond source available, this how Free Software is designed and dreamed to work in practice.
Which is also present in some research software. You buy the source, but can't share/distribute. Only compile, customize and use.
https://github.com/DrTimothyAldenDavis/SuiteSparse/blob/mast...
Maybe that's what the commenter meant but that's not what that language means. Open source licenses are, well, licenses, and if you use them then you're licensing your code. The L in GPL stands for license, for example.
There's no need to be pedantic over a word here.
"Strict/Strong Copyright" makes sense.
Strong and weak copyleft is a different matter, those terms do make sense. Maybe that's the source of confusion here.
People generally ask "why" people give away their life research for free, and I wanted to clarify that as a person who works in an HPC Center and makes research himself.
GitHub does a ton of work to formalize license management, it's odd that copilot ignored all of that.
1) People are disingenuous all the time, intentionally and otherwise. Doesn't necessarily mean that they are "bad". The reality is that people must frequently act cynically for the purposes of self-preservation.
2) With this sentence, I am only referring to the question of whether or not the code from co-pilot posted by Tim Davis "is" the original code. In my view, the transformations present in the code are on the level of what clang-format is capable of. I would argue that if you run code through clang-format, it is functionally "the same" as the original. I believe that this is probably obvious to nearly all experienced programmers, hence my point that I think that if someone is arguing to the contrary, they are being disingenuous. For people with skin in the game (like Graveley, quoted in the article), I would argue that they are being disingenuous in order to cover their asses. For people here on HN, I would argue that they are doing so out of naivety.
Obviously, you're free to disagree with any of this, and anything in my original post. If you feel like I am accusing you of being "bad", that is not my intention.
Perhaps I meant a normal debate and not something horrific and obvious.
People somehow arguing that both things can be untrue are delusional imo. If copilot is ok, then obfuscating code in order to get around a license is ok.
Copilot can produce agpl code verbatim, and per the license copilot should also be agpl, but that’s not going to happen practically, so copilot will probably be shut down.
This is to say nothing of commercial code.
Best way to think of it is like - if copilot was a group of people manually writing and sending code would it be ok? No.
Why should big tech be able to effectively reproduce the work of the little guy without following the licenses they set?
They aren't really confused about the licensing. Microsoft suggests that no matter what the license is for training code, generated code is not infringing — either because Copilot training is "fair use", because the similarity isn't strong enough, or because the amount copied is de minimis. However, the risk is primarily borne by the users, not by Microsoft.
They claim no such thing. They claim that Copilot itself is not infringing any license by being trained on code - that their own use of the code for AI training purposes is fair use and that the Copilot NN is not a derivative work of any code in the training set.
But they explicitly warn you that code that Copilot generates may infringe others' copyright, and that you should use "standard IP scanning tools" to ensure you are not infringing copyright by using the code Copilot spit out.
How can it be fair use to train a dataset whose output is infringing copyright? It's obvious how: it can't be.
If Copilot is infringing just by training on copyrighted data, then so is Stable Diffusion.
Microsoft's position basically distills to laundering:
At the entrance to the black box, there are assets being appropriated in a way that is difficult to criticize. At the exit, the same assets - trivially manipulated - are presented as if they were invented by the box itself.
It doesn't matter what the box is or does. It's clearly deriving its output from its input. An appeal to the complexity of the box as its own moral ground leaves us with nowhere to stand.
In your analogy, does copyright infringement occur when the box is created, or does it only occur when the box is used?
This seems like a very important, but unresolved, legal distinction. If the former, then Microsoft is liable. If the latter, then only users of the box are liable.
Imagine I took a hose, and put a sprinkler on one end and left that in my yard; then attached the other end to the spigot of my neighbor's house.
My neighbor comes to me to complain, "You're stealing my water!"
"Oh, but I've tangled the hose so much that even I can't prove where the water is coming from! Furthermore, I've gotten permission from our other neighbor Dave, and you can plainly see that I'm using his water. Maybe some of yours has been mixed in, but the courts have deemed that 'fair use'. After all, I can be responsible for tracking down where every drop of water comes from!"
"You attached the hose right here to my spigot! Water is flowing out of my house into your hose!"
"As I stated before, no one - not even I - can prove whose water is flowing through my hose. You can't compel me to untangle my hoses!"
---
This situation is obvious. I have left my hose tangled in order to dodge responsibility for taking your water. When someone does this with money, it's called "laundering". The only difference here is that it's an asset (copyrighted source code) and not a currency (dollars) being laundered.
It's further complicated by the fact that you can't "take away" source code. But the entire premise of copyright is to pretend that you can: to violate copyright is to use the assets created by someone else, thereby ignoring their arbitrary monopoly.
Fair use allows us to accidentally or incidentally use copyrighted works, but that clearly isn't what Microsoft has done. They are just using all the code they can get their hands on, and laundering it through ML.
I think this is overly simplistic. People routinely write code that bears a lot of similarity to existing code (there's only so many ways you can write "for x in my_list: print(x)") so there's some threshold for originality that's applied. Looking at it from the other side, minor tweaks to otherwise interesting and novel bits of code shouldn't exempt you from the original licensing obligations.
There's a happy medium in there somewhere (though admittedly it's hard to define).
Looking at it from the other side, minor tweaks to otherwise interesting and novel comments shouldn't exempt you from scrutiny.
There might be a happy medium in there somewhere, but it's hard to know for sure.
Github cannot independently create code, we know this. Comparing Github to humans writing code is disingenuous at best.
This is what the current terms of service of Copilot say: you as the user are responsible for ensuring you are not breaking anyone's copyright when you accept an auto-completion from Copilot. How you would know that it just spit out someone else's code is unspecified, of course.
Of course, similar tools that make it easy to infringe others' copyright are less accepted when they don't come from corporate behemoths (cough Popcorn Time cough), but such is the world we live: copyright protection for me but not for thee.
This is the key point.
If they were confident that the "generated" code did not violate any copyright or licenses then they would be using Copilot internally or offering indemnity for commercial users of Copilot.
Is Microsoft or Github using Copilot for the proprietary software products they sell? (that is, are they "dog-fooding" Copilot?)
What corporate legal team would be willing to sign off on the use of Copilot for a software product that is sold and distributed?
If the "generated" code is licensed under the GPL and it gets mixed into a proprietary software product then the terms of the GPL have been violated and the proprietary code falls under the GPL.
Their enterprise offering is not available yet, so we don't if that is the case or not
Yeah, that's my question. Are MS devs allowed to use Copilot? For, say, developing Windows? Edge? Office? If no, then I'm gonna continue to stay the hell away from it.
In any other industry, MS WOULD offer indemnity protection or it would be dead-in-the-water and no company would even attempt what MS is attempting to do.
Big tech really needs to be reigned in.
> The code, functions, and other output returned to you by GitHub Copilot are called “Suggestions.” GitHub does not claim any rights in Suggestions, and you retain ownership of and responsibility for Your Code, including Suggestions you include in Your Code.
It's not super clear ("retain"?) but I understand it as claiming that you own the copyright on the suggestions you get when you run Copilot.
Source: https://docs.github.com/en/site-policy/github-terms/github-t...
> Other than the filter, what other measures can I take to assess code suggested by GitHub Copilot?
> You should take the same precautions as you would with any code you write that uses material you did not independently originate. These include rigorous testing, IP scanning [IP == intellectual property, in this context - my note], and checking for security vulnerabilities. You should make sure your IDE or editor does not automatically compile or run generated code before you review it.
I believe the "you retain ownership and responsibility for Your Code, including Suggestions" is the part that covers this in the terms of service.
This will just encourage stolen (and now closed source) code.
Why should Microsoft be free to use the work of the little guy and not vice versa? It really needs to go both ways or neither.
So, it's not as black and white as people seem to insist things must be by their limited & flawed understanding of copyright law as it exists in various countries (the world is bigger than the US, and people use co-pilot outside the US).
If you ever copied code from a book or a web page, you are doing exactly the same thing as github co-pilot is doing. And it's fine and legal to do so because that is fair use. That's why fair use exist: so we can actually use the information that is shared with us via copyrighted works. You can't restrict fair use with a license; so the original license is not relevant here. It applies to distributing the whole copyrighted work, not bits and pieces of it.
The only legal question that might exist around Github co-pilot is whether it can be defended as fair use. It probably is but it would need to be challenged in a court to get that verified. So, we might know in a decade or so if somebody actually feels like spending a few million on kick-starting that process (best of luck!). And you'd actually have to go after users of github co-pilot as Github neatly deflects responsibility to the end user. All they do is show you some code.
But that says nothing about its output being fair use, and that definitely won't hold up in court anymore than the output of a human retyping a novel word-for-word would hold up as fair use.
Oh ok so if I come upon a piece of copyrighted microsoft code all I have to do is rename the variables and modify some whitespace and it's no longer owned by them. Lol.
In the end the conclusion from the SC was that Google's use was fair use regardless of the question of copyrightability, which they didn't decide on (the previous court, at Oracle's appeal, had found that APIs are in fact copyrightable and that Google's use did not fall under fair use of that copyrighted API as they weren't 100% compatible with other JVMs).
One thing in life is never to get into a position where you need to defend yourself against this scenario which is why I would never even consider using the tool.
The outcome is that introducing the tool actually tangibly increases business risk in a changing climate.
Copilot, as far as I know, tend to be used to generate snippets and small algorithms that, by themselves, are not a full system/product/game.
There still may be some merit to your argument, but I believe the opposite side should be a bit steelman'ed beforehand.
It's not just one snippet, it's thousands of them, every day being slurped up and copied across code bases everywhere license free.
If the pirate bay can get jail time for facilitating mass copyright infringement, maybe Microsoft needs a little jail time too.
It's been said in this thread already:
Either licenses matter or they don't. Pick one.
I disagree that comparing Copilot with Pirate Bay is a good argument to make your case (maybe Dall-e is a closer comparison if resulting artwork could contain licensed artwork from the training set). You'd probably have other fronts to use and present a stronger argument.
And let me remind you that I have not argumented on behalf of any side, so I do not think this thread requires arguments against arguments I have never brought forward (it feels a bit more emotional than constructive).
This is an argument. A bad one, but an argument none the less.
Copilot seems to be copying code snippets from the internet wholesale. It's not using AI to generate snippets which would possibly be ok.
Copyright doesn't care about the size of the abuse or the size of the works involved.
Microsoft is creating these copies and even facilitates their redistribution on GitHub.
I know legally it's more likely the users of copilot that are at risk here but we said the same thing about torrenting.
Look at Napster, LimeWire or the pirate bay. The outcomes Microsoft deserves here is pretty cut and dry.
so if your work is transformative, it can be exempt as fair use. But i'm not sure what degree counts as transformative - i do think that training an AI model with millions of repos worth of code is transformative.
For example, when Google reimplemented Java for Android, they didn't copy the OpenJDK code. They just wrote their own code that fulfilled what the methods did. If they'd taken the OpenJDK code and just renamed a variable or two in each method, I think Oracle would have had a very strong case.
With Copilot, if it's generating code that is a near-identical copy, that's kinda just copy-paste with extra steps. Let's say that I create a dumb copilot. This will just look at a method signature and fill out the body if it finds an exact match in its database - there's no AI/ML involved. I "train" it on the OpenJDK source - meaning I create a hash of method signature -> method body. I then start writing Java method signatures and it fills in the bodies with the body it "learned" would be good for that method signature.
I think one of the things to always think about is "if the AI were dumber, would it just be copy-pasting?" If the stable diffusion were just taking a copyrighted photo and applying an instagram filter to it, that would be copyright infringement. If I ask for "Robert Downey Jr in a suit" and it gives me back a Getty Images photo run through a filter, that doesn't seem ok to use. If it is truly generating something new and unique from its training dataset, that does seem ok (to me, IANAL).
The problem is that Github Copilot seems to be generating things that are almost identical. In some ways, it seems like they should modify Copilot to catch things that are substantially the same as their training dataset, but that might be difficult.
But I think the process used to create it makes a difference. Copyright doesn't protect against independent creation. If two people independently write the same code, that's ok. However, Copilot isn't independently creating code if it's recreating code that's in its training set.
One of the things I remember reading about that trial was that one of the Oracle lawyers was adamant that a method called "rangeCheck" had been copied and that it was proof that Google was able to get Android out into the market faster than if they hadn't copied code. The method was used inside the standard library (but not exported as part of the API) to check if the index of an array was out of bounds and throw the proper exception (so basically, if index < 0 throw one thing, if index >= length throw a different one, otherwise do nothing). The judge who heard the original case (and the second one at that level after the appeals court sent it back down to judge whether it was fair-use) had programmed in other languages in the past and decided to learn Java before the trial, so he was able to tell that this was such a trivial bit of code that it wasn't at all surprising that the code was the same and that even if it were copied, it would not expedite Android's appearance in the market in any noticeable way, leading to this amazing quote:
"I have done, and still do, a significant amount of programming in other languages. I've written blocks of code like rangeCheck a hundred times before. I could do it, you could do it. The idea that someone would copy that when they could do it themselves just as fast, it was an accident. There's no way you could say that was speeding them along to the marketplace. You're one of the best lawyers in America --how could you even make that kind of argument?"
Obviously most judges won't have this level of technical experience (I remember reading about a different case where a judge upheld an objection to the prosecution zooming into a photo on an iPad due to worrying that it might manipulate the image inaccurately, but then didn't have a problem with them wheeling in a giant TV instead of displaying the imaged scaled up on that), but I think the law has a capacity for nuance about these sort of technical issues that might surprise some programmers, especially with regards to the fact that intent and context can sometimes matter as much or even more than the end of result of something.
https://en.wikipedia.org/wiki/Structure,_sequence_and_organi...
They started by stating "we love open source". Created tools to demonstrate their commitment and then when the community lowered their guard just a little bit, boom, back to old practices.
This is a rant but you get the idea. The level of corporate malice and greed is almost comical and it was perfectly executed.
Start saying you love OSS. Build tools (vscode) to attract devs and keep selling the message "for the community" (WoW reference). Create a Linux environnement to stop bleeding devs and keep selling the message. Buy the largest code repository in the world and keep selling the message. The community starts letting its guard down since it seems they are changing. Then, with no surprises on why this happens, switch gears and start going back to old habits. Rest assured this was not "this is what's best for the community" decision nor it was decided 2 years ago.
We tend to forget that corporations only answer to shareholders. That's it. Everything is and will be, one way or another, aligned around them.
> Developers may wonder: is this AI that generates code, or AI that searches the web for open source code that might be suitable?
Developers who wonder this never tried/researched about copilot.
What it's actually usefull for is to generate code thightly integrated into yours, or doing boilerplate/repetitive stuff based on what you have already written.
Example : I have a base exception class and classes that inherit this class all with different names.
I'm writing error handling logic and only wrote the case for the first error class, with an error message and everything.
Copilot is able to generate the cases for all the other classes, with unique and descriptive error messages for each one.
If you only use it for this kind of stuff, it saves time and you don't have to worry about licenses/code laundering for simple cases like this.
I don't think of it as a tool that solves problems for me, I see it as a glorified autocomplete that understands code and is skilled at applying patterns from previous lines (e.g. in config files) and other files in your project in a way that usually makes sense.
This is a contrived example, but just to illustrate the kind of pattern matching I'm talking about, say I'm writing the following code:
VAR_PROJECT_A_VAR_A=$(curl http://localhost/PROJECT-A/VAR-A)
VAR_PROJECT_A_VAR_B=$(curl http://localhost/PROJECT-A/VAR-B)
VAR_PROJECT_X_VAR_Q5_X=
It will auto-complete it with the following, as expected:
$(curl http://localhost/PROJECT-X/VAR-Q5-X)
It feels like Copilot understands the context in variety of programming languages and configuration formats. If you're writing a Docker Compose file and define a Redis service that contains the flag `--port 8888`, it will automatically offer to autocomplete your exposed ports with 8888:8888. Without the --port flag it will suggest 6379, which is the default port for Redis.
These examples are trivial, but I find that these savings really add up and cut down on the amount of boilerplate and otherwise uninteresting busywork you have to do.
Of course the copyright-violating aspects of Copilot are inexcusable, but I still find Copilot to be useful in other ways.
The issue I'm more bothered with is: the actual cost of copyright/license enforcement. Do we as society really want/need to be accepting and sustaining these costs? The time wasted?
I don't buy the argument that without strong IP laws no good code will be shared publicly: great code is already shared via VERY permissive licenses (and not so permissive ones and yet _stolen_ without others knowledge). Entire enterprises exists adding to the shared corpus of code.
Also, look at other industries where ideas are copied/remixed without attribution all the time: fashion. It never dissuaded new fashion designers or companies to be formed.
If everyone has full access to all code, then it's not possible to plagiarize in secret. It's open for everyone to see and to call out.
Then you don't need software licenses anymore.
What I mean here:
1. Person 1 writes Library 1 with License 1, and posts it to GitHub under Person 1's account.
2. Person 2 steals code from Person 1, or parts of it, without carrying over any of the original license information. This code is then posted to GitHub under Person 2's account, without any links to the original repository.
3. Copilot is trained on data from Person 2's repository and is able to reproduce it despite the original code being under license.
Seriously.
Terms of service don’t cease existing because they’re inconvenient.
This is not really a new problem that has suddenly appeared with these new AIs...
If this is a lightly edited copy, chances are pretty good there’s hundreds of equally lightly edited copies sitting around on GitHub. It shouldn’t be possible to overfit the model to this degree from a single example.
Microsoft should train Copilot on their own code. Theoretically they know where all that code came from and they haven't got a lot of copyright problems internally that would be revealed by Copilot. If they won't do that, that should factor into what we think of their arguments about what Copilot means for training corpus authors.
Normally, if person 2 steals your code and hosts it on Github while infringing the original license, person 1 can send a DMCA notice to Github to take the code down. A similar process could be involved for copilot too.
If copilot would be a DMCA safe harbor, then they would need to comply would DMCA notices and stop distributing the offending code in a timely manner. It might be implemented by elaborate filters (which are probably not a 100% accurate), or batching up DMCA notices, and retraining their model regularly with offending repos/code excluded.
Of course, Microsoft drag their feet and say that copilot never infringes on copyright. They don't want to do any of this.
https://en.wikipedia.org/wiki/Abstraction-Filtration-Compari...
This allows for verbatim copies if they are utilitarian in nature!
As for why we should allow verbatim copies of utilitarian features... First, let's preface this with the substantial similarity of the structure, sequence and organization as established in Whelan v. Jaslow which amongst other things says that you cannot merely change the variable names if the expressive structure of the code remains the same. Now let's imagine 10,000 software developers who all implement Dijkstra's algorithm in C and then run it through clang-format. Aside from variable names, isn't it safe to assume that many of the implementations are going to be exactly the same?
Now, this doesn’t mean that GitHub is not in violation of other copyright claims, such as clearly expressive parts like comments and more!
I'm using "pattern" quite loosely here as recent ruckuses come from the StableDiffusion image synthesis and CoPilot - n=2.
I have no idea what the ethical and legal long term ramifications are.
It seems, with copilot it's mixed. Sometimes it regurgitates blocks of existing code, other times it can apply context and patterns to create boilerplate code.
Maybe we are a little too short sighted and defensive while facing our industrial revolution in the form of the AI revolution, which is coming for our jobs, or at least for the way we are used to do our jobs.
Those pesky copyright violations could be manually 'fixed', or, Microsoft could go and prove with their data that programmers are all just reinventing the wheel and are actually reimplementing the most basic stuff over and over. Meanwhile, they might be correlating live data from Azure and git histories from many projects, and have Copilot refactor your project based on load. Unthinkable?
You can't convince me that typing
get item() {
and it autocompleting it to get item() {
return (itemId: string) => {
return this.items.find((item) => {
return item.id === itemId;
}) || null;
};
}
is using copyrighted code, especially since I've written scores of nearly identical functions on my vuex (vuex-module-decorators) stores long before I touched CoPilot. Maybe it's "stealing" from my own code? Good! That's exactly what I want, follow my previous patterns. And honestly if it auto-suggests that to someone else then I don't care.Sure, if I write a `doReallyComplicatedMathmaticFormula()` and it autocompletes super complicated code then that's one thing but I've literally not seen that once (or maybe I just don't write any code that would cause it to inject larger chunks?). Also if it's something really "basic" like the Haversine formula (determine distance between 2 sets of coordinates) then I'm fine with it injecting an implementation of that, is that really copyrightable? There are only so many different variable names/ordering you can do and still be using the Haversine formula.
Almost all the code it's suggested has been in my style (naming, spacing, etc) and referenced variable/property names that are either super generic OR super specific to my codebase (I seriously doubt anyone else has some of these variable names we came up with, if you do then my apologies).
That's exactly that's happened and discussed in the submitted article. For reference the code reproduced (literally) is: https://github.com/DrTimothyAldenDavis/SuiteSparse/blob/dev2....
We are all getting trained on someone's else code and up to a certain point replicating it (along withe some mix and matching).
Let's make an open source version, this should be for everybody.
How many times have we written code to solve a problem, got a new job at another company and come across the same problem and ended up writing very similar code. That's not even copying code written by another person, but still potentially copyright violation (depending on how similar that new version of the code is).
I bet the same people would be fine if this code was included verbatim with appropriate attribution and licenses. But that’s not the case.
His opinion, as well as the current state of law: His code could and shall be used pretty freely, but a comment referring to the original license/author is needed.
The context: Having an option to "leave out copyrighted code", that does not do what the option suggests, is a flaw, both in a functional and secondly with law-as-is, in a legal way.
The problem is that code is text-based and obviously some kind of writing while mostly being a utilitarian invention. The courts of course recognize this distinction (because unlike the average software developer, lawyers and judges have studied the law) so they have come up with a number of ways to filter out what is and isn't covered by copyright with regards to software.
The deal with patents is that you make the way your invention works public knowledge. You document the principle and in exchange you get a limited-time monopoly on its implementation. An alternative for this is a trade secret. As long as you guard the secret, you can remain the only one profiting off it. Of course, some things are impractical to keep secret. It's easier to keep a method of making a fizzy drink under wraps than how the gears are laid out in the gadget you sell.
You also can't really make the contents of a book a secret. Anyone can just look at their copy, arrange words in the same order as you did and sell your book. It also makes no sense to patent the words. The contents are already inherently public knowledge to a certain extent and now people could just buy the book from a patent office. And the clerks don't like filing novel-length applications. Music is not quite as easy to make a 1-to-1 copy of, especially before wide adoption of sound recording technology, but a lot of people can carry a tune well enough to plagiarize one. And sometimes you get a child prodigy Mozart illegally transcribing Miserere.
Since it has been deemed desirable [if not unanimously] that writers and artists also have a period of monopoly rights over their work, copyright needs to work differently from patents.
Software has interesting properties. Unlike a book, you can distribute a program but keep its "recipe" a secret. We'll disregard people who are capable of working with compiled binaries; consider them modern-day transcribers of Miserere. Proprietary code is typically guarded just as any trade secret would be. Perhaps software should have been covered by patents instead of copyright from the beginning. After the patent expires, the algorithm would be unencumbered and public knowledge.
But that's not the world we live in. Software is considered to be like a book, not like a pocket watch. You could mail me a printout of the entire Windows 11 source and there's hardly anything I could ever legally do with it. I could change it for my own amusement I guess, like I can take a red pen and change all the names in my copy of Postmodernism for beginners.
I’m totally against software patents and totally in favor of software copyright, fwiw.
Open Source software can exist in a world without copyright; that's basically what permissive licenses are. Proprietary software can exist in a world without copyright; just don't release the source code, only distribute build artifacts. But free software is dead without copyright.
Most permissive licenses — and most popular permissive licenses — include a copyright statement and conditions which would not be enforceable without copyright. They usually require redistribution of the original copyright notice and license text to be distributed along with copies or derivatives of the work.
That doesn't mean there can't be other legal mechanisms that would also accomplish software freedom.
Closed-source software should just be illegal. Then you don't need copyright to ensure freedom. With copyright gone, you are now guaranteed free to reuse any code. The law should not prohibit copying, it should only require attribution (i.e., prohibit plagiarism).
For a prominent example, WINE goes to extreme ends to ensure their code isn't derived from any prior direct knowledge of Microsoft code. Replicating functions of released binaries with original code is fair game, but reverse engineering or copying Microsoft code is a hard no-go.
As far as AI is concerned in all this, it's probably even easier to distinguish the unacceptable. Whereas humans can argue for plausible deniability, we clearly know what data an AI is fed to generate subsequent output. So unless an AI is certified to never have eaten licensed materials, literally everything it produces will be license infringements.
The training data doesn’t exist at runtime, and if the developers are competent then it only learned a bit or two from each input snippet. So, though even though it’s certainly a highly capable compression algorithm, verbatim copies shouldn’t be… possible.
My best guess for this bug is that a lot of already plagiarised the code. AI has a habit of holding up a mirror to humanity. Sometimes we don’t like what we see.
I don't think current legal frameworks are equipped to handle these types of generative works. There will need to be either new laws, or at least some legal precedence established. For example, it would make total sense for stuff posted publicly to optionally have a "not to be used to train models" license. But models are being trained on public data that pre-dates anyone thinking that is something that they should worry about.
Addressing this exact premise is the purpose of copyleft licenses like the GPL. I would expect GitHub of all companies to understand this, and it's a spit in the face to disregard it, a spit in the face of the entire FOSS community to which they owe their success.
Disclosure: founder of a GitHub competitor
Most software licenses are not full carte-blanche do-whatever copyright waivers. Sure, it would be hypocritical to get up in arms about copilot emitting CC0/WTFPL/Unlicense code, but comparatively little software is released under such licenses.
The GPL (in all its versions) is almost as notable for the specific conditions placed on the freedoms it offers as it is for those freedoms themselves. One possible reason to release your code under GPL is that it you don't want your work used in proprietary developer tools. It doesn't really matter then if the code is used in a proprietary developer tool's training data set instead of its own program code.
Even permissive licenses typically come with some conditions attached. If I release a program under MIT license and someone then uses portions of that code in their own project, I expect to see my name and the MIT license included in some way. Perhaps a line in the README, an ACKNOWLEDGEMENTS.txt or a comment like
// This function taken from quuxifier (https://example.org/software/quux)
// ⓒ bitofhope 2022, used under the MIT license. See doc/licenses/MIT or
// https://spdx.org/licenses/MIT.html for details.
The same, in my opinion, applies if recognizable (for some definition and threshold of recognizable, which can and does get deep into lawyer territory) portions of that code are emitted by a convolution network. If Copilot can write my function, why can't it write my name and choice of license too?Even if you are a full copyright abolitionist, you can still point out the hypocrisy from the opposite side. Github's parent company Microsoft is notoriously protective of its proprietary code and its copyright. The company has historically been explicitly hostile to the free software community and despite its later unilateral declaration of love for open source continues to profit from their proprietary code, including Copilot. If you spend a decade or a few calling open source a cancer and campaigning against it, you can expect that to come bite you in the ass when you later try to sell a neural network trained on a vast corpus of free and open source software.
Or, well, no, "weird" is not the correct name.
What will benefit the world more?
I'd say allowing to train
But on the other hand I'd probably be pissed off if somebody took my work without credit
Hard to say, maybe link to original code would be fine?
Do not annoy copilot users with credit spam and actually give credits
> I'd say allowing to train
If we widen the conversation a little, I would argue that it would benefit the world to publicly release the source code of Windows 11. Not only would Jr. programmers get a huge real world code base to learn from, but I'm sure many exploits could be found and fixed in short order. So why are we limiting the discussion to violating just FOSS copyrights?
I guess when a computer does it people get creeped out. Or the computer is not smart enough to change the variables names. :)
As a result it sure seems like Copilot or similar applications of DL/ML is a red herring a general criticism. If someone is upset their elegant solution to a specific problem is being replicated, that's not about DL/ML; and there are unlimited opportunities for reverse engineering or reimplementing or transforming or rewriting or whatever you like, which affect the same algorithm etc.
If you want a software patent, take up the fight in that domain. If you don't want people reusing your logic, don't put it in public view or for that matter anywhere an interested and motivated party might disassemble or deobfuscate it.
It’s pretty hard to prove that the code was copied from an individual, so good luck to them.
I think one of the things the FSF kind of got right is that copyrights (and patents) on code are kind of unnatural and difficult to enforce.
If it only finds one, it can list it, but highlight it with the associated licenses.
Copilot will copy that college student's code, even if it looks similar to mine.
If a valuable tool can disappear due to copyright issues I'm not using that tool.
Isn't this still a copyright violation? I'm not a lawyer, but I always thought that swapping in a few synonyms wasn't enough to sidestep copyright issues.
That's a pretty big hash-list you've got there...
Edit: which ToS paragraph are you referring to? I haven't been able to find any grant of rights beyond the ones needed to use GitHub.
> This license includes the right to do things like copy it to our database and make backups; show it to you and other users; parse it into a search index or otherwise analyze it on our servers;
I think using it to train an AI model falls into the "analyze it" part.
And bout the mirroring, I'd say what happens is the same as whatever happens if someone sold you something stolen. I don't know exactly but I think the same kind of law applies.
As a thought experiment I wonder what the reaction would be if one day it was revealed that copilot is actually just minimum wage coders in Lagos who strip open source code of their attribution, rename a bunch of variables and reformat the code and then sell those snippets as a service at the behest of a gigantic company.
“It looks like your copyrighted code because it learnt from reading your code”
I could be totally wrong, but I have a hard time believing GitHub didn’t pay attention to licenses when training CP. I would put my money on some random person not respecting a license (like me, many times) and putting it up without thinking about the downstream effects.
IANAL, but if someone naively uploads my copyrighted code and just checks the MIT License box without knowing better, I’m not about to sue them over it. I’m not sure “but this 13 year old kid learning to code for the first time uploaded it under a permissive license and it got gobbled up by our crawler so it’s not our fault your honour” would be an effective defence in court, either.
That all said, I’ve largely just described any modern search engine which has been working in this way for years. So.. shrug?
Will no doubt be litigated at some point down the road.
Instead, the goal of the Free Software movement would be to enshrine the 4 freedoms into law - for example, making it illegal to distribute software without making the source available, and to make it illegal to distribute hardware devices whose software can't be changed by the end user.
Having copyright + the GPL is more useful than no copyright at all.
That's an inaccurate understanding, and is not supported by that history page, or by more recent actions from the FSF and Richard Stallman. Without copyright, there would be no copyleft - which would mean that companies would be entirely free to take any libre software, package it with their own user-hostile changes, and bundle it in an obfuscated binary. They would also be free to do what TiVo did, and add hardware constraints to prevent you from running modified versions of "their" software, even if "their" software is GNU utilities.
The framework of copyright is actually quite important for the free software movement to be able to enforce its goals and try to force companies to contribute to free software.
Copyright doesn't, for sure.
You play privately, but you have absolutely no right to copy and redistribute the work to the masses, or even do whatever you please with it.
Not so with a piece of code that you use privately.
I wouldn't be surprised if lawyers/courts/lawmakers determine that current generative AI are on the "infringing" rather than the "inventive" side of things.
I think there is potential for this to change as AI gets better:
https://kitsunesoftware.wordpress.com/2022/10/09/an-end-to-c...
Does it really matter? Once you embed copyrighted works into your learning dataset, any model you ship and any "creation" of that model is by definition a derived work, must have a license from the original authors and must obey the terms of that license. The extent of derivation is immaterial.
The AI crowd is on a hard collision course with an army of gavels.
The theory used here though is that fair use permits transformative use. Normally copilot output should be sufficiently transformative. This particular situation appears to be a sort of bug (overfit).
Note that AI models already work somewhat similarly to how humans work (becoming bug-for-bug compatible at times even :-P ). We may need laws to be amended in the opposite direction even, else it might become illegal for humans to learn too.
Of course everything a human learns and produces is a derivative of real world input, some of it copyrighted. But people agreed that this transformation is a form of fair use, as long as the derived work is sufficiently creative - where the key implicit assumption is that creativity is a human ability that warrants protection. Both the creator and the learner enjoy protection because they are human.
In the AI case, the output of a program will certainly not receive protection by default - it will be protected only as a consequence of its human creator or user rights. It doesn't really matter how transformative and analogous to the human mind the program is, the output only deserves protection solely for its unique human created parts - the internal working and the extent of the craft embedded in a GPT3 text prompt. The first is unrelated directly to the work produced, and the second is laughable vs the artworks embedded in the model.
If in the future AI advances to the point where it can be granted natural rights, this could change, but we are far far from that point.
Animals have (some, limited) natural rights, but can't hold copyright. Corporations don't have "natural" rights but can hold copyright.
If you have code under a restrictive license but we acquire it and use it privately for our own needs in trade secret processes how would you even know?
And if that isn't good enough, I'm sure the copyright holders can lobby for whatever to make it easier to find, and they'll only be thwarted in this if the governments are convinced that the interests of the lobbyists are no longer aligned with the government's interests.
In cases like this, the benefits of the technology are so huge, and the downsides to the original code author so tiny. Was he actually going to license that code to someone who said 'nah, I'll recreate it with copilot instead'?
Very few people successfully license anything less than 100,000 lines of code anyway.
When this starts impacting emplyment opportunities for devs, get back to us with the "tiny" downsides.
Benefits to whom? I don't recall Microsoft making this technology free for OSS development.
Suppose you're right and 5 years from now everyone is using co-pilot. To remain competitive with proprietary software, OSS developers will have to use it also but they will still be paying Microsoft $10/month for the privilege. Meanwhile none of them were compensated for making this possible in the first place.
Actually, they did exactly that... Anyone with a decent number of commits to any opensource project gets a copilot license for free.