It's been a massive productivity improvement to our senior devs, and I got so used to it that it's an annoyance when Copilot doesn't respond.
It's been a massive productivity improvement to our senior devs, and I got so used to it that it's an annoyance when Copilot doesn't respond.
It would be because it's illegal and violates the licenses, desires, and intentions of the thousands of workers who wrote the code in its corpus.
You act like Microsoft is trying to do a public service and people are angry about it. The reality is that they're taking billions of hours of work and using it to build a product that only they control.
If they re-released Copilot as FOSS, a lot of the valid criticisms would evaporate.
I wonder how many people on HN would be on the side of the creators if we were talking about content created by Walt Disney and whether pirating was ethical?
Disney used its power to distort copyright laws in a self-serving manner. As an individual you don't have equal power to oppose them (you were supposed to have in a democracy, but lobbying is legal and corporations are people).
Disney is a huge corporation that won't even notice if you pirate a movie, which you may not even have been able to pay for anyway, because of their region-locked twisted maze of distribution and DRM.
OTOH you may be screwed if you're a creator making a living from your work, and a big corp can just take it without paying, launder it through "… in the style of $YOURNAME" query, and say they own it now, because unlike your copyright, their Terms Of Service apply.
Even people who think copyright shouldn't exist may rely on using copyright — against itself. You can't unilaterally say "I don't believe in copyright", because the law doesn't care, but if you license something as copyleft, then the law does care about your anti-copyright license.
People probably would have less of a problem if microsoft breached the license on 100 year old code.
But life of the creator + 70 years?
Reminds me of this:
https://arstechnica.com/uncategorized/2007/07/research-optim...
Without copyright, it would be perfectly legitimate and legal for someone to not follow a license (because it would bear no legal weight because of the lack of copyright).
I would argue that open source is best served by strong copyright protections that allow the people created the software to make sure that further changes to it are released back to the community. Weakening copyright law means that it is that much easier for big companies to co-opt some software and not need to release the changes back.
As long as Microsoft can and will wield those laws against me? Darn tootin'
This is also why I fully expect Steamboat Willie to fall out of copyright protection in January 2024 - right on schedule. There's a few countries that have supra-EU copyright terms, but none of them are dealmakers. Nobody is demanding we match Mexico's life+100 terms, for example.
I am 100% on the side of content creators. Regardless of who they are .
The courts tend to take a dim view of theft. Which is what this is.
The article clearly lays out that multiple requests for sound legal basis have gone unanswered . It simply doesn’t exist and Microsoft is operating on a forgiveness vs permission model.
Licensing is 100% about permissions. Clear and explicit enumeration of the permissions (or lack thereof ) for a work.
This class action lawsuit should surprise nobody. It’s a class that is sick and tired of being exploited.
Do not take my work that I contributed with explicit permissions and use it in a way I didn’t grant permission for. Full stop. It isn’t complicated.
You wouldn’t download a car and all that jazz….
https://docs.github.com/en/site-policy/github-terms/github-t...
so much animosity over rights you gave GitHub when you put your code there. "Theft" gimme a break. you license your code to GitHub so they can show it to others. This is separate to the stated license in your code. Nowhere in that terms of service document are the means that the code is shown to users specified.
Also MS themselves don't even claim that training is covered by their terms of service, they claim it is fair use.
Or you are mistaken and "the service" of github, includes all features available on the website including copilot.
Even if you're right and a court rules against them, what's to stop them changing the terms to become compliant?
Moreover terms have been largely unchanged for years AFAIK. If someone agreed to the license years ago, they can't have agreed to copilot use. Also copilot is not a service on their website, it is a separate service and they charge for it, also contradicting the terms.
What does separate service mean? What would Copilot look like if it was not separate?
Elaborate on "not a service on their website", as it is available and listed as a feature "Github Copilot" on their website.
Is the contradiction related to payment for the service, or just because you think it is separate?
Since I thought you were arguing that Githubs own terms prevent them from using the public repositories in Copilot, this is what I argued against.
If you think fair use is involved, then that's the end of the line. If MS claims fair use, then until a court says otherwise, it is. Anyone who thinks their copyright is being violated can get an injunction tomorrow.
Maybe some type of hint shown inside your code when it’s shown at github.com. There is already a text editor.
they claim the training of their model is covered by fair use, but they did not say that was the justification they were using. They don't need to claim fair use.
It's pretty clear from the Terms of Use that they can use code hosted on github.com to provide any service they like, so long as it is a GitHub service. They don't need fair use, they already have the rights to do what they are doing.
Fair use is a major part of copyright law. I do not have to ask permission to use your work.
For you to win in court you have to overcome fair use, you have to overcome innocent infringer, you have to overcome no damages.
Anyone leaving comments saying that there's an obvious way a court would rule on a copyright case involving those 3 things is wrong.
If you use my licensed work then yes, you do need to follow the terms of the license .
The issue of license / contract / copyright is messy. It doesn’t ever seem (in the USA anyway ) to be definitively answered / “solved.”
I chose AGPL v3 only on purpose.
Co pilot and users thereof (so now two levels removed ) utilizing code in whatever work are stealing my work (unless it’s AGPL v3 licensed ). The adding of intermediaries (and the most likely unknown and with no way to know) infringing is going to be very difficult to mitigate. It’s like truly unknowingly buying stolen property,
If I use it under fair use, there is nothing you can do about it.
The fourth factor of fair use is the effect on the value of your work. If I'm not affecting the works value because there is no market for it, because it has no commercial value, you are going to have a very hard time defeating this argument in court.
That changes nothing at all. A FOSS license is, in a legal sense, no different then a proprietary one. If it’s illegal it’s illegal regardless of whether it’s FOSS.
If Copilot was FOSS there'd probably be a few absolutists complaining, but they'd be mostly ignored.
Why haven't they uploaded Windows source to Copilot?
Just how much code reproduced violates copyright?
If, instead of Copilot, Bob was giving me code to copy, and it was a AGPL codebase, am I still subject to the AGPL?
The problem is that their product sometimes produces verbatim copies of licensed works, without attaching licensing information. This not only goes against the licenses under which the original authors made these works available. It can also put the product's users in danger of anything from bad publicity to a copyright lawsuit.
CoPilot is a very interesting research project. It's not yet an acceptably mature product though.
Copilot makes source code much more open, if you think about it. It implements code reuse in a different way than classes and libraries. It offers its skills equally to everyone, skills learned from everyone.
As for the cost of the API - it's expensive to run large language models, I think the price is justified. But there are free models if you like to run your own.
For a commercial version, run it on Microsoft's internal code, the code they actually own!
Yeah, the vocal few.
Do you think I give a rats ass that Copilot is duplicating my OS code?
I have to imagine most people are completely ambivalent. Of course I have no proof, I just can’t imagine anything else.
The lines probably fall somewhere along the MIT vs GPL camps.
"Ambivalent" means "of two minds," but I'm going to assume you meant that you're indifferent.
If people are/were indifferent, their licenses should reflect that. They overwhelmingly don't.
Regardless, Microsoft is legally bound to obey the licenses.
I would accept a claim of license violation if someone used copilot to autocomplete so many methods from one specific project that you have recreated that original project.
I still think it is a matter of scope. It can still be the case that a relatively small module is not cool to lift, but I think in this case we are still talking about such small subsets of functionality that it is completely divorced from the original software. Like, I could see it if a specific method were really key in some way to a unique application, a very novel solution to a difficult problem - but if that were the case, how can an AI possibly use that for a training model? In other words, the auto-suggestions of an AI are going to be common coding solutions to common coding problems that the AI has seen hundreds of thousands of times. That individual proprietary GPL, unique and novel solution is not really the stuff of an AI suggestion. In other words, the code that co-pilot is going to suggest is going to be non-unique, generic, and not really specific to the overall application at all.
> If people are/were indifferent, their licenses should reflect that. They overwhelmingly don't.
Apparently, overwhelmingly they do. At least if the licenses used are any indication.
https://github.blog/2015-03-09-open-source-license-usage-on-...
I'm almost entirely certain you're wrong about the desires bit. 99% of the developers who wrote that code won't mind.
This is completely unsubstantiated. I for one would mind microsoft profiting from the closed source code I wrote.
If I’m writing code for a query optimizer, the SQL Server solution isn’t going to magically show up.
It doesn’t mean you have the right to use any of the code it generates but Copilot itself isn’t illegal in any meaningful sense.
I don’t think it’s accidental that this product is specifically Github Copilot.
But even then I think this is legal overkill. If you use the search box on Github they will display snippets of code from public repositories without the license. Same as what Sourcegraph does same as Copilot does. Nobody here is arguing ripgrep is violating the license by displaying matches without the corresponding license.
Yes if you use a tool to violate copyright it’s copyright infringement. If you prod Midjourney into outputting near exact Starry Night that’s on you too.
So far no one has made a compelling case that Copilot itself is violating copyright.
The people who comment on something are disproportionately those who care a great deal.
What proportion of its capability is derived from the labor of people who don't like it? I get your point about feeling like an unwilling contributor while github/MS harvests revenue from people who like it. But there's an implication here of being in a critical majority, which I am not convinced is the case.
Reread your post. Doesn't it sound scary? You are blocked from even thinking and crafting because a specific web service is down.
Even if Google is down you can go direct to Stackoverflow and MDN, and have a choice of information sources.
Also what is "productivity" ... as in features built / month or lines of code / month?
People's expectations have already been set by this technology, and they are only going to want more. Also, AI researchers are still publishing their work out in the open for anyone to reproduce.
If there was a Copilot model out in the wild like with Stable Diffusion then this ceases to be a valid question, regardless of the model's legality. All it takes is a single leak or decision by another entity to release their own code generation model.
Saves a lot of "hey how do I do this simple thing again?" memory loss issues.
I would hate to work at a place where advanced-but-untrustworthy autocomplete would, at all, impact the productivity of a senior engineer.
Not only does this indicate that your senior engineers' productivity is measured poorly (lines of code), but also that your senior engineers are paid to type, rather than to think.
Life's a lot easier when you can just copy whoever did the hard work without crediting/paying/etc for it.
Incorrect suggestions all the time