GitHub confirmed using all public code for training copilot regardless license
cybre.space
cybre.space
Automate the rote software engineering so it's built upon a shared, healthy, secure, etc codebase... Leaving only the harder, more-fun 1% for us to noodle on instead.
I might start with a simple for loop:
for(int i = 0; i < limit; i++){
...
}
And then decide that I can make it more concise: for(int i: [0 .. limit]){
...
}
But I still need the base form, for less likely scenarios: for(int i = 1; i < limit; i <<= 1){
...
}
A language or its libraries can support new cleaner forms for fundamental structures and frequently repeated boilerplate constructs - but it increases library/compiler complexity, cannot always be predicted ahead of time and often sacrifices power for specificity and apparent language simplicity.At some point in the future we won't need auto complete because the only thing we'll be coding is the model itself, and then we'll provide rules (in some format, not a language) and the UI will know what to do. We see a lot of that happening in the no-code movement today.
As long as you're smart enough that you are still required to at least delegate some tasks to a machine, you're still worth keeping around - even if you're not always busy.
I think "commoditizing" programming would be great for humanity.
There are lots algorithms that can just be copy and pasted. Of course many of those get turned into a library. You still have to know how to string those together.
I suppose if we assume it retains it, then TFA's claim creates a further question of how the hell anyone using copilot can conform to this unholy mixture of all the licences. But that's a big assumption at this point.
A legal case could be made that they are assisting in ip thief by selling a product that makes copyright breaking easier (we have seen those cases) the majority of legal issues will be on the user. Imagine the Google / Oracle legal battle replayed with this tool involved..
Yes, sometimes code is returned that is a verbatim reproduction of the training data. This can be prevented if need be.
What I really don't understand is how some people are complaining about GPL'ed code being used for training.
What's the difference between a machine looking at the code and learning from it and a human being doing the same. As long as the code isn't patented, there's no reason why I shouldn't be able to look at GPL'ed code and implement the idea using my own code.
In other words, is - according to those who think using GPL'ed code for ML training - every implementation a derived work if I looked at GPL'ed code that implemented the same algorithm? Where's the line that separates plagiarism from original work? Is there even such a line? Does it matter whether the GPL'ed code is encoded in human neurons or network weights after looking at it and if so, why?
It could be like a dev reading public code enough to understand it, but then coming up with her own implementation. Not saying copilot works that way - I haven’t tested it yet. But could be one nuance here.
What about images?
You can claim that it was proposed by the tool and that you didn't know the license terms, but lack of knowledge of the actual license still doesn't mean that the license is not relevant.
I would rather not use such a tool than worry about infringing someone's license.
And you'll be screwed.
Probably would have been smarter to just use bsd. Perhaps they tried that and failed Is it too late to rerun the model? It feels like it.
One set of terms doesn't invalidate another, at least in EU.
Secondly, if I release my code under a license that doesn't permit this sort of thing, and someone else publishes the code on GitHub, I haven't agreed to GitHub's terms and the person who did can't give GitHub any more rights over it than I gave them.
Not every piece of code on GitHub was written by people who agreed to the TOS, it's quite easy for open source code to move around. Suppose the maintainer of a project moves it from GitLab or whatever, with your contributions intact. The maintainer only has your code because you contributed it under the AGPL or whatever- they can't "license" it to GitHub under more lenient terms than that!
> You grant us and our legal successors the right to store, archive, parse, and display Your Content, and make incidental copies, as necessary to provide the Service, including improving the Service over time. This license includes the right to do things like copy it to our database and make backups; show it to you and other users; parse it into a search index or otherwise analyze it on our servers; share it with other users; and perform it, in case Your Content is something like music or video.
> This license does not grant GitHub the right to sell Your Content. It also does not grant GitHub the right to otherwise distribute or use Your Content outside of our provision of the Service
The core of it (IMO, IANAL) comes down to "as necessary to provide the Service". It claims a right to "otherwise analyze it on our servers", but the clause "as necessary to provide the Service, including improving the Service over time" likely still applies to that.
So it comes down to whether Copilot still counts as part of their "Service" that the user agreed to when using the site, or if it's different enough that it's unreasonable to apply that proviso.
[1] https://docs.github.com/en/github/site-policy/github-terms-o... (thanks themanmaran)