All public GitHub code was used in training Copilot
twitter.com
twitter.com
Take Google as an example, running Google Photos for free for several years. And now that this has sucked in a trillion photos, the AI job is done, and they likely have the best image recognition AI in existence.
Which is of course still peanuts compared to training a super AI on the entire web.
My point here is that only companies the size of Google and Microsoft have the resources to do this type of planetary scale AI. They can afford the super expensive AI engineers, have the computing power and own the data or will forcefully get access to it. We will even freely give it to them.
Any "lesser" AI produced from smaller companies trying to compete are obsolete, and the better one accelerates away. There is no second-best in AI, only winners.
If we predict that ultimately AI will change virtually every aspect of society, these companies will become omnipresent, "everything companies". God companies.
As per usual, it will be packaged as an extra convenience for you. And you will embrace it and actively help realize this scenario.
They use photos from the web for training, and then user photos are only used for the actual indexing.
Yes sophisticated AI tech concentrates power for those who already have power.
And the technology we all (presumably readers of HN) create can enhance the impact of the user. And this can result in unfair circumstances, in reality.
Law and force can prevent disproportionate use of power. Of course one must define the law, which may be done AFTER the offense has been committed. Further, if those who make the laws are corrupted by those with e.g. this AI tech power, then no effective law may be enacted and the hypothetical abuse will continue.
If there were offline image recognition we could train on our own data privately, could the results of those trainings be merged to come up with better recognition on average than any one person could do themselves with their own photos?
In other words, would it be possible for us to share the results of training, and build better models, without sharing the photos themselves?
Transfer learning best use cases are for fast prototypes or for ml tasks that do not need state of the art performance.
By keeping the training data itself private, distributed and outsourced, you might be able to get otherwise unachievable levels of performance.
What I'm building into PhotoStructure is typically called "transfer learning."
https://en.wikipedia.org/wiki/Transfer_learning
PhotoStructure is entirely self-hosted, including model training and application: the public domain base models (trained on huge datasets) are fetched and cached locally.
By design, none of your data (or even metadata) leaves your server.
(I expect to ship this in an upcoming beta next month.)
I run windows. It can't ever be secure, anyone who wanted to hack me could.
Scrambling the data really makes things worse as any accident requiring recovery of my data is also probably going to lose the encryption key.
The only time I ever lost any significant chunk of data (a persons lifetime set of photos!) was because Windows encrypted data at rest, and thus it couldn't be recovered after a disk crash.
Unless there is some corporate or legal requirement to do so, I'll never encrypt a whole disk, or backup.
I wish backup tools like Duplicity would warn you about the risks of encrypting backups instead of warning the user if they disable encryption, because encryption has the possibility of rendering all those backups useless when the moment to use them finally comes.
I have a similar feeling that large swathes of my digital life would be rendered permanently inaccessible if 2FA was enabled and my device was rendered inoperable. (That's why I keep meticulous physical backups of emergency keys.) I think 2FA and the like should be considered a tradeoff with its own inherit risks and benefits, instead of a universally better option than randomly generated 80-character passwords alone.
... why?
i'd hate encrypting too if I threw away all best-practices regarding it -- losing a key with the failed system is a "problem exists between chair and keyboard" type of issue.
Encryption protects your data from yourself, from your adversaries, from serendipitous grey-moral types, and from the prying eyes of over-zealous data-collection conglomerates.
You seem experienced in the field, so I won't presume what your best practices are -- but to be enthusiastic against encryption is a form of cheer-leading that I think I cannot ethically support; the longer I live and the more pervasive companies get to be with their data collection policies then the more powerful and required tools like encryption seem to become.
Apple does all face recognition and image processing stuff on the edge. On your iPhone or Mac.
I wondered why my phone got frighteningly hot while charging sometimes. Then I saw the note after adding some faces manually for it to recognize, which was in the line of "Your phone will update faces when the phone is charging". My all photos are backed up to iCloud, btw.
Every now and then you get someone to think about an old problem on a clean sheet of paper and you might get a better result with less training data / investment.
This suggests that seeing the future a bit ahead of the rest of the world, and then assembling a motivated all-star team is (perhaps in the short term at least) one way of out-competing the "super AI" of the giants.
Don't let the name fool you, OpenAI is anything but Open.
In fairness it's not quite as good, but, it's good enough for the searches I've wanted to do so far and gets better all the time. And they're adding searching text in photos this release. I'm happy to wait a little for this better implementation.
https://a16z.com/2019/05/09/data-network-effects-moats/
https://a16z.com/2020/02/16/the-new-business-of-ai-and-how-i...
Google is actually pretty crappy at reverse image searches.
https://www.bellingcat.com/resources/how-tos/2019/12/26/guid...
(Until, of course, they force you out and give company to some crony oligarch. But that idea is also not unknown to Yandex, I believe?)
Also, I think you are overdramatizing this. Governments used to be omnipresent (maybe still are), in a different way, more threatening to individuals and probably as threatening to societies as "everything companies" could be.
What we currently call AI is very from AGI, and it's not clear that sitting on piles of proprietary data gives an edge towards AGI. If the goal is human-level intelligence, that has been demonstrably achieved with the far lesser resources of the public school system. :)
Current DL systems need huge amount of data, because they are very primitive: they work with immediate associations, so they require seeing data very similar to all possible inputs to generalize well.
As we develop more sophisticated systems, I expect that the leverage from data will tip over to engineering finesse, and nothing is better at fostering great engineering than the permissionless tinkering environment of open source.
Seems unlikely human education costs less than AI education in total.
Pretending that the scientifically managed public school system, that attempts to manufacture uniform educated humans on a conveyer belt, is responsible for human education is fairly ridiculous.
Children have a remarkable capacity to learn, and do so automatically through free play and exploration until public education wrings that curiosity out of them and turns education into a job.
Humans get educated despite the public education system, not because of it.
Say what now? There may be places on Earth that practice scientific management, there are definitely some that pretend to, but IME public school systems are neither.
You can read for yourself: https://files.eric.ed.gov/fulltext/ED566616.pdf https://radicalpedagogy.icaap.org/content/issue3_2/rees.html
The problem with the huge models like GPT-3 is that they are too expensive even to run by regular people, not train.
It is yandex who now collects massive amounts of data to improve their image search now, while google apparently doesn't.
Yandex is a giant, for sure, but google is, like, 10 times bigger and still doesn't provide the best service.
But I'm not too worried here because everyone gets access to larger datasets every year, and it gets cheaper to process every year, so whatever Microsoft or Google is capable of doing now, smaller companies will be capable of doing in a few years.
Where Microsoft does have an “unfair advantage” is in their marketing and sales firepower. Replicating their B2B and B2C sales channels is indeed very expensive. GitHub will be able to monetise Copilot by some upselling campaign. Then again, startups regularly manage to break into markets that are supposedly locked down by the likes of Microsoft.
I see a lot of people comparing human learning to machine learning in the comments, but there is a huge difference - we don't distribute copies of humans
By comparison, Copilot is even more obviously fair use.
I've had this conversation quite a few times lately, and the non-obvious thing for many developers is that fair use is an exception to copyright itself.
A license is a grant of permission (with some terms) to use a copyrighted work.
This snippet from the Linux kernel doesn't make my comment here or the website Hacker News a GPL derivative work:
ret = vmbus_sendpacket(dev->channel, init_pkt,
sizeof(struct nvsp_message),
(unsigned long)init_pkt, VM_PKT_DATA_INBAND,
VMBUS_DATA_PACKET_FLAG_COMPLETION_REQUESTED);
This snippet from an AGPL licensed project, Bitwarden, does not compel dang or pg to release the Hacker News source code: await _sendRepository.ReplaceAsync(send);
await _pushService.PushSyncSendUpdateAsync(send);
return (await _sendFileStorageService.GetSendFileDownloadUrlAsync(send, fileId), false, false);
Fair use is an exception to copyright itself. A license cannot remove your right to fair use.The Free Software Foundation agrees (https://www.gnu.org/licenses/gpl-faq.en.html#GPLFairUse)
> Yes, you do. “Fair use” is use that is allowed without any special permission. Since you don't need the developers' permission for such use, you can do it regardless of what the developers said about it—in the license or elsewhere, whether that license be the GNU GPL or any other free software license.
> Note, however, that there is no world-wide principle of fair use; what kinds of use are considered “fair” varies from country to country.
(And even this verbatim copying from FSF.org for the purpose of education is... Fair use!)
That case required that the output be transformative, in that "words in books are being used in a way they have not been used before".
Copilot only fits the transformative aspect if it is not directly reciting code, that already exists in the form that it is redistributing. So long as it does so, it fails to meet the criteria.
However - Copilot directly recites code. That is _very unlikely_ to fall under fair use.
Redistributing the exact same code, in the same form, for the same purpose, probably means that Copilot, and thus the people responsible for it, are infringing.
You make that statement as an absolute, but in the interests of clarity, all evidence so far shows that it directly recites code very rarely indeed. Even the Quake example had to be prompted by the specific variable names used in the original code.
In practice, the output code is heavily influenced by your own context — the comments you include, the variable names you use, even the name of the file you are editing — and with use it’s obvious that the code is almost certainly not a direct recitation of any existing code.
_Once_ is enough for it to be infringing. The law is not very forgiving when you try and handwave it away.
I tend to think it would be covered (provided it there were relatively small snippets and not entire functions).
I'm not American, but like others around here — I was just restricting the discussion to American law for simplicity's sake.
* Minus a few countries/regions targeted by US sanctions, I assume, though they've gradually broadened their services in sanctioned countries with the necessary licenses from OFAC.
I don't know the answer. I was only surprised that the commenter seemed dead sure that any and all copying (no matter how small) would be infringing.
That just doesn't correlate with my understanding of how Fair Use works: The "amount" of the infringement is one (of several) factors in determining if something falls under Fair Use:
>The third factor assesses the amount and substantiality of the copyrighted work that has been used. In general, the less that is used in relation to the whole, the more likely the use will be considered fair.
If it's so fair use, why not train it on all Microsoft code, regardless of license (in addition to GitHub.com) ? Would Microsoft employees be fine with Copilot re-creating "from memory" portions of Windows to use in WINE ?
Now, if you knew the code you wanted Copilot to generate you could certainly type it character by character and you might save yourself a few keystrokes with the TAB key, but it's going to be much MUCH easier to simply copy the whole codebase as files, and now you're right back where you started.
As Francois Chollet points out in this talk, ultimately deep neural network models are locally sensitive hash tables, so the examples of people pulling out source code is an inherent shortcoming of deep learning models in general. Give the right 'key' and you can 'recall' the value you are looking for.
Sounds like that wouldn't be difficult to fix? Transform the code to an intermediate representation (https://en.wikipedia.org/wiki/Intermediate_representation) as a pre-processing stage, which ditches any non-essential structure of the code and eliminates comments, variable names, etc., before running the learning algorithms on it. Et voila, much like a human learning something and reimplementing it, only essential code is generated without any possibility of accidentally regurgitating verbatim snippets of the source data.
(I guess these are going to depend a LOT on the jurisdiction that you're in ?)
Where is your ego when you're dead and gone? Where could we be if the majority of human advancement we're not tightly clutched as trade secrets?
As someone who has done paid software engineering (yes, you can feel free to call me a hack or sell out if you wish), I've come to find that the salary I've pulled over the years has not gone to me... But keeping a roof over those I love, helping other people's projects grow, giving people a shot, etc.
My time on the other hand, gets dumped into implementing the same handful of processes doing the same damn thing, but different this time, because you can't just bloody make "Here ya go, here's your Enterprise-in-a-box".
I'd like people more people able to solve novel problems than necessarily need to retread the same path over and over. Some degree of that will always have to be done to keep the skills fresh in the population, but we could do way better at marshaling that split, and I'm convinced part of what necessitates it is creating artificial barriers through things like enforced implementation monopolization. Yes. It ensures a minimum level of novelty and variance across populations, but it also does terribly at not consuming the finite amount of human capacity for truly novel thought to innovate.
It may make societies that function based on greed and economic/fiscal measures work, but I'm not convinced other incentive structures won't keep the rolling stone of innovation from accruing moss.
(Copyright has went IMHO overboard with its duration, we should scale to back to the original 14 years renewable once, just like patents, but copyright doesn't apply to processes anyway, and so arguably it shouldn't apply to software that can't claim to have any artistic merit.)
With no copyright/copyleft, how do you enforce the rule that derived works must provide access to the source code? I’ve never heard that copyleft was a stepping stone—rather, it’s the stick that fully realizes the four freedoms.
Absent this, I don't think there's a case. The courts have given extraordinarily wide latitude to fair use and ML algorithms are routinely trained on copyrighted works, photos, etc. without a license.
I understand that this feels more personal because it involves our field, but artists and authors have expressed the same sentiment when neural nets began making pictures and sentences.
The question here is no different than "Is GPT-3 an unlicensed, unlawfully created derivative work of millions, if not billions of people?"
No, I'm quite confident it is not.
It doesn't need to be substantial. In Google v. Oracle a 9-line function was found to be infringing.
The Supreme Court did hold that the 11,500 lines of API code copied verbatim constituted fair use.
Yes, because it was _transformative_, in a clear way. Because an API is only an interface. Which makes that part of that decision largely irrelevant to the topic at hand.
> Google’s limited copying of the API is a transformative use. Google copied only what was needed to allow programmers to work in a different compu-ting environment without discarding a portion of a familiar program-ming language. Google’s purpose was to create a different task-related system for a different computing environment (smartphones) and tocreate a platform—the Android platform—that would help achieve and popularize that objective.
> If I recall correctly, the nine line question wasn't decided by the supreme court, but the API question was.
It was already decided earlier, and Google did not contest it, choosing instead to negotiate a zero payment settlement with Oracle over the rangeCheck function. There was no need for the Supreme Court to hear it.
If they felt the nine line function made Google's entire library an unlicensed derivative work, they would have pressed their case.
That's not the case. It wasn't an out-of-court-settlement, but an agreement about the damages being sought, the court had already found it to be infringing, and that was part of the ruling.
But none of that changes that 9-lines is substantial enough to be infringing. It isn't necessary to be a large body of work.
> If they felt the nine line function made Google's entire library an unlicensed derivative work, they would have pressed their case.
No... It means the rangeCheck function was infringing. The implication you seem to have inferred here wouldn't be inferred by any kind of plagiarism case.
If Copilot is infringing, I suspect it's correctable (by GitHub) by adding a bloom filter or something like it to filter out verbatim snippets of GPL or other copyleft code. (And this actually sounds like something corporate users would want even if it was entirely fair use because of their intense aversion to the GPL, anyhow.)
I don't think your argument is as strong as you're making it out to be.
This goes to the "substantial" test for fair use. Clips from a film can contain core plot points, quotes from a book can contain vital passages to understanding a character, screen captures and scrapes of a website can contain huge amounts of textual detail, but depending on the four factors for fair use, still be fair use. (There have been exceptions though.)
The reaction on Hacker News to a machine producing code trained on their works is no different than the reactions artists and writers have had to other ML models. I suspect many of us are biased because it strikes at what we do and we think that our copyrights (because we have so many neat licenses) are special. They are not.
I think it would need to get to that level of "Copilot will emit a kernel module" before it's not obviously fair use.
After all, Google Books will happily convey to me whole pages from copyrighted works, page after page after page.
https://www.google.com/books/edition/Capital_in_the_Twenty_F...
it's anything but obvious. https://www.copyright.gov/fair-use/
> there is no formula to ensure that a predetermined percentage or amount of a work—or specific number of words, lines, pages, copies—may be used without permission.
9 lines of very run-of-the-mill code in Oracle / Google weren't considered fair use.
1. The act of training Copilot on public code
2. The resulting use of Copilot to generate presumably new code
#1 is arguably close to the Authors Guild v. Google case. You are literally transforming the input code into an entirely new thing: a series of statistical parameters determining what functioning code "looks like". You can use this information to generate a whole bunch of novel and useful code sequences, not just by feeding it parts of it's training data and acting shocked that it remembered what it saw. That smells like fair use to me.
#2 is where things get more dicey - just because it's legal to train an ML system on copyrighted data wouldn't mean that it's resulting output is non-infringing. The network itself is fair use, but the code it generates would be used in an ordinary commercial context, so you wouldn't be able to make a fair use argument here. This is the difference between scanning a bunch of books into a search engine, versus copying a paragraph out of the search engine and into your own work.
(More generally: Fair use is non-transitive. Each reuse triggers a new fair use analysis of every prior work in the chain, because each fair reuse creates a new copyright around what you added, but the original copyright also still remains.)
Not sure I see it that way.
If I take your hard work that you clearly marked with a GPL license and then make money from it, not quite directly, but very closely, how is that fair use? Or legal?
Copying and storing a book isn't recreating another book from it. Copilot is creating new stuff from the contents of the "books" in this case.
Edit: I misunderstood fair use as it turns out...
You can wipe your ass with the GPL license if your use of the product falls within Fair Use.
You can actually take snippets from commercial movies and post them onto YouTube if your YouTube video is transformative enough for your usage to be considered fair use. Well, theoretically at least - in reality YouTube might automatically copyright strike it.
>Copying and storing a book isn't recreating another book from it.
That doesn't mean that GitHub has to redistribute Copilot under GPL. However, the end user could potentially have to if they use Copilot to generate new code that happens to copy GPL code verbatim.
Is Copilot fair use? It's reading code, generating other code (some verbatim) and making money from it all while not having to release its source code to the world?
> That doesn't mean that GitHub has to redistribute Copilot under GPL
I wasn't saying that was the case: some of the code that Copilot used may not allow redistribution under GPL.
But let's say that all of the code it scanned was GPL for the sake of argument. Why would they not have to distribute their Copilot source yet, if I use it to generate some code, I'd have to distribute mine?
My spidey-sense it tingling at that one!
Again, fair use is an exception to copyright protection. If something is fair use, the license does not apply. The fact that Copilot does not release its source code is related only to a specific term of a specific license, which does not apply if Copilot is indeed fair use.
Stackoverflow on the other hand is much trickier question...
I do not find that to be obvious at all.
So following your own argument, even if Copilot is allowed, using it still risks you falling under GPL
For your browser analogy, that would mean that the "browser" is the copilot code, while the weights would be some data derived from GPL'd works, perhaps a screenshot of the browser showing the code.
I'd think that the weights/screenshot in this analogy would have to abide by the GPL license. In a vacuum, I would not think that the copilot code had to be licensed under GPL, but it might be different in this case since the copilot code is necessary to make use of the weights.
But then again, the weights are sitting on some server, so GPL might not apply anyway. Not sure about AGPL and other licenses though. There is likely some illegal incompatibility between licenses in there.
No let’s substitute a different database of for the code that isn’t SO. It doesn’t really matter if that database is a literal RDBMS, a giant git repo or is encoded as a neural net. All copilot is going to do is perform a search in that database, find a result and paste it in. The burden of licensing is still on me to not use GPL code and possibly on the person hosting the database.
The gotcha here is that copilot’s database is a neural network. If you take GPL code and feed it as training data to a neural network to create essentially a lookup table along with non-GPL code did you just create a derived work? It is unclear to me whether you did or not. In particular, can they neural network itself be considered “source code”?
Some good responses in sibling comments already, but I don't see the narrow answer here, which is: No, because no distribution of the browser took place.
If you created a weird version of the browser in which a specific URL is hardcoded to show the GPL'd code instead of the result of an HTTP request, and you then distributed that browser to others, then I believe that yes, you'd have to do so under the GPL. (You might get away with it under fair use if the amount of GPL'd code is small, etc.)
Not sure if you meant to reply to me but I agree with you: you can't compare what Google did to what Copilot does.
Suggesting code is generating code
Neither. Someone else did, and published it. Copilot copied the dialog and suggested it.
> If you change names in those lines of dialogue to fit your story, do you now gain credit for writing those lines?
It depends. Talking generalities isn't productive or interesting. Can you give an example and we can discuss specifics?
> Suggesting code is generating code
This isn't even superficially true
But the other, more direct question is ... what about the instances where Copilot doesn't come up with a learned mishmash result? What happens when Copilot just gives you a straight up answer from it's learning data, verbatim?
Then you, as a dev, end up with a bunch of code that is effectively copied, via a 'copying tool', which is GPL'd?
It's that specific case that to me sticks out as the 'most concerning part'.
Please correct me if I'm wrong.
Fair use is an exception to copyright and, by definition, copyright licenses.
Google didn't create new books from the contents of existing ones (whether you agree that they should have been allowed to store the books or not) but Copilot is creating new code/apps from existing ones.
Edit: I guess my understanding of fair use was wrong. I stand corrected.
Copilot producing new, novel works (which may contain short verbatim snippets of GPL works) is a strong argument for transformativeness.
I don't know how a court would decide this, but I do think the facts in future GPT-3 cases are sufficiently different from Author's Guild that I could see it going any way. Plus, I think the prevalence of GPT-3 and the ramifications of the ruling one way or another could lead some future case to be heard by the Supreme Court. A similar case could come up in California, or another state where the 2nd Circuit Artist Guild case isn't precedent.
Define short
Fair use is a defense for cases of copyright infringement, which means you're starting of from a case of copyright infringement, which sort-of muckys up the whole "innocent until proven guilty" thing. And considering it's a weighted test, it's hardly very cut-and-dry at that.
If I'm Google, and I scan your code and return a link to it when people ask to find code like that (but show an ad next to that link for someone else's code that might solve their problem too), that's fair use and legal. My search engine has probably stored your code in a partial format, and that's fine.
[1] https://www.gnu.org/licenses/gpl-faq.en.html#DoesTheGPLAllow...
For commercial use and derivative works?
Authors won't incorporate snippets of books into new works unless they're reviews. Copilot is different.
If anything, the ways in which Copilot is different aid Microsoft/GitHub's argument for fair use. Because Copilot creates novel new works, that gives them a strong argument their system is more transformative than Google Books, which just presents verbatim copies of books.
Copilot does none of that. If all the ML companies are so sure this is fair use I encourage them to train an AI on Disney movies to generate short cartoon snippets based on some description. There sure would be a court case.
Of course they do, previous works are quoted all the time.
Citing your source is not a get out jail free card for copyright infringement, it doesn't really matter.
No, but it's a requirement of the license stackoverflow.com uses, which is unfortunate, for code (as opposed to text, where a quote can be easily attributed).
Plagiarism is not the same as infringement.
They couldn't do it with a license, which only imposes conditions for the license to be valid. Fair use applies even if the copier has no license at all.
Potentially they could do it with a contract. A license is not a contract and imposes no covenants on the parties.
The question on this one will be about the difference between Microsoft/Github's product and a programmer using copilot's code:
"If I feed the entire code base to a machine, and it copies small snippets to different people, do we add the copies up, or just look at the final product?"
When fair use is an issue, the courts look at the facts in context each time. These are obviously different facts than scanning books for populating a search index and rendering previews; and each side is going to argue that the facts are similar or that they are dissimilar. How the court sees it is going to be the key question.
1. a fascinating Supreme Court opinion.
2. a frustrating ruling because SCOTUS doesn't understand software and code.
3. the type of anti-anticlimactically(?) narrow ruling typical of the Roberts court.
While our Congresspersons can't seem to wrap their minds around technology/social media, I think SCOTUS would understand this one enough to avoid (2).
If Github had a service that automatically mirrored public repositories on Gitlab, that would be equivalent to the example you gave.
But Github is taking content under specific licenses to build something new for commercial use.
I'm not sure if what Github does falls under Fair Use, but I don't know that it matters. I can read fifty books and then write my own, which would certainly rely—consciously or not—on what I had read. Is that a copyright violation? It doesn't seem like it is but maybe it is and until now has been impossible to prosecute?
The end user is.
By this logic any and all neural nets that draw pictures are copyright infringing as well.
What's more, if any of the code implements a patent, fair use does not cover patent law, and relying on fair use rather than a copyright license does not benefit from any patent use grant that may be included in the copyright license. If a codebase infringes a patent due to Copilot automatically adding the code, I can easily imagine GitHub being attributed shared contributory liability for the infringement by a court.
Not a lawyer, just a former law student and law feel layman who has paid attention to these subjects.
What a weird autocorrect typo. This should have read "law geek layman." (And it initially autocorrected again as I was typing this paragraph.)
This mostly has to do with the nature of the wishy-washy nature of the 4 part Fair Use test, which, unlike decent legal tests, doesn't actually have discrete answers. The judge looks at the 4 questions, talks about them while waving her hands, and makes a decision.
Comparing to, e.g., Patent, where you actually do have yes-or-no questions. Clean Booleans. Is it Novel? Is it Non-Obvious? Is it Useful? If any of the above is "No", then no patent for you.
As for the execution of Fair Use, while I haven't gone too deep into Software, I can assure that for music, the thing is just a silly holy-hell mess; confirmed most recently by the "Blurred Lines" case, where NO DIRECT COPYING (e.g. sampling or melody taking) was alleged, merely that the song sounded really similar to "Got to give it up" and that was enough.
So then, I'd say everything either is, or should be, up in the air, when it comes to Fair Use and software.
Those questions for patents are barely more clear-cut than copyright fair use tests, there is lots of room for disagreement.
It's definitely true that a fair use defense against copyright infringement varies a lot by the field of work and norms can develop which are relevant to court cases. The music field is a mess, the "Blurred Lines" judgement was total bullshit. But the software field is not without its own copyright history and norms so there's no reason to expect everything to go to hell.
The big guns like Microsoft, Google, Oracle, do this sort of thing as a matter of course in their business activities, they have the lawyers, the money, and the ear of members of parliaments, senators etc.
Whereas an individual or small business probably wants to conduct themselves within a more narrow set of adherences.
Is (was?) a swipe gesture novel? Is it non-obvious?
All that said, the one thing I'd add about fair use is that it isn't permission to use anything you like, but rather a defense in a legal proceeding about copyright. It's pretty much all about being able to reference copyrighted material with the law later coming in and making final decisions on whether or not that reference went too far. (IE, copying all of a disney movie and saying "What's up with this!" vs copying 1 scene and saying "This is totally messed up and here's why".)
That was a big part of the google oracle lawsuit.
The unauthorized copy arises when someone gets the work out of the model.
Of course if you make a model explicitly for the purpose of evading copyright then the courts will see through that ploy.
You are correct about (US specific) the fair use exception, but it is in no way as clear as you suggest that what copilot is doing entirely falls under fair use. Fair use is always constrained.
I suspect some variant of this sort of thing will have to be tested in court before the arguments are really clear.
I disagree that Copilot is "more obviously fair use.", some parts might be, but we have seen clear examples (i.e verbatim code reproduction) that would not be.
I dont believe the question of "is this fair use" is as clear as you believe it to be
I think the factor most at risk in a fair use test with Copilot is whether it ever suggests verbatim, code that could be considered the "heart" of the original work. The John Carmack example that's popped up here at least gets closer to this question, it was a relatively small amount but it was doing something very clever and important.
One can imagine a project that has thousands of lines of code to create a GUI, handle error conditions, etc. that's built around a relatively small function; if Copilot spat out that function in my code, it might not be fair use because it's the "heart" of the original work. Additionally, its inclusion in another project could affect the potential market for the original, another fair use test.
But Copilot suggesting a "heart" is unlikely, something that would have to be ruled on in a case-by-case basis and not a reason to shut it down entirely. Companies that are risk-averse could forbid developers from using Copilot.
I agree with you that the relative importance of the copied code to the end product would be (or should be) the crux of the issue for the courts in determining infringement.
This overall interpretation most closely adheres to the spirit and intent of Fair Use as I understand it.
...and if you're outside the USA?
To me, seeing youtube-dl's case as fair use is so much easier than using hundreds of thousands source code files without permission in order to build a proprietary product.
My point was however that I'm just utterly failing to see how the youtube-dl test thing could be more of a copyright problem than this entire thing based on millions of others' works that is Copilot.
Would it be 'fair use' for the devlopers to simply copy code from those repos - even just 10 lines, and claim 'fair use' - i.e. circumventing Copilot?
Even if Copilot is 'fair use' ... does that mean the results are 'fair use' on the part of AutoPilot users?
And a bigger question: is your interpretation of those statues and case law enough to make the answer unambiguous?
I don't have legal background, but I do have an operating background with lawyers and tech ... and my 'gut' says that anyone using Copilot is opening themselves up to lawsuits.
If the code you put in your software comes, via Copilot, but that code is verbatim from some kind of GPL's (or worse, proprietary) ... there's a good chance you could get sued if someone gets the inclination.
Maybe it's because of my personal experience, but I can just see corporate lawyers banning Copilot straight up as the risks are simply now worth the upside. That's now what we like to hear in the classically liberal sense i.e. 'share and innovate' ... but gosh it doesn't feel like a happy legal situation to me.
Looking forward to people with more insight sharing on this important topic.
Only a lawyer (and truly, only a court) could answer that question.
If you copy 100 lines of code that amounts to no more than a trivial implementation in a popular language of how to invert a binary tree, it's likely fair use.
If you copy 10 lines of code that are highly novel, have never been written before, and solve a problem no one outside the authors have solved... It may not be fair use to copy that.
Other people who have replied have mentioned "the heart" of a work. The US Supreme Court has held that even de minimis - "minimal", to be brief - copying can sometimes be infringement if you copied the "heart" of a work.
You can be violating copyright without plagiarizing, so long as you cite your source, but if you copy a copyright-protected work in an illegal way when doing so.
And you can be plagiarizing without violating copyright, if you have the permission of the copyright holder to use their content, or if the content is in the public domain and not protected by copyright, or if it's legal under fair use -- but you pass it off as your own work.
Two entirely separate things. You can get expelled from school for plaguriism without violating anyone's copyright, or prosecuted for copyright without committing any academic dishonesty.
You can indeed have the legal right to make use of content, under fair use or anything else, but it can still be plagiarism. That you have a fair use right does not mean "Oh so that means you are allowed to turn it in to your professor and get an A and the law says you must be allowed to do this and nobody can say otherwise!" -- no.
Exactly the point I came to make.
The Authors’ Guild is a US entity, and so is Google, so only US law applies. And thus, we have the Fair Use exception.
But developers sharing code on GitHub come from and live all over the world.
Now, Github’s ToS do include the usual provision stating that US & California law applies, et cætera, et cætera [1], but… and even they acknowledge it may be the case, such provisions usually aren’t considered legal outside of the US.
So… developers from outside the US, in countries with less lenient exceptions to copyright, definitely could sue them.
Identifying these countries and finding those developers, however, is a different matter altogether.
[1]: https://docs.github.com/en/github/site-policy/github-terms-o...
That's a reference to factor four of the fair use test, "the effect of the use upon the potential market for or value of the copyrighted work." (17 USC 107).
None of the factors are dispositive, however. For example, a scathing book review that quotes a passage to show how bad the writing is might eviscerate sales of the book, but such a use is usually protected. For a counter-example, see Harper & Row v. Nation Enterprises 471 U.S. 539 (1985).
I'm really out of my depth in giving my own opinion here, but I'm not sure that either the "distribution != derivative" characterization, or that "parsing GPL => derivative of GPL" really locks this thing down. The bit that I can't follow with the "distribution != derivative" argument is that the copilot is actually performing distribution rather than "design". I would have said that copilot's core function is generating implementations, which to me does not seem like distribution. This isn't a "search" product, and it's not trying to be one. It is attempting to do design work, and I could see a case where that distinction matters.
For Copilot itself, I do see the case for fair use, though it gets fuzzy should Microsoft ever start commercializing the feature. Nevertheless it remains to be seen whether ML training fits the same public policy benefits public libraries and free debate leverages to enable the fair use defense.
For Copilot users, I don't see an easy defense. In your hypothetical, this would be akin to me going on Google books and copying snippets of copyrighted works for my own book. In the case of Google books, they explicitly call out the limits on how the material they publish can be used. I'm contrast, Copilot seems to be designed to encourage such copying, making it more worry some in comparison.
A book completely written by pasting passages of other books would actually be a pretty interesting transformative work.
While software is in this limbo between copyrights and patents...
Copilot will not write an entire software module, it will provide you with snippets. I see using GPL code for training fair use. If a developer reads the source code of a project to take inspiration and possibly copy some small parts does it violate the license?
More precisely, fair use is an affirmative defense to an claim of copyright infringement. A fair use defense basically says, "Yes, I am copying your copyrighted material and I don't have a license (or am exceeding a licensed use), but my usage is allowed under the fair use doctrine (codified in 17 USC 107 in US law)."
And copyright itself is an exception to the normal state of things : the public domain, copyright being only a temporary monopoly.
If relating this to how humans learn, books and other sources are used to inform understanding and human knowledge. One can purchase or borrow a book without actually owning the copyright to it. Indeed, a given passage may be later quoted verbatim, provided it is accompanied with a reference to its source.
Otherwise, a verbatim use without attribution in authored context is considered plagiarism.
So, sure one can use a multitude of material for the training. Yet, once it gets to the use of the acquired "knowledge" - proper attribution is due for any "authentic enough" pieces.
What is authentic enough in this case is not easy to define, however.
At some point Neural Nets like GameGAM might be good enough to duplicate (and optimize) a commercial game. Can you then release your version of the game? Do you just need to make a few tweaks? Are we going to get a double standard because commercial interests are opposed depending on the use case?
It would be pretty funny if Microsoft as a game publisher lobbies to prevent their IP being used w/ something like GameGAN, but then takes the opposing stand point for something like their CoPilot! Although I'm sure it'll be spun as "These things are completely different!".
Maybe we will some day, but for now this isn't the case, where the law is concerned :
https://ilr.law.uiowa.edu/print/volume-101-issue-2/copyright...
There's no point to copilot without training data, some but not all of the training data was (A)GPL. There's no point to github without hosting code, some but not all of the code it hosts is A(GPL).
The code in either cases is data or content, it has not actually been incorporated into the copilot or github product.
GitHub's TOS include granting them a separate license (i.e., not the GPL) to reproduce the GPL code in limited ways that are necessary for providing the hosting service. This means commonsense things like displaying the source text on a webpage, copying the data between servers, and so on.
Sure one could argue that Copilot learned in the way a human does. There is nothing that prevents one from learning from copyrighted work, but snippets delivered verbatim from such works are surely a copyright violation.
Of course the interesting part is that the user not only has no idea what that license is but also where the code came from and if it is in fact copied verbatim. It's unlikely a court would agree that putting licensed code through a machine strips the licensing requirements of the code, of course, but that doesn't seem to be Microsoft's problem.
I think Microsoft's use of public code hosted on GitHub is covered by the terms of service but if this use includes granting a license more permissive than the license indicated on the code itself, this would probably put every GitHub user who ever committed less permissively licensed code to GitHub that they didn't control in violation of those licenses.
There's really only three ways this can go:
1) Machine learning does legally become a license-stripping black box, which would allow creating a machine generated commons by feeding arbitrary copyrighted works into sloppy AIs that mostly just replicate their input without changes.
2) Copyright law is extended to consider the output of machine learning as derived works from its inputs, massively extending the reach of copyright and creating massive headaches for everyone (e.g. depending on the exact ruling this would effectively make it impossible to reproduce a digital artwork as merely rendering it on a screen would create a derived work).
3) The original licenses are upheld and remain in effect, rendering the output of Copilot useless by creating a massive legal headache for anyone trying not to violate copyright.
I think outcome 2 is unlikely but 1 and 3 aren't mutually exclusive.
A bit of a tangent and it’s fictional, but I really have to recommend the tale of MMAcevedo. https://qntm.org/mmacevedo
I know HN loves a good "well actually" and Microsoft is always suspect, but let's leave the idea of code laundering to the Oracle lawyers. Let hackers continue to play and solve interesting problems.
Copilot should be inspiring people to figure out how to do better than it, not making hackers get up in arms trying to slap it down.
It only happens at boss level when tech giants litigate IP issues.
Also people rarely do it; I've caught maybe a couple instances of it in my career and I never really thought too much about them again. This tool helps make it a lot easier and more common. I have a feeling other people chiming in are also in the camp of "Oh, this is going to be a thing now, huh?"
I also can't help but to think that my negative opinion of it isn't solely based on this provenance issue. While it's cool it seems questionable about how practical it is. If the value was more clear I think I could stomach the risk a bit better.
One of the (many) problems is that GitHub/Microsoft already benefit from runaway network effects so it’s difficult to “do better”. Where will you get all of that training code if not off GitHub?
The real answer to this is to yank your projects from GitHub now while you search for alternatives.
Whether or not this move is “legal”, it should serve as a wake up call that GH is not actually a service we should be empowering. This incident is just one example of why that’s a bad idea.
You would have a much stronger case if they had taken your code from elsewhere.
I don't care if MS copies my hobby projects exactly, but I'm not sure my employer(defense contractor) would even be allowed to use a tool like this.
I think it looks cool though. I will probably try it out if it is ever available for free and works for the languages I use.
Personally, I think that in the age of AI programming any notions of code licensing should be abolished. There is no copyright for genes in nature or memes in culture; similarly, these shouldn't be copyright for code.
I still think we're a long way from that. Copilot will help write code quicker, but it's not doing anything you couldn't do with a Google search and copy/paste. Once developers move beyond the jr. level, writing code tends to become the least of their worries.
Writing the code is easy, understanding how that code will affect the rest of the system is hard.
It's just a smarter tab-completion.
https://analyticsindiamag.com/open-ai-gpt-3-code-generator-a... has a bunch of videos of this in action.
Since when are humans not a part of nature?
I feel like this comment misunderstands what a software developer is doing. Copilot isn't going to understand the underlying problem to be solved. It's not going to know about the specific domain and what makes sense and what doesn't.
We're not going to see developers replaced in our lifetime. For that you need actual intelligence - which is very different from the monkey see monkey do AI of today.
Having a semi-intelligent monkey that can fetch obvious things off the shelf, build very basic control structures, and do the boring little housekeeping tasks is bad for the craft of programming but very good for the good-enough-solution situation. I can see it having the same impact as cheap and widely available digital cameras; anyone can be a kinda decent photographer now, but if you want to be a professional you're probably going to have to work a lot harder to stand out, whether that's by development of craft, development of narrow technical expertise and fancy equipment, or development of excellent business skills.
Photography is a good analogy - with everyone having fancy cameras you could think that a photographer is now not necessary. But yes there are still photographers about - they see things that the average person doesn't. The camera doesn't tell them what type of photos to take, what composition the photo should have or what poses a model should have.
New genetic sequences are patentable, not copyrightable, but that because of the process involved in creating new genetic sequences more then the genes themselves.
Sure naturally occurring genes aren't patentable, but it's not like we have code growing on trees. So that's a terrible comparison.
If they made the trained model public (and also trained it on private code) the response would be completely different.
I expect that we'll need new copyright law to protect creators from this kind of thing (specifically, to give creators an option to make their work public without allowing arbitrary ML to be trained on it). Otherwise the formula for ML based fair use is "$$$ + my things = your things" which is always a recipe for tension.
If you're asking about the moral reaction here, I think it depends on how one views Copilot. Does Copilot create basically original code that just happens to include a few small snippets? Or does Copilot actually generate a large portion of lightly changed code when it's not spitting out verbatim copies of the code? I mean, if you tell Copilot, "make me a QT compatible, crossplatform windowing library" and it spits out a slightly modified version of the QT source code and if someone started distributing that with a very cheap commercial license, that would be a problem for the QT company, which licenses their code commercial or GPL (and as QT a library, the QT GPL forces user to also release their code GPL if they release it, so it's a big restriction). So in the worst case scenario, you can something ethically dubious as well as legally dubious.
Copilot should be inspiring people to figure out how to do better than it, not making hackers get up in arms trying to slap it down.
Why can't we do both? I mean, I am quite interested in AI and it's progress and I also think it's important to note the way that AI "launders" a lot of things (launders bias, launder source code, etc). AI scanning of job applications has all sorts of unfortunate effects, etc. etc. But my critique of the applications doesn't make me uninterested in the theory, they're two different things.
Still, some of the moral outrage here has to do with it coming from Github, and thus Microsoft. Software startup Kite has largely gone under the radar so far, but they launched this back in 2016. Github's late to the game. But look at the difference (and similarities) in responses to their product launch posts here.
https://news.ycombinator.com/item?id=11497111 and https://news.ycombinator.com/item?id=19018037
Maybe Github isn't violating the licenses of the programmers who host on them. Maybe Copilot doesn't just spit out code that belongs to other people. Those are matters of interpretation and debate.
But if Github was doing this with Copilot, virtually an open source programmer would have a reason to be upset. Open source programmers don't give their code out for free they license it. This is a legal position, not a feeling. "Intellectual property" may be a pox on the world but asking open source developers to abandon their licenses to ... closed source developers, is legitimately a violation.
And before the spitting out source code problem appeared, I recall quite a few positive responses to Copilot. Lots of people still seem excited. And yeah, people are looking at the downside given Microsoft's long abusive history but hey, MS did those thing.
but then again I migrated away from github as soon as MS bought it
still, it's a matter of principle
"Hackers" "playing" and ignoring copyright is fine, but Copilot isn't promoted as a toy, it's promoted as a tool for professional software development. And in that framing it is about as dangerous as an untrained intern with access to the production server.
Want Linux to run on your thing? You must publish driver source then or you're violating copyright law. This was less a big deal before device vendors ratcheted the pathological behavior up to 11 with smartphones and that's why far more people seem to react far more strongly now.
First, you might choose to distribute your code under a copyleft license to advance the OSS ecosystem. Second, the older you get, the more experience you accumulate, paradoxically the harder it is for you to find a job or advance your career in this industry—so, to maintain at least some source of motivation for tech companies to hire you, you may choose to make some of the source available, but reserve all the rights to it.
You’re fine making the source of your tool or library open for anyone to pass through the lens of their own consciousness and learn from it, but not to use as is for own benefit.
Now with GitHub Copilot suddenly you see the results of your labour you’ve previously made (under the above assumptions) public being passed through some black box, magically stripped from your license’s protections, and used to provide ready-made solutions to everyone from kids cheating at college tests to well-paid senior engineers simply lacking your expertise.
I hope it’s easy to spot how engineer’s interests in the above example are not necessarily aligned with GitHub’s, how this may be perceived as an unfair move disadvantaging veteran rank-and-file software engineers while benefitting corporate elites and investors, and subsequently has the potential to disincentivize source code sharing and deal a blow to OSS ecosystem as a whole.
I have conjured up two scenarios here:
Let's say I use copilot to generate a bunch of code for an app, something substantial, and it regurgitates a load of bits and pieces from many sources it got from GitHub, I'd assume there won't be any attribution in it... it will be as if Copilot made the code itself (I know it sort of does but lets not split hairs!). I'm guessing the prevailing theory (from GiitHub anyway) is that I'm legitimately allowed to do this.
Now, let's say I generated all that code by manually copying and pasting chunks of code from a whole bunch of repos, whether they are open source, unlicensed, whatever. Would I not be ripe for legal issues? I could potentially find all the code that copilot generated and just copy and paste it from each of the sources and not mention that in my license. What if I told everyone "yeah, I just copied and pasted this from loads of Github repos and didn't put any attribution in my code". I'd assume that (morality aside) I'd be asking for trouble!
Am I missing something? Am I misunderstanding the situation, or the capabilities of copilot?
Superficially at least, Copilot (from my understanding) is "copying" code, letting me use it in my app, and making money from it.
I'm just trying to wrap my head around it.
Let's be clear, I am not a lawyer, but it seems... strange!
- Only a very small proportion of Copilot generated code is reproduced verbatim, so if you specifically built a product just from copied-verbatim code, your act of selecting and combining those pieces of copyrighted code would be creating a derivative work.
- GitHub is not selling the copyrighted code, they are selling the tool itself. Google is literally the same thing: you could theoretically create a product by googling for prefixes of copyrighted code and then copying the remainder straight out of the search results. It's you who would be violating copyright, not google.
The strange bit is how they are allowed to use other peoples code to create derivative works (this is how I see it from my non-legal perspective anyway).
Even if it's legal (to the letter of the law, not the spirit) it leaves a sour taste.
Famous code: https://en.wikipedia.org/wiki/Fast_inverse_square_root#Overv...
GitHub claims they didn't find any "recitations" that appeared fewer than 10 times in the training data. That doesn't mean it's a completely solved issue (some code may be repeated in many repositories but always GPL, and there are limitations to how they detect recitations), but from rare cases of generating already-common solutions people seem to be concluding that all it does it copy paste.
[0]: https://slate.com/technology/2016/08/in-copyright-law-comput...
Edit: > There's a decent bit of caselaw indicating that computers reading and using a copyrighted work simply "don't count" in terms of copyright infringement.
That means their computer can read any code it wants, do whatever it wants with the code, then they can monetise that by giving YOU the code. Would they then be indemnified by saying "no Microsoft human read or used this code"?
However, if you then use the code and look at it, does that make you liable?
I would be shocked if GitHub's lawyers didn't argue that using copyrighted material as training data for an AI model is highly transformative. There may be snippets available from the original but they are completely divorced from their original context and virtually unrecognizable unless they happen to be famous like the Quake inverse square root algorithm. And I think GitHub's lawyers would also argue that Copilot's use does not affect the _original_ market -- e.g. it does not hurt Quake's sales if their algorithm is anonymously used in a probably totally unrelated codebase.
Your counterexample would probably fail both tests -- it's not transformative use if your software hands out complete pieces of copyrighted software, and it would definitely affect the market if Copilot gave me the entire source code of Quake for my own game.
[0]: https://fairuse.stanford.edu/overview/fair-use/four-factors
That being said, creating a transformative work from something else is considered fair use. So, for example, if I read a whole bunch of books and then, heavily influenced by them, create my own, similar book, that would be fair use I suppose... that makes sense.
But, where does the derivative works come in? Where do you draw the line?
If I am heavily influenced by billions of lines of other people's GPL code (ala Copilot!), then I create my own tool from it and keep my code hidden, does that not mean I am abusing the GPL license?
I have read variations of "computers don't commit copyright" more times than I can count in the past few days.
How is Copilot different from a compiler? (Please give me the legal answer, not the technical answer. I now the difference between Copilot and a compiler, technically.)
Isn't a compiler a computer program? How is its output covered by copyright?
Am I fundamentally misunderstanding something here?
A compiler is run on original sources. I don't see any analogy here at all.
* They both produce software as output.
* They both transform their input.
* They both can combine different works to create a derivative work of each work. (Compilers do this with optimizations, especially inlining with link-time optimization.)
They really do the same things, and yet, we say that the output of compilers is still under the license that the source code had. Why not Copilot?
Because the sources used for input do not belong to the person operating the tool.
If you say that doesn't matter, then you are saying open source licenses don't matter because the same thing applies - I could just run a tool (compiler) on someone else's code, and ignore the terms of their license when I redistribute the binary.
If I take some code I don’t have a license for, feed it to a compiler (perhaps with some -O4 option that uses deep learning because buzzwords), then is the resulting binary covered under fair use, and therefore free of all license restrictions?
If not, then how is what Copilot is doing any different?
No, the binary is not free of license restrictions. Read any open source license - there are terms under which you can redistribute a binary made from the code. For GPL you have to make all your sources available under the same terms for example. For MIT you have to include attribution. For Apache you have to attribute and agree not to file any patents on the work in Apache licensed project you use. This has been upheld in many court cases - though it is not always easy to find litigants who can fund the cases the licenses are sound.
Authoring is the act that causes a work to be copyrightable. In most jurisdictions, authoring a work automatically causes copyright to subsist in the work to some degree. The purpose of the copyright system is to encourage people to author new, original works, by rewarding those who do with exclusive rights. It is well-known that only humans can author a work. Computers simply cannot do it. If your computer (by some kind of integer overflow UB miracle) accidentally prints out a beautiful artwork, NOBODY has exclusive copyright over it, and anyone may reproduce it without limitation. Same goes for that monkey who took a selfie.
What a compiler does, on the other hand, is adapt a work. Adapting a work is not authoring it. Sometimes when you adapt a work, you also author some original work yourself, like when you translate a book into another language. When a compiler (not a linker) transforms source code, it absolutely, 100% definitely does NOT add any original work; the executable or .so/.a/.dylib/.dll file is simply an adaptation of the original work. The copyright-holder of the source code is the copyright-holder of the machine code. An adaptation is also known as a "derivative work".
(Side note; copyleft licenses boil down to some variation of "if you adapt this, you have to share everything in the derivative work, not just the bits you copied.")
Adaptation is a form of reproduction. It's copying. "Distribution" also often involves copying, at least on the internet. (Selling or giving away a book you have purchased does not constitute copying.) Copying is one of the exclusive rights you have when you own the copyright in a work, that you may then license out.
It gets more complicated when the computer uses fancy ML methods to produce images/text out of things it has seen/read. You can't simplify the law around that to a simple adage digestible enough to share memetically on HN and Twitter. One thing is certain: if the computer did it, by itself, then no original work was authored in the process. That poses a problem for people who write the name of a function and get CoPilot to write the rest; if you do that, you are not the author of that part of the program. If you use it more interactively that's a different story.
There is, however, always a question of whether the copyright in the original works the computer used still subsists in the output.
My rough framing of the licensing issues around CoPilot is therefore as follows:
1. The source code to CoPilot is an original work, and the copyright is owned by GitHub.
2. When GH trained CoPilot's models on other people's works, was that copying? (This one is partially answered. It can spit out verbatim fragments, so it must be copying to some extent, rather than e.g. actually learning how to code from first principles by reading.) If it was not all copying, how much of it was copying and how much of it was something else? What else was it?
3. If GH adapted the originals, what is the derivative work? (I.E. where does the copyright subsist now? Is is a blob of random fragments of code with some weights to a neural network?)
4. Which works is it an adaptation of? You might think "all of them, and for each one, all of the code" but I'm not so sure. For example, imagine the ML blob contains many fragments, but some are shorter than others. If your program has "int x;" in it, and CoPilot can name a variable "x", you can hardly claim that as your own. I'm most interested in whether the mere fact of CoPilot having digested ALL of it, having fed this into the mix and producing a ML blob based on all that information, means that the ML blob is a derivative work of all of them. Or whether there is some question of degree.
5. Fair use. Was it fair use to train the model? Is it, separately or not, fair use to create a commercial product from the model and sell it? Fair use cares about commercial use, nature of the copied work, amount of copying in relation to the whole, and the effect on the market for / value of the copied work. Massive question.
6. If not fair use, then GH is subject to the licenses and how they regulate use of the works. What license conditions must GH comply with when they deal with the derivative work, and how? Many will be tempted to jump straight to this question and say GH must release the source code to CoPilot. I'm not yet convinced that e.g. GPL would require this. I can't believe I'm writing this, but is the ML blob statically or dynamically linked? Lol.
7. Final question, is there some way to separate out works which were copied with no fair use (or not copied at all), from works which were copied with no fair use? People are worried about code laundering, e.g. typing the preamble to a kernel function and reproducing it in full. In that situation, it is fairly obvious that the end user has ultimately copied code from the kernel and needs to abide by GPL 2.0; moreover if they're using CoPilot to write out large swathes of text they will naturally be alert to this possibility and wary of using its output. But think of the converse: if there is no way to get CoPilot to reproduce something you wrote, what's the substance of your complaint? Is CoPilot's model really a derivative of your work, any more than me, having read your code, being better at coding now? Strategically, if you wanted to get GH to distribute the model in full, you might only need one copyleft-licensed, verbatim-reproducible work's owner to complain. But then they would just remove the complainant's code. You might be looking at forcing them to have a "do not use in CoPilot" button or something.
Also, I loved this quote:
> Copying is one of the exclusive rights you have when you own the copyright in a work, that you may then license out.
I've been paying attention to software copyright topics for more than twenty years and never thought of it in exactly these terms. Its right there in the name - the right to copy it - and determine the terms under which others can copy it is exactly what a copyright is!
Copyright on art gets more interesting / fuzzier. The key part is substantial similarity - https://en.wikipedia.org/wiki/Substantial_similarity and https://www.photoattorney.com/copyright-infringement-for-sub...
Rather than text, my AI copyright hypothetical... consider a model created based on sunset photographs. You take a regular photograph, pass it through the model, and it transforms it into a sunset. The model was trained on copyrighted works but the model is considered fair use.
Now, I go and take a photograph from some location during the day and then pass it through the transformer and get a sunset. Yea me! Unbeknownst to me, that location is a favorite location for photographers and there were sunsets from that location used in the training data. My photograph, transformed to look like a sunset is now similar to one of them in the training data.
Is my transformed photograph a derivative work of the one in the training data to which it bears similarity to? How would a judge feel about it? How does the photographer who's photograph was used in the training data feel?
Maybe it needs to be paired with another network / hunk of code that checks for verbatim copying?
No. Copilot is a technical preview. In the final release, if it reproduces code verbatim, it'll tell you and present the correct license.
https://docs.github.com/en/github/copilot/research-recitatio...
Are you doing that? If not, then I wouldn't use GitHub's use as justification to engage in copyright infringement.
Will it check the LICENSE file? Simply having a LICENSE file is not a declaration that all the code in that repo is under that LICENSE.
What if specific lines/files are specified to be under different licenses?
What if the publisher of the repo is publishing it under an incorrect license in bad faith?
Will github be responsible if it tells me the wrong license?
I don't see this as fundamentally different. It's unlikely that the Free Software Foundation is going to track you down for including some GNU code in your single-user repo. If you used their stuff in a popular commercial project and they got wind of it, you might expect to receive a cease and desist at best.
See also : Napster, including how it was condemned for facilitating copyright infringement (what Microsoft is risking here, though the offense is likely to be much milder, of course).
I mean, sure you don't copy an entire file, but you tend to copy a snippet, or in the end you look at how is done and you done the exact same way (that is the same of copying it!)
I would say there is not a problem in there.
That's some nasty walled garden terms... I wonder how much these kinds of ToS are actually legal ?
So would you say that it's publicly visible?
Isn't that what autopilot is doing here? The system is merely learning how to code, and then applying it's learnings on other programming problems. It's not like it's writing software to specifically compete with other programs.
That's at best a technical issue. What way too many people claim, however, is that the machine isn't even allowed to look at GPL'ed code for some reason, while humans are.
I'd like to learn the reasoning behind that.
Why would those be the same thing? It's a matter of scale. Just like how people are allowed to read websites, but scraping is often disallowed.
Hosting code on Github explicitly allows this type of usage (scraping) according to their TOS so I have to ask again - why the sudden complains?
Are we still talking about a shortcoming of the ML model, which very occasionally spits out a few lines of copied code or should we include search engines into this, because they do the exact same thing by design?
robots.txt, for example, has a non-binding, purely advisory character as well and Common Crawl [0] (also used for training GPT-3) publishes a dataset that by definition contains GPL'ed code as well, no matter where it's hosted. So is that off-limits now, too?
For instance, a parrot doesn't learn to speak, it learns to imitate speech.
We just need to bring the music industry into this!
For example: Let’s train a network on Beatles music to generate new Beatles songs. I’m pretty sure music lawyers will find a way to prove that the trained network is violating the label’s copyright, as they always manage to do that.
And then we just need to use the precedent and argue that music is the same thing as code.
Does Copilot infringe Google's patent(s) on the Transformer architecture? If so, then Google could potentially sue them for royalties, at least.
Further, couldn't this Copilot thing backfire for Github because customer trust is more valuable than AI training data right now? If folks don't feel they can trust Github, seems like they could move their work to other version control systems like Gitlab or Bitbucket...
Same today with the licensing system.
The people making the machine that learned (and recites) beatles songs aren't infringing though (most likely). It's those that use the machine to create and distribute the new works that are.
Same here. No one will be able to say that Copilot itself is a "derived work" or somehow uses the code in a way similar to a computer program (Although such claims have already been made - I highly doubt that's the case). But those that produce a whole file full of GPL code verbatim (Which will be rare, but WILL happen), are at risk of violating the license terms if they distribute it under the wrong license.
(Maybe bring oracle into this :D)
Asking seriously. It's really unclear to me where law and/or ethics put the boundaries. Also, I'd guess it's probably country dependent.
GitHub scanning billions of code files to build commercial software is different than you learning at human pace, even if they're both "learning" and in the end they both produce commercial software.
The human activity most like training an ML system is memorizing a text by reciting from memory, checking against the original, adjusting, and repeating until there are acceptably few mistakes.
And if a human did so for thousands of texts then publicly repeated those texts, they would be violating copyright too.
The main problem that has been the topic is a simpler one - about the produced work. If you exactly reproduce someone's existing code (doesn't matter if you copy by flipping bits one by one or which technology you use), isn't it a copyright violation?
I'm kind of imagining a Rube Goldberg machine that spells out the quake invsqrt function in the sand, now...
With programming, there's the further complication what constitutes a work. But quakes invsqrt certainly qualifies, just like that one function from the Oracle vs Google case.
sometimes? it's enough of an issue that companies explicitly avoid it by having two teams.
If the ML model essentially is just a very sophisticated search and helps you choose what to copy and helps you modify it to fit your code then it's 100% infringement. If it is actually writing code then maybe not.
Does the code even matter at all? If I start with a copy of some existing code, how much do I have to change it to no longer constitute a license violation? Can I ever reach this point or would the violation already be in the fact that I started with a copy no matter what happens later? Does intention matter? Can I unintentionally violate a license?
But I think we don't have to do all the work, I am pretty sure this has already been considered at length by philosophers and jurist.
That depends, if you end up writing copies of the code you've studied then yes. You are on thin ice. Plagarization is definitely something that you can do with computer code. There have been several high profile cases around this in arts. As far as I can see it usually ends up being a question about how much of the work is similar, how similar it is, and how unique what was similar is. And added wrinkle in programing is that some things can be done in only one way, or at least any reasonable programmer will do it in only one way. So for example a swap(var1, var2) function can usually only be done in one way, and therefor you would not get in trouble if your and someone else swap function are the same.
I've been following the discussion about Copilot, and one issue that comes up again and again is that people seem to think that since Copilot is new, the law will treat it, and the code it writes differently, than what it would you or a copy machine. I think that is naive, my impression is that courts care more about what you did, not how you did it, and if you think Copilot can be used to do an end run around the law. Prepare to be disappointed.
So if Copilot memorize code and spits out copies of that code, then it is at best skating on thin ice, or at worst doing a license violation. If the code it is copying is unique, then it definitely is heading into problematic territory. I'm fairly sure sure someone in legal at Github is very unhappy about the quake fast inverse square root function.
As for fronted/open source etc... sure, if you don't care about copyright and licensing, use it.
Well, there's also the xor way to be pedantic :)
var1 = var1 ^ var2
var2 = var2 ^ var1
var1 = var1 ^ var2
But yeah, not too much wiggle room there. var1 += var2;
var2 = var1 - var2;
var1 -= var2;
And another: var1 ^= var2 ^= var1 ^= var2;
Assembly even has an instruction for it: xchg eax, ecxMost humans don't do so unintentionally though.
I'm not the biggest fan of copyright law as currently written, but I wouldn't say that MS's desire to file off the serial numbers on every piece of public code for their own profit is a good impetus to rewrite the law.
Yes, a bit? It depend. Using such things for advertisement would likely cause anger if people start to recognize images of the training set the AI was trained on.
As someone who has taught students in ICT a quick rule of thumb was that I picked a piece of text that I suspected, wrapped it in doublequotes and put it into a search engine.
9/10 times - possibly more - of the times I had that feeling it was true. 17 year olds don't write like seasoned reporters most of the time.
Obviously there needs to be some independent tought in there as well, but for teenagers I put the line at not copying verbatim, and to cite sources.
As we've seen demonstrated again and again copilot breaks both my minimum standard rules for teenagers: it copies verbatim and it doesn't cite sources.
I say that is pretty bad.
If the system had actually learned the structure and applied what it had learned to recreate the same it would be a whole different story.
But in this case it is obvious that the AI isn't writing the code - at least not all the time, it is instead choosing what to copy - verbatim.
GPT-3 is no more 'intelligent' in the human sense than it is 'learning' in the human sense.
Do you mean that the terms, algorithms, concepts, and applications found in the field labelled "Artificial Intelligence" should not be called as such?
I have a feeling you are simply playing a semantic game, though, in which case we are likely to talk past each other.
Edit: I suspect you may be conflating artificial general intelligence[0] with AI
[0]: https://en.wikipedia.org/wiki/Artificial_general_intelligenc...
I still don't see any problem with that. If it's larger sections (e.g. entire NON-TRIVIAL function bodies), those can be filtered or correctly attributed after inference. So that's just a technicality.
Smaller snippets and trivial or mechanical implementations (generated code, API calls, API access patterns) aren't subject to any kind of protection anyway.
int main(int argc, char* argv[]) {
Lines like that hold no intellectual value and can be found in GPL'ed code. It can be argued that that's a verbatim reproduction, yet it's not a violation of any kind in any reasonable context.Where do you draw the line and how would you be able to - automatically even! - decide what does and does not represent a significant verbatim reproduction?
Today copilot does what it does.
I've never heard Microsoft defend anyone running afoul of some of their licensing details with "they can fix it later, it is just a technicality".
I think this should go both ways? No?
> Smaller snippets and trivial or mechanical implementations (generated code, API calls, API access patterns) aren't subject to any kind of protection anyway.
int main(int argc, char* argv[]) {
> Lines like that hold no intellectual value and can be found in GPL'ed code. It can be argued that that's a verbatim reproduction, yet it's not a violation of any kind in any reasonable context.Totally agree. Edit: otherwise we'd all be in serious trouble.
> Where do you draw the line and how would you be able to - automatically even! - decide what does and does not represent a significant verbatim reproduction?
I am not a lawyer but I guess many can agree that somewhere before copying functions verbatim, comments literally copied as well for good measure, somewhere before that point there is a line.
On the other hand: if there was significant evidence that the AI was doing creative work, not just (or partially just) copying then I think I would say it was OK even if it arrived at that knowledge by reading copyrighted works.
Edit: how could we know if it was doing creative work? First because it wouldn't be literally the same. Literal copying is liter copying regardless of if it is done using Xerox, paid writers, infinite monkeys om infinite typewriters, "AI" or actual strong AI.
After that it becomes a bit more fuzzy as more possibilities open up:
- for student works I look at how well adapted it is to the question at hand: a good answer from Stackoverflow, attributed properly and adapted to the coding style of the code base? Absolutely OK. Copying together a bunch of stuff from examples in the frameworks website? Fine. Reading through all the docs and look at how a number of high profile projects have done it in their open source solution, updating the README.md with info on why this solution was chosen? Now you are looking for a top grade in my class.
(of course IBM will probably not want you to work on their compiler though if you admit that you've studied OpenJDKs, or so I have heard.)
It's also not a commercially released product yet, but a technical preview, so uncovering and addressing issues like that is exactly what pre-release versions are for.
I'd say it succeeded greatly in sparking a discussion about these issues.
... will you defend it just because I claim it is a tech preview?
That's a straw man argument and you know it.
Code snippets are in no way shape or form comparable to entire software products and CoPilot neither installs anything nor is its intention to knowingly violate licences or copyright law.
Disingenuous straw manning like this doesn't help the discussion and only serves to distract from actual issues.
It is absolutely not in my opinion and that particular idea did not cross my mind at all so the idea that I knew it is patently double false.
But let me try to be constructive here and be even more precise:
Would it be OK if I launched a tech preview of my AI poem writer companion that would copy lines but also complete stanzas from famous poets, rock bands and singer-songwriters?
Yes it would be if it only happened ~0.1% of the time and if quoting verbatim wasn't the intended function of the system but merely a side-effect. In fact, that's what artists sometimes do deliberately.
It's what happens with other GANs as well and all that needs to happen is to educate users about the possibility of this. As long as you don't take ownership of the output produced by your AI (and neither do Microsoft), it's at the discretion of the user what they use the generated content for and in which context.
It has been demonstrated that training data can be extracted from any large NLP model [0] so this wouldn't come as a surprise either.
[0] https://arxiv.org/abs/2012.07805
https://towardsdatascience.com/openai-gpt-leaking-your-data-...
Idxs[i] += (Imm >> ((i * HalfLaneElts) % 8)) & ((1 << HalfLaneElts) - 1);
double r2 = fma(u*v, fma(v, fma(v, fma(v, ca_4, ca_3), ca_2), ca_1), -correction);
seed ^= hasher(v) + 0x9e3779b9 + (seed << 6) + (seed >> 2);
qint32 val = d + (((fromX << 8) + 0xff - lx) * dd >> 8);
even if it's one line, it likely took some non-negligible thinking time from the programmerMathematics and physics equations are not copyrightable.
There are many good answers from the legal side. I would also attack this side: the way human beings learn is entirely different from the way ML models are trained. We don't do gradient descent to find the slope of data points and find the most likely next bit of code.
We humans create rational models of the code and of the world, and use deduction from those models to create code. This is extremely visible in the way we can explain the reason behind our code, and in the way we are aware of the difference between copying code we've seen before vs writing new code. It's also visible in that we can be told rules and produce code that obeys those rules that doesn't resemble any code ever written before.
The difference is also easily quantifiable: humans learn to program after seeing vastly fewer code examples than Co-pilot needed, and we are much better at it.
One day, we will design an AI that does learn more similarly to how humans learn, and that day your question will be far more interesting. But we are far from such problems.
"You" are just not paying attention to it in that moment.
Does the copilot still learn from new repos? Can I post github enterprise code publicly to let it learn from it?
Serious answers only please
Of course using code generated by Copilot from those would still be illegal.
See also : Napster (and other p2p), the bitcoin blockchain allegedly containing illegal numbers...
That said, if enough people are bitten by this, i'm not sure what happens -- does anyone know of a relevant case. One somewhat relevant case that caused mass pain was the SCO Linux Dispute
The official answer from Github that they take all input on purpose doesn't play in their favor.
For example, if I purchase certain corporate Linux licenses, i'm protected against being sued if something in the distribution ends up having misappropriated code.
Check out the SCO Linux Dispute for how bad things can get for corporations: https://en.wikipedia.org/wiki/SCO%E2%80%93Linux_disputes
The answer to what will happen when companies are bitten, is that there will be series of lawsuits involving various parties (including GitHub), dragging on for a while and costing a fortune. The court will decide everything in the end (who's responsible for what, who cover the fees, who own the IP, etc...).
The SCO case was rather frivolous, I don't think there is much to take from it, except that if a US company is determined to sue and they have a billion dollar to go on forever, there's nothing stopping them to and it's a lot of troubles.
Which is relevant I suppose. It's only a matter of time until there's a major case putting GitHub Copilot in the spotlight and an aggressive company with deep pocket (think Oracle and the likes). We will certainly be reading about it everywhere the day it starts.
whereas for open source it's a disaster
Finally, it beat the "cancer that attaches itself in an intellectual property sense to everything it touches" after all those years, with its own tools!
Now it's safe to touch.
Remember that Amazon won off the back of open source. Now all the open source servers and databases are Amazon products.
https://github.com/PubDom/Windows-Server-2003/blob/master/co...
#ifdef _MAC
# include <string.h>
# pragma segment ClipBrd
// On the Macintosh, the clipboard is always open. We define a macro for
// OpenClipboard that returns TRUE. When this is used for error checking,
// the compiler should optimize away any code that depends on testing this,
// since it is a constant.
# define OpenClipboard(x) TRUE
// On the Macintosh, the clipboard is not closed. To make all code behave
// as if everything is OK, we define a macro for CloseClipboard that returns
// TRUE. When this is used for error checking, the compiler should optimize
// away any code that depends on testing this, since it is a constant.
# define CloseClipboard() TRUE
#endif // _MAC
Just the kind of trick co-pilot should help us with?This is exactly what I meant.
>none of the things you enumerate are hosted on GH.
Plenty of them on GH, if not src then magnet links
What exactly is the difference between a machine learning patterns and techniques from looking at code and people doing it?
Is every programer who ever gazed at GPL'ed code guilty of plagiarism and licensing violations because everything they write has to be considered derivative work now?
That's a fair point. ML models don't seem memorise all the code they've seen either, it seems. Plus while the argument of human limitations applies to the vast majority of people, what about those with eidetic memory?
> what happens when Copilot produces certain code verbatim?
There are several options: suppress the result, annotate with a proper reference or mark the snipped as GPL'ed.
There are technical solutions to this question, but it's also important to ask to which degree this is necessary.
Is a search engine that returns code snippets regardless of license also a tool that needs to be discussed the same way? After all, code samples from StackOverflow or RosettaCode are copied on a regular basis and not every example provides a proper reference as to where it's been taken from.
So maybe a hint like "may contain results based on GPL'ed code" suffices? I don't know, but that's a question best deferred to software copyright law experts.
This idea that a snippet of a code is a work seems crazy to me. I thought we went through this with SCO already.
https://stackoverflow.com/help/licensing
I'm guessing most uses of stack overflow snippets are violating the license (no attribution, no share alike of the "remix" - which would probably be the entire program).
ShareAlike — If you remix, transform, or build upon the material, you must distribute your contributions under the same license as the original.
The idea that programmers taking snippets from stackexchange or co-pilot etc meaning they have a derivative work seems like total insanity.
We need the legal right to do things like host Your Content, publish it, and share it. You grant us and our legal successors the right to store, archive, parse, and display Your Content, and make incidental copies, as necessary to provide the Service, including improving the Service over time. This license includes the right to do things like copy it to our database and make backups; show it to you and other users; parse it into a search index or otherwise analyze it on our servers; share it with other users; and perform it, in case Your Content is something like music or video.
This license does not grant GitHub the right to sell Your Content. It also does not grant GitHub the right to otherwise distribute or use Your Content outside of our provision of the Service, except that as part of the right to archive Your Content, GitHub may permit our partners to store and archive Your Content in public repositories in connection with the GitHub Arctic Code Vault and GitHub Archive Program.
So it wouldn't include just any AI copy pasters. Only the ones that are provided by GitHub.
I could previously mirror GPL code, because the GPL granted me the rights I need to grant GitHub as part of their ToS; but if they change their ToS, or if the meaning is changed by them adding vastly different features to their Service, this becomes a problem.
Whether you're allowed to upload GPL code to GitHub or not depends on whatever their Service is at the moment, since the terms say you grant them all the rights "necessary to provide the Service".
[1] I’m not a lawyer.
[2] Still not a lawyer.
* most people here are unhappy
* most laywers will say it's fine (it very probably passed MS ones)
I can understand that. Copyright was not created with AI/ML in mind, even as a random stray thought. Those were not even words at the time.
So the question is: If we change the law and require trained algorithms to only work on licenses that permit this, and to output the "minimum common license" somehow, what are the repercussion on other application of copyright?
Because the consensus here seems to be that this looks a lot like a de-licensor with extra steps
Copilot isn't a search engine any more than any other language model is. It can sometimes output data from the training set verbatim as most AI models do from time to time, but that is the exception not the rule.
Whether modern autoregressive language models can be called "inteligent" is debatable, but they're certainly far beyond what you'd get from a simple search engine.
Since neural networks are pattern matching based on the training input it is a derivative work of the training set. It says it right in first thing that comes up in auto regressive language models use the training input plus context to predict what the next word would be.
Now here where the fun begins if they try this in court. If you claim it’s generating new work then who owns the copyright? You may not realize how big of a deal this is but there was a court case you can lookup where a monkey took a selfy and the person who camera the monkey used tried to claim copyright and lost.
The more questionable line is if someone happens to inadvertently reproduce entire paragraphs of Twilight: Breaking Dawn word-for-word using GPT-3 and then sells it, that might be a violation even if they didn't realize they were doing it.
Copilot is the same thing. Creating a product that makes suggestions that it learned from reading other people's work is fine. Now if you write code using Copilot and happen to reproduce some part of glibc down to the variable names, and don't release it under GPL, you might be in trouble. But Copilot won't be.
Another example is the photo generation ML algorithms that exist. They generate photos of random "people" (imaginary AI-generated people) by using actual photos of real people. If one eye or nose is verbatim copied from the actual photo to the generated photo, is the entire output now illegal or plagiarism? One might argue it's just an eye, the rest of the picture is completely different, the original photographer doesn't need to grant permission for that use.
Any analogies we make with this, be it text generation, image generation, even video generation, seems like it falls under the same conclusion: so far we've thought all of this was perfectly fine. I don't see why code-generation is any different. A function is just a tiny part of a project. It's not necessarily more important than the composition of a photograph, or a phrase in a book. We as programmers assign meaning to it, we know it takes time to craft it and it might be unique, but likewise a novelist may have spent weeks on a specific 10 word phrase that was reproduced verbatim, in a text of 500 pages.
The more I look at this the more it seems copyright, and IP law in general, is the main problem. Copyleft and OS licenses wouldn't be needed if it wasn't for the aggressive nature of IP law. I don't see the need to defend far more strict interpretations of it because it has now touched our field.
Let's use this thought experiment: Imagine that Github's Copilot was just a massive array of all the lines of code from every github project, with some (magical automated whatever) tagging and indexing on each function, and a search engine on top of that.
Now imagine that copilot simply finds the closest search result, and then when you press a button, it inserts the line from the array, and press it again and you get the next line, etc.
Now hopefully nobody here thinks such a system would fulfil either the spirit or the law of any half-restrictive license. Yet that is a perfectly valid implementation of Copilot's aim - and it sounds like it's not that far from what actually happens, maybe with a bit of variable name munging.
So my question is this: If you could build a line between the system I describe above and the system of human learning, where a human learns the patterns and can genuinely produce novel structures and patterns and even programming languages that it has never seen before.
At what point along that line would you say that Copilot is close enough to human to not be violating licenses that require attribution?
Devs were hoping for stars and network effects rather than listening to those of us feeling uncomfortable taking all traffic to gh. Something like Copilot or even a coding bot was predicted two years ago already.
(The question of whether this is sufficiently transformative to count as fair use is still wide open)
I'd suggest it makes it more interesting. If it's self hosted, then the hoster can choose to impose restrictions on server aceess, including no automated scraping, rather than trying to impose licensing on the code itself.
But I think more and more companies, particularly those in highly regulated industries, are deciding that the benefit of controlling the data — access, security, privacy, and understanding who, exactly, it’s being shared with — outweighs the risks of someone else having that control.
I definitely understand why people pick a license that disallows use someone doesn't agree with. Imagine baking cookies for your friends, and one of them reselling them. The material effect is the same to you, you gave away your cookies, but sometimes you make/do something for a certain group of people and not for other to make a profit of your work.
But a great deal of the value that's come from open source generally has been that open source licenses haven't imposed the sort of usage-based restrictions (e.g. free for educational use only) that were fairly common in the PC world.
And, to your example, in the case of software the incremental copy that your friend sold cost you absolutely nothing. So it comes down to a purely emotional response to someone else making money off something you made.
Exactly, as I said, the material situation is the same. But we all are emotional beings, you would do certain things for your family you wouldn't for strangers. I don't think this case is any different.
I personally don't work for free for a company, but I do charity work for free. Working for a company in the time I work for a charity would "cost me absolutely nothing" if I already spend the time anyway, but everyone understands the difference.
Edit: commercializing of the derived work is one explicit consideration used by US law in making a fair use determination. That said, even if it weren’t commercialized it may still be infringement and I believe it is.
We already have examples of copilot blatantly plagiarizing code
https://www.theverge.com/2018/2/8/16992626/apple-github-dmca...
It isn't any kind of copyright infringement. The AI is not copying and pasting code that is has found, it is rewriting the code from scratch on its own.
We keep trying to take old ways and meld them to the internet, and its just not appropriate and it doesn't work.
https://twitter.com/kylpeacock/status/1410749018183933952
https://twitter.com/mitsuhiko/status/1410886329924194309
https://twitter.com/pkell7/status/1411058236321681414/photo/...
On the plus side, a large body of work effectively becomes public domain. On the negative side, copyleft licenses lose their teeth. You probably see more power shift to those with big budgets. You probably see fewer things made source available, because you either have the public license or the private license now. This feels like a bad path but I'm not convinced the end result isn't better still.
I can setup my drone to detect me and attempt to crash into me. AI would be quite poor, probably would attempt to crash at any human. Would it be my fault it didn't crash into me and someone lost eyes?
Can I setup torrent box that automatically downloads and seeds all detected links from public trackers? Would I be responsible for it?
Not a lawyer or even particularly well informed
edit: I am reminded of the monkey selfie, in which it was ruled that a non-human cannot create copyrightable works. https://en.wikipedia.org/wiki/Monkey_selfie_copyright_disput...
1) Can you be liable for violating copyright if you have never seen the work?
2) Can a non-human be held accountable for violating copyright?
3) Can github be held liable for an end user using their tool to violate copyright?
https://en.wikipedia.org/wiki/Substantial_similarity
wikipedia states: Generally, copying cannot be proven without some evidence of access; however, in the seminal case on striking similarity, Arnstein v. Porter, the Second Circuit stated that even absent a finding of access, copying can be established when the similarities between two works are "so striking as to preclude the possibility that the plaintiff and defendant independently arrived at the same result."
This is a different situation in which exact replication can be reasonably occurred without access to the original.
Secondly, can you actually claim Github has violated copyright if it doesn't have any claims to the work in question?
I think it's totally plausible that they win this in the long run.
2,3) Seems pretty settled at this point, look at the cases around the VCR and copy machine. In general the one using the machine is liable. The creator of the machine can be held liable if there aren't substantial non infringing uses.
2) Someone using a copy machine is knowingly copying a specific work.
Many people on HN assert this based on the Authors Guild vs. Google case, but it's quite important to keep in mind that that case was about Google creating a search algorithm, which is not generating "new" output.
We are talking about a very different kind of system here and in many other cases. Claiming the Authors Guild case sets precedent for these very different systems seems unbased to me.
This is a very bold assumption, one that I assume will not hold in the court of law in all cases. I think the nuanced question is: to train a model that does what, exactly.
Let's say distributing meth recipes is illegal[1], can one legally side-step that by training a model that spits out the meth recipe instead? No court will bother with the distinction, causation is well-trod ground.
1. As an example - not sure if its illegal. You may replace with classified nuclear weapon schematics if you like.
I think most people are more concerned about whether the user of Copilot would be liable for using copyrighted code generated by Copilot.
If one of the three largest record labels uses their own catalog to train an AI, copyright seems less important to discuss. I suspect the discussion would be a bit different if a company scraped youtube and used that as a training set for AI music and successfully sold it.
Were those products sold to help people write commercial pop music faster? If not, I don't think your point is valid.
Best case scenario, they explained in advance on the GH blog they're going to be doing some work on ML and coding, and they'd like people to opt into their profile being read via a flag setting/or put a file in the repo that gives permission like robots.txt? Second best case scenario, same as first but opt out vs opt in, and least ideal would be something like not doing the first two, however, when they announced it, explained in detail how the model was trained and what was used, why, and when- kinda thing?
Is that generally about right, or..?
Not really, consider for example repositories mirrored to Github.
It seems unclear who has the rights to grant this permission anyways (with free software licenses). Probably the copyright holder? Who that is might also be complicated.
This is the complicated bit: All open-source licenses grant you permission to redistribute the code (usually with stipulations like having to include the license), so you are almost always allowed to upload the code to Github.
What it doesn't mean however is that you're the copyright holder of that code, you're merely redistributing work that somebody else has ownership of.
So who gets to decide what Github is allowed to do with it?
I expect this will end up in courts and we won't get a definite answer before that.
(Not sure for the cases where there is no license and therefore normal copyright applies, but AFAIK this isn't the case for any code on Github, which automatically gets an open source licence ?
EDIT : Code in public repositories seems to be "forkable" on Github itself but not copyable (to elsewhere). That's some nasty walled garden stuff right there, I wonder how legal that ToS is ? I could see how this could make them to incentivize people to stop using other licenses on Github, to not have to deal with this license mess... EEE yet again ?)
> Version 3, 29 June 2007
> Copyright © 2007 Free Software Foundation, Inc. <https://fsf.org/>
> Everyone is permitted to copy and distribute verbatim copies of this license document, but changing it is not allowed.
Like, don't use it if you're worried about violating licenses, but I don't see how Microsoft could get in trouble for the tool. It doesn't write and publish code by itself.
In short, github gets to make the license violator bot and push the violations off onto the small fry who actually use it? No thanks.
Forbidding the training of AI on public code would definitely be a step too far though.
Edit: I'd also like if they provide a tool for checking if your code matches copyrighted code too close so you can confirm if you are violating anything or not when you use copilot.
My simplistic view is that the following is legally equivalent:
input -> ai network -> output
input -> huffman coding -> output
So, whilst:
* compressing and decompressing a copyright work is permissible;
* output and weights are deterministic transformations of the inputs;
thus:
* not eligible for copyright (lacking creativity); and
* are derivative works of the inputs;
copyrighted input -> compiler -> copyrighted output
So portions of a derivative work are covered by the original copyright, and other portions may be under a distinct copyright as a derivative work, and several copyrights may apply to a work as a whole.
In the case of a Huffman transform, the transformed work does not meet the "creativity" requirements to be eligible for copyright, over that of the original works.
That may be true but I fail to see how any process that produces the same content that was input into it somehow strips the license. If the generated code is novel, then there is no copyright and it is just the output of the tool. If the code is a copy, but non-creative (example a trivial function) then it isn't covered by copyright in the source anyways, so the output is not protected by copyright either. However if the output is a copy and creative I don't think it matters how complicated your copying process was. What matters is that the code was copied and you need to obey copyright.
Again, I don't think that novel code generated from being trained on copyrighted code is the problem. I think it is just the verbatim (or minimally transformed) copying that is the issue.
We already have the same fuzzy line for writing. Am I forbidden from ever reading other author's books because I might accidentally "generate exact copies" of some of the sentences? Clearly not, that is how people learn a language. Does that mean I am allowed to copy the whole book? Also clearly not.
Where do you draw the line? Somewhere.
There's absolutely nothing different whether the creator is ML or a human.
Generally, if you train an ML network to generate an almost exact copy of a thousand lines, it's obviously not fair use. If it's five simple lines, it obviously is fair use. If it's somewhere in between, there are a lot of different factors that need to be weighed in a fair use decision, which you can easily look up.
Putting the (imho) big licensing problems aside, what about the software patents?
Apache and GPL have patent protection clauses.
Does this mean that anyone using copilot might somehow get code that implements something patented, but protected by license, except they did not get proper permission through the Apache/GPL license?
...I kind of hate myself for saying this, but... Patent trolls to the rescue?
But what if the system memorizes entire functions? What if a human does so? What if you change all the variable names? What if you rearrange the control flow a bit? What if you just change the spacing? What if two humans write the exact same code independently? Is every for loop with i from 0 to n a license violation?
I am not picking any side, but the problem is certainly much more nuanced then either side of the argument wants to paint it.
Copilot will replicate entire functions, including comments, from licensed code
I think it is important to point out that not all Copilot output is on the plagiarizing side of the spectrum. However it does on occasion produce plagiarized code. And most importantly there is no indication when this occurs.
In general using the idea is fine, whether it is AI or human written. I think the major concern here is when the code is copied verbatim, or near verbatim. (AKA the produced code is not "transformative" upon the original)
> But what if the system memorizes entire functions? What if a human does so?
In both of these cases I believe it would be a copyright concern. It is not strictly defined, and it depends on the complexity of the function. If you memorized (|a| a + 1) I doubt any court would call that copying a creative work. But if you memorized the quake fast inverse square root it is likely protected under copyright, even if you changed the variable names and formatting.
It seems clear to me that GitHub Copilot is capable of producing code that is copyrighted and needs to be used according to the copyright owner's license. Worse still, it doesn't appear of capable of knowing when it is doing that, and what the source is.
I hope that even if it wouldn't work, it puts enough doubt in companies' minds that they wouldn't want to use a model trained by code under those licenses.
[1]: https://gavinhoward.com/2021/07/poisoning-github-copilot-and...
1. Do you need a license to use materials for training, or to use the output model?
2. If so, does the code's license allow this?
GitHub is claiming 'no' for #1, that they do not need any sort of license to the training materials. This is reasonably standard in ML; it's also how GPT-3 etc were trained.
Now, whether a court will agree with their interpretation is an interesting question, but if they are correct then #2 doesn't come into play.
"Dear copilot, I'm writing a Unix-like operating system...."
You aren’t allowed to sell someone’s likeness without their permission. You don’t need an AI for this if you create a portrait of Clooney and sell it or make any use that isn’t covered by fair use he can sue you.
Depending on the composition of the picture for example if Clooney is naked and say Putin is riding in the “bitch seat” of the saddle then you also are quite likely be open for a libel suit as well.
>For example, in Hustler Magazine v. Falwell (1988), Chief Justice William H. Rehnquist, writing for a unanimous court, stated that a parody depicting the Reverend Jerry Falwell as a drunken, incestuous son could not be defamation since it was an obvious parody, not intended as a statement of fact. To find otherwise, the Court said, was to endanger First Amendment protection for every artist, political cartoonist, and comedian who used satire to criticize public figures.
The US system isn’t the only one on the planet you know, the UK still has political cartoonists despite a very different definition for what defamation is which the example above can fall under.
https://fossbytes.com/github-copilot-generating-functional-a...
Most of the comments are waxing about philosophical possibilities of copilot copying GPL.
Reality is clear in the case, it's copy pasting thousands of characters of GPL code with no modifications. Copyright violation clear as day
Also, given that the author of this tweet called me a "bootlicker" last year in response to a somewhat lengthy nuanced post about GitHub, I'm gonna go out on a limb and say that they're not all that interested in a meaningful conversation on this in the first place but are rather on a quest to "prove" GitHub is evil.
>oh my gods. they literally have no shame about this.
Then continues with
>it's official, obeying copyright is only for the plebs and proles, rich people and big companies can do whatever they want
and
> GitHub, and by extension @Microsoft , knows that copyright is essentially worthless for individuals and small community projects. THAT is why they're all buddy-buddy with free software types; they never intended to respect our rights in the first place
At any rate, it's not even clear to me if me publishing code written with copilot (or even with a random tool that will wget from github) puts the blame on the toolmaker or on me. This post, however, doesn't attempt to look at that but uses language that paints GH/MS as doing something illegal (and evil) that others wouldn't even get away with but not caring about it.
A non rich individual has basically zero chance of challenging GitHub on these blatant violations, and they know it.
> At any rate, it's not even clear to me if me publishing code written with copilot (or even with a random tool that will wget from github) puts the blame on the toolmaker or on me.
It really depends on the license, which GitHub apparently doesn't care about at all.
The legality lies on what the user does with the code.
I'm sure there's a "for i in range(0, n):" somewhere in a GPL repo, and yet having that in my code doesn't make it GPL.
It’s literally being used to stifle research as we speak, but for some completely insane reason we are protecting a handful of publishers as a cartel…
It really is so simple: “don’t be a bigoted fascist”, but just have a glance at the fascists decrying my stance as a degenerate liberal, in the answering comments
The US seems well enough informed. As mentioned in the following report "AI tools are diffusing broadly and rapidly" and "AI is the quintessential “dual use” technology—it can be used for civilian and military purposes.".
https://www.nscai.gov/wp-content/uploads/2021/03/Full-Report...
I'm fully expecting that if I begin a story and put it on my blog or on github, and if I go away for a couple years, I'll see it completed for me when I return. I can use foresight to my advantage or I can pretend like it's still the 1990s as if placing some text at the top of the code I exposed publicly is going to prevent people from training on it.
One thing for sure though, I don't think a large company such as Microsoft should be profiting from training their language model on open-source code.
The best way to release Copilot in my opinion would be to make the entire thing open source and have separate models, even a private paid-for model so long it's trained on their own code.
An open source model trained on code for specific licenses sounds fine, but then the model should also follow that same license as the code it was trained on.
There's just something deeply unsettling about having a computer complete your thoughts for you without being able to question how or why.
Same thinking probably applies to GitHub Copilot and copyright
So by your logic ABC, CBS, Fox, and NBC have all been plagiarizing and violating copyright for doing so? I’m not sure if there’s been a legal challenge/precedent set in that case yet, but that seems like a more apples to apples comparison than the Google Books metaphor being used.
Disclosure: I work at GitHub but am not involved in CoPilot
I think this will end up with a large class action lawsuit for sure, tho I really think it’s a toss up as to who would win it. This conversation was bound to happen eventually and we’re in uncharted territory here.
I think it’s going to hinge on whether machine learning is considered equivalent in abstraction to human learning, which will be quite an interesting legal, technological, and philosophical precedent to set if it goes that way.
Why would they distinguish between licenses if there's no legal need to?
Licenses are only restrictions on top of fair use. Licenses can't restrict fair use.
It would be interesting if someone takes them to court and a judge definitively rules on fair use in this particular case. Or I don't know if there's enough precedent here that the case would never even make it to trial. But with a team of top-paid Microsoft lawyers that gave this the green light, I'm pretty sure they're quite confident of the legality of it.
The model is said to spit out code verbatim 0.1% of the time, a low number, but if copilot is used a lot, it means you are going to find a lot of copied code in people's projects, and these project owners may be breaching copyright. I don't think "but, copilot..." will be an excuse.
Here is a (probably unrealistic) scenario illustrating it. I am playing a copyright troll here:
- Release plenty of generic code and put it on GitHub under a restrictive license
- Have the copilot bot scan it
- wait some time
- scan public codebases and do an exact match for my code
- sue project owner that contain my code
I see the use of copilot more of a minefield for me than as a liability for Microsoft.
The sticky part may be: does GitHub T/C overrule these licenses?
It might have an interesting feedback effect that some licenses which are more popular would presumably have better Copilot recommendations, which would produce better and thus more popular code for those licenses. Although maybe this happens already.
[1]: https://gavinhoward.com/2021/07/poisoning-github-copilot-and...
Edit: done. They are under the CC-BY-ND license now.
Copyright is broad, licenses are minimal. This must be the case otherwise they would not be very effective at protecting the work of creators. There is no explicit allowance for what GitHub is doing in most licenses so they do not have general permission to do so.
What my licenses are supposed to do is sow even more doubt in companies' minds about models trained on my code.
I don't see how they can keep this clause, and then have a service that recites/redistributes code, based on a model that has already ingested said code.
> This license does not grant GitHub the right to sell Your Content. It also does not grant GitHub the right to otherwise distribute or use Your Content outside of our provision of the Service, except that as part of the right to archive Your Content, GitHub may permit our partners to store and archive Your Content in public repositories in connection with the GitHub Arctic Code Vault and GitHub Archive Program. [1]
Copilot is distributed verbatim code when it regurgitates, which seems a pretty clear violation of this clause. (If it wasn't regurgitating, they'd have caselaw for fair use. But... It is.)
[0] https://docs.github.com/en/github/site-policy/github-terms-o...
[1] https://docs.github.com/en/github/site-policy/github-terms-o...
If you don't want others to use your code then the solution is very simple. Keep it on a secure private server and don't publicly release it.
Currently, if you claim to be a copyright owner GitHub can respond to a DMCA takedown by removing the repository. This might require them to retrain the entire model.
One option for GitHub might be to maintain a blocklist of various code snippets, and if there is a substring match, just don't make the suggestion.
I warned against hosting source code on GitHub and going all in on GitHub Actions, mainly for them being unreliable for the past year. [1] (They go down every month). Now Copilot has gone and trained on every single public repo on GitHub as admitted right in this post, regardless of the copyright.
Maybe for organisations with serious projects, perhaps now's the time to leave GitHub and self-host your own somewhere else?
And in the end, the output in itself is not really an issue. It is just a machine outputting random lines it encountered on internet.
The problem is from the user side: Ok, you got random lines from random places. If you do nothing about it, then no issue. But if you try to use, publish, sell the code, then you are in deep shit. But somehow it's your fault.
For GitHub, the problem is more to be sued by "customers" that assumed that the generated code was safe to use when it is not the case.
And, as a general comment, I think that this case is very illustrative about the misconceptions about AI and machine learning for the general public:
Here you can see that you don't really have an intelligent system that can learn and then create something new and innovative from scratch. But it is just a machine that copy code it already saw based on correlation with similarities in your current code.
Optimistically, Copilot could be a wake up call for thinking more deeply about how the winnings of data-dependent technologies (ultimately, dependent on the labor of people who do things like write open source code) are concentrated--or shared more broadly.
This longer blog post goes into more of a labor framing on the topic: https://www.psagroup.org/blogposts/101
(For the record, I certainly think Copilot could be very good for programmers in general and am not arguing against its existence -- just arguing that this is a high profile case study, useful for thinking about data-dependent tech in general)
No more "exploitation" of labor.
I think she's barking up the wrong tree here. If she's looking for organizations interested in eliminating fair use, RIAA, MPA, and AAP are more likely allies.
> How much of licensed code is scattered across private projects?
Whether or not copyright violations regularly occur is not (directly) relevant to whether or not it is illegal. People download copyrighted movies without licenses all the time and it still isn't legal.
On other hand said AI have no idea of who the code belongs to and it's able to reproduce it perfectly.
The above thread is a dupe of this discussion but with interesting discussions already in place before being marked as a dupe.
The training is not a copyright violation. That seems to be settled case law. Whether the verbatim copying as a result of that training is a copyright violation I think is less tested.
Let’s flip the domains. Say we had an ML algorithm that could auto generate news stories and it at some point (not all the time) copied verbatim a Wall Street Journal article and posted it to a blog. Copyright violation?
With copilot, we’re sometimes seeing “paragraphs” of source lines copying verbatim, so this analogy is not such a stretch.
I think we need to think about how much our sharing culture in programming has tinted our view of the legality of this enterprise.
I agree. It'd be a nice gesture to reach out to the creators of the training data, like is usual with web scrapers. But collecting and analyzing data publicly available on the web is ok.
>So the license doesn’t matter unless there’s some “can’t use this for ML training” license that I don’t know about (and doesn’t seem to be legal).
I disagree. While Copilot is, at heart, a ML model, the copyright trouble comes from its usage. It consumes copyright code (ok), analyzes copyright code (still ok), and then produces code which sometimes is a copy of copyright code (not ok). The only way it'd be ok is if Copilot followed all licensing requirements when it produced copies of other works.
Personally, I won't touch it for work until either Copilot abides by the licenses or there's robust case law.
I don’t think this is practical. And who notifies people of scraping content? I would’ve annoyed if I got spam from sites that scraped my content.
>I don’t think this is practical.
I don't like people ignoring things just because they're impractical for ML. That leads to crap like automated account banning without possiblity of talking to a living customer service representative.
Obviously most ( I think ) lawyers seems to be siding with Microsoft on fair use. But most owner of the code seems to think they are infringing on their work.
Then there is the international issue because one court cant decide for everyone else.
I think the issue is important enough I wonder if we could somehow crowdfund it for a court trial or something.
This simply feels like anti-Microsot people flocking to what they see as some exposed Microsoft flesh for social media biting.
Maybe this could spark a discussion to change the current rules that allow them to do this, but questioning the current legality to me is a waste of time.
As this boils down to legal arguments, are there any clauses (maybe disputed) in the ToS allowing github/MS usage of public repos for such purpose?
Would it even be legally possible to override a software license as a repo provider like "by using this service, you agree to..."?
What about deep learning-artwork trained on google searches?
We enter a new era…
https://en.m.wikipedia.org/wiki/Open_source_license_litigati...
I think my point still stands.
you can't go copying anything and everything just because nobody has told you that you can't. and I feel that's part of the purpose behind GPL. force a license on derivative code so that at least there's clear rights moving forwards.
I would hope GitHub (and Microsoft) did the legal work to cover this, and not just ploughed ahead with the plan to drown any legal challenges. From my perspective, they're doing the latter.
* An algorithm (or person) ingesting lots of code and then later spitting out that same input, does not free anyone from the copyrights of the input.
* An algorithm (or person) that ingests lots of code, finds commonalities, synthesizes that into something new, and produces something well beyond mere copying is producing something new, likely without any legal tie to the original.
Right now, it looks like most of what co-pilot does is closer to the latter, but sometimes it does some things that are closer to the former? I can't see any reason why they wouldn't be able to fix it to avoid regurgitating its input, however, with something like a bloom filter, so I expect a long-term there's a way to do it that falls entirely within fair use?
So now we know its ALL public repos ... how long until all the opponents of this tool have a giant repo full of syntactically correct code that employs terrible design patterns and is thoroughly obfuscated? I'm not going to waste my time on this personally but there are certainly those who will. Someone will invent a tool that perverts perfectly good code in the process and probably have a good laugh.
Personally, while I recognize some people might find it useful, I don't much care for it. No, I haven't tried it yet either. Ive never sampled escargot either and I know I don't care for it all the same. Maybe it's wonderful, I'll never know - but I do know that I simply don't like the idea of it. Call it an objection on General Principal if you like.
So remember kids, If you're not PAYING then you are the product.
Bottom line - private repos are cheap and you should use them rather than freebie public stuff.
Of course, theres a huge irony in that Github is also making the tool that enables the widespread plagarism....
> Short version: You own content you create, but you allow us certain rights to it, so that we can display and share the content you post. You still have control over your content, and responsibility for it, and the rights you grant us are limited to those we need to provide the service. We have the right to remove content or close Accounts if we need to.
[1] https://docs.github.com/en/github/site-policy/github-terms-o...
They have done an excellent job and succeeded in their goal.
Now, with copilot they are about to lose it all.
> regardless of license
Copilot is currently a technical preview. Github has already said they intend to detect verbatim code and notify the user and present the correct license. That'll be in the final release.
Don't use the technical preview for anything except demoing a cool concept. It's not ready for that yet because it will reproduce licensed code and not tell you.
Please provide a source for this.
Replace "GPL" with the most restrictive license that's on GitHub, but you get the point.
They're kinda shooting themselves in the foot, because this reduces the commercial potential of the tool to almost nothing.
I think this may be the crux of this whole kerfuffle.
If you're the author isn't it on you if you infringe?
If not then perhaps you and GitHub/Microsoft share authorship/culpability?
Who has the copyright to a piece of text generated by a tool? Or art generated by a model?
Imagine a scenario where you'd love to have access to a large number of my digital widgets, but they're expensive to make or buy, and a large number of them is really expensive. So you train an ML model on my things you can't afford to buy. It's still expensive, but that's a one time cost. Spend $5M training GPT-3, it's fine. Now you can sample from the space of my digital widgets. You have gotten a large number of widgets, just by throwing money at AWS. With money, you have converted my widgets into your widgets, and I'll never see a cent of it.
That's the issue. Content is expensive and it's still needed. Traditionally, I make content and if you want to benefit from my labor, you pay me. In the future, if you want to benefit from my labor, you pay AWS instead.
tl;dr The most significant equation for generative models is "$$$ + my stuff = your stuff"
A human learns by looking at all public code
A robot learns by looking at all public code
(Okay, I have some reservations to above comment, but for discussions sake, that’s what I’m going with)
"You either take part in open source or you don’t." I disagree. You can allow your software to be used and post the source code, but it is yours so you get some say in your intentions. Forking is what you're looking for. However, once you fork it, you still owe credit to those that did the heavy lifting before making whatever tweak it is you made and want to call it your own. There's nothing wrong with the original developers getting credit for the work they did. There's nothing wrong with the original devs willing to let other people use their work as long as it is used in the same spirit it was provided (FOSS). That also does not mean the original devs are wrong for wanting evilCorps that want to use their freesoftware to be included/distributed in their packages they sell and profit from to be restrictive.