I do not agree with Github's use of copyrighted code as training for Copilot
thelig.ht
thelig.ht
Copilot regurgitating Quake code, including sweary comments - https://news.ycombinator.com/item?id=27710287 - July 2021 (625 comments)
GitHub Copilot as open source code laundering? - https://news.ycombinator.com/item?id=27687450 - June 2021 (449 comments)
Also ongoing, and more or less a duplicate of this one:
GitHub scraped your code. And they plan to charge you - https://news.ycombinator.com/item?id=27724008 - July 2021 (148 comments)
Original thread:
GitHub Copilot - https://news.ycombinator.com/item?id=27676266 - June 2021 (1255 comments)
https://twitter.com/luis_in_brief/status/1410242882523459585...
And this is a longer article about how IP and AI interact:
https://ilr.law.uiowa.edu/print/volume-101-issue-2/copyright...
I am not a lawyer, but I am capable of summarizing the thoughts of lawyers, so my take is that in general, fair use allows AI to be trained on copyrighted material, and humans who use this AI are not responsible for minor copyright infringement that happens accidentally as a result. However, this has not been tested in court in detail, so the consensus could change, and if you were extremely risk-averse you might want to avoid Copilot.
A key quote from the second link:
Copyright has concluded that reading by robots doesn’t count. Infringement is for humans only; when computers do it, it’s fair use.
Personally, I think law should allow Copilot. As a human, I am allowed to read copyrighted code and learn from it. An AI should be allowed to do the same thing. And nobody cares if my ten-line "how to invert a binary tree" snippet is the same as someone else's. Nobody is really being hurt when a new tool makes it easier to copy little bits of code from the internet.
This would be interesting to test with AI and pop music.
If you’re doing it from an analog format you bought for your own use (format shifting), it is fair use.
And suddenly there is a world of robots composing, writing and painting for other robots. With us humans left out.
There should be a /s at the end, but legal world sometimes produces such convolutions. See, for example, the interpretation of the Commerce Clause in Gonzales v. Reich.
Quite the opposite. We all get a tiny bit better with good information like this. This is what the internet should be for, evolving, learning from past mistakes, information availability.
If the discussion was “I clicked this button and got someone’s entire chat platform” that would be different. Words and sentences aren’t copy written, books are, so when exactly are a collection of words a book?
There is nuance, and the linked page has none. But that’s fine, that guy is free to pull his content off GitHub. This seems like a useful feature for other people who want to make things first and foremost.
If that were true, then 20 people could each steal a single chapter from a book, and one of the people could combine those 20 chapters into a new copyright-free book. That's clearly false.
And for your strawman, the assembly of uncopywritable components into a copywriten work, would still be a violation.
So we agree that copywriter is somewhere between the paragraphs and chapters and the book. So why are tiny code excerpts “a problem”?
Of course not. Reading some copyrighted code can have you entirely excluded from some jobs - you can't become a wine contributor if it can be shown you ever read Windows source code and most likely conversely. Likewise, you can't ever write GPL VST 2 audio plug-ins if you ever had access to the official Steinberg VST2 SDK. Etc etc...
Did people forget why black box reverse engineering of software ever came to be ?
In general, you're absolutely allowed to learn programming techniques from anywhere. You can contribute software almost anywhere even if you've read Windows source code. Re-using everything you've learned, in your own creative creation, is part of fair use.
Your example is the very specific scenario where you're attempting to replicate an entire program, API, etc., to identical specifications. That's obviously not fair use. You're not dealing with little bits and pieces, you're dealing with an entire finished product.
But that doesn't mean it's the only thing it does or even that it does it frequently. It's like calling a human a parrot because he completed a line from a famous poem when the previous speaker left it unfinished.
The same argument was brought up with GPT too and has been long debunked. The authors (and others) checked samples against the training corpus and it only rarely copies unless you prod it to.
There also are cases where this happens unintentionally, but those are not the norm.
Perhaps Copilot behaves very differently from my own model, but I strongly suspect that the examples that have been going around twitter are outliers. Github's study agrees: https://docs.github.com/en/github/copilot/research-recitatio... (though of course this should be replicated independently).
I'm not really sure why we should consider Copilot legally different from a fancy pen – if you use it to write infringing code then that's infringement by the user, not the pen. This leaves the practical question of how often it will do so, and my impression is that it's not often.
That's the very reason why AI technologies can be useful in augmenting human intelligence; they see problems in a different light, can find alternate solutions, and generally just don't think like we do. There are many paths to a correct result and they needn't be isomorphic. Think of how a mathematical theorem may be proved in multiple ways, but the core logical implication of the proof within the larger context is still the same.
A human brain with an unlimited supply of pencils and paper, then.
This is wrong, this is not what Turing completeness is. It applies to computational models, not hardware.
You could argue that a stack is missing in my simplified model of the human brain, which would be correct. I used the simple model in allusion to the Chinese room thought experiment which doesn't require anything more than a dictionary.
[1]: https://en.wikipedia.org/wiki/Turing_completeness#Formal_def...
In computability theory, several closely related terms are used to describe the computational power of a computational system (such as an abstract machine or programming language)
Github could make a blacklist and tell Copilot never to suggest that code. Problem solved. You use one of the other 9 suggestions.
These aren't your crazy uncle's Markov chain chatbots. They're sophisticated bayesian models trained to approximate the functions that produced the content used in training.
At least from a copyright point of few.
TL;DR: Having right, and having a easy defense in a law suite are not the same.
BUT separating it makes defending any law-suite against them because of copyright and patent law much easier. It also prevents any employee from "copying GPL(or similar) code verbatim from memory"(1) (or even worse the clipboard) sure the employee "should" not do it but by separating them you can be more sure they don't, and in turn makes it easier to defent in curt especially wrt. "independent creation".
There is also patent law shenanigans.
(1): Which is what GitHub Copilot is sometimes doing IMHO.
No - google's 9 lines of sorting algorithm (iirc) copied from Oracle's implementation were not considered fair use in the Google / Oracle debacle.
Likewise SCO claimed that 80 copied lines (in the entirety of the Linux source code) were a copyright violation, even if we never had a legal answer to this.
The Supreme Court decided Google v. Oracle was fair use. It was 3 months ago:
https://en.wikipedia.org/wiki/Google_LLC_v._Oracle_America,_...
That's the highest form of precedent, the question has now been effectively settled (unless Congress ever changes the law).
Edit: added a dummy hash to end of URL so HN parses it correctly (thanks @thewakalix below)
> With respect to Oracle’s claim for relief for copyright infringement, judgment is entered in favor of Google and against Oracle except as follows: the rangeCheck code in TimSort.java and ComparableTimSort.java, and the eight decompiled files (seven “Impl.java” files and one“ACL” file), as to which judgment for Oracle and against Google is entered in the amount of zero dollars (as per the parties’ stipulation).
But I'm happy about all the new GPL programs created by Copilot
> Of course not. Reading some copyrighted code can make you entirely excluded from some jobs - you can't become a wine contributor if it can be shown you ever read Windows source code and most likely conversely.
You can of course read the code. The consequences are thus increased limitations, like you say.
What you mention is not an absolute restriction from reading copyrighted material. You perhaps have to cease other activities as a result.
That's not a law. That's a cautionary decision made by those companies or projects to make it more difficult for competitors to argue that code was copied.
Those projects could hire people familiar with competitor code and assign them to competing projects if they wanted. The contributors could, in theory, write new code without using proprietary knowledge from their other companies. In practice, that's actually really difficult to do and even more difficult to prove in court, so companies choose the safe option and avoid hiring anyone with that knowledge altogether.
Now the question is whether or not GitHub's AI can be argued to have proprietary knowledge contained within. If your goal is to avoid any possibility that any court could argue that GitHub copilot funneled proprietary code (accessible to GitHub copilot) into your project, then you'd want to forbid contributors from using CoPilot.
If humans did that, it would be hard to argue they didn't outright copy the source.
When a machine does it, does it matter if the machine literally copied it from sources, or first transformed it into an isomorphic model in its "head" before regurgitating it back?
If yes, why doesn't parsing the source into an AST and then rendering it back also insulate you from abiding a copyright?
You've hit the nail on the head here. If this is okay, then neural nets are simply machines for laundering IP. We don't worry about people memorizing proprietary source code and "accidentally" using it because it's virtually impossible for a human to do that without realizing it. But it's trivial for a neural net to do it, so comparisons to humans applying their knowledge are flawed.
I'm not sure why it's different, but that's a common concern with music. For example: https://www.reddit.com/r/WeAreTheMusicMakers/comments/4v8u8d...
I mean, if you used CoPilot on one computer, stared at it intensely for 1 hour, closed that computer, and then typed out code in the other computer that you were contributing from, you technically didn't use it for the contribution, you just used CoPilot for your education only.
Intellectual property is itself a flawed concept in many ways. It's like asking someone to do physics research but forbidding them from using anything that Einstein wrote.
No category of intellectual property covers thoughts, so the question has no relevance to the preceding statement.
Does it have flaws and can it be improved upon? Sure. I think society underweights what improvements to the patent system in particular could do. But such ideas are so niche they are hardly even written down, let alone debated at large. Society has bigger issues on its mind.
Like any evolved system IP law encounters new challenges over time and will be expected to evolve again, which it will surely do. A simple fix for Copilot is surely to just exclude all non-Apache2/BSD/MIT licensed code. Although there might technically still be advertising clause related issues, in practice hardly anyone cares enough to go to court over that.
Patents should have reduced with product lifecycles, copyright should be a similar period; maybe 10-14 years.
My personal opinion.
If you just used it for inspiration, that's fine; if the way it was coded is a result of technical constraints, that's fine too; if the code is generic it's not distinctive enough to acquire copyright in the first place.
Okay, so it's not law, it's just a policy compelled by preceding legal judgements. Case law, perhaps.
and they made those decisions based on the need to be able to argue in court that code was not copied.
>then you'd want to forbid contributors from using CoPilot
Right, the whole thing about arguing if copilot spits out a ten line function verbatim is not really what will be the problem, the problem is a human programmer still needs to run copilot and they will be the ones shown in the codebase as the author of the code (they could of course put a comment 'I got this bit from copilot' but might be cumbersome and anyway would hardly work as proof), although I suppose it would be not just proprietary code but code with an incompatible license.
> and they made those decisions based on the need to be able to argue in court that code was not copied.
Yeah, but only to make it easier for them to argue it; the letter of the law doesn't require it. You could argue that "Sure, I read Windows source code once -- but that was years ago and I can't remember shit of it, so anything I wrote now is my own invention." That might be harder to get the court to accept as a fact, but it's not a prima facie legal impossibility.
Cautionary decision =/= actual law.
What provision of copyright law are you referring to? Are you conflating copyright law with arbitrary organizational policies?
They said
> Reading some copyrighted code can have you entirely excluded from some jobs
And they're right. It's because of corporate policies. They never said it was because of a law - you imagined that out of nothing.
@jcelerier flatly contradicted the statement that copyright doesn’t prevent you from reading something.
You’re right that @jceleier didn’t say their example was law, that’s because the example is a straw man in the context of what @lacker wrote.
And, who says improving or clarifying a comment is poor form? What is the edit button for, and why is it available once replies have been posted?
I think you added
> Which “it” are you referring to?...
Because I have a tab open and can see the old one!
Edit - I’m adding another point as an edit to show another way to communicate. Would any of your points been lost had you done something similar?
No that’s not true. I did not edit my posts after reading their reply, and the false accusation was that I changed my comment after it was replied to.
I didn’t challenge whether the question was in good faith, but I’ll just note that the relevant discussion of copyright got dropped in favor of an ad-hominem attack.
My question of which “it” was being referred to is a legitimate question that I believe clarified the intent of my comment, and I added it to make clear I was talking about what @lacker said, not what @jcelerier wrote.
> Edit - I’m adding another point as an edit to show another way to communicate. Would any of your points been lost had you done something similar?
This doesn’t answer my question of why an edit should not be made before I see any replies, nor of why any edit is “poor form” and according to whom. I made my edit immediately. I’m well aware of the practice of calling out edits with a note, I’ve done it many times. I don’t feel the need to call out every typo or clarification with an explicit note, especially when edited very soon after the original comment.
Replies exist before you read them.
If that's the case, it should be easy to kill a project like wine - just send every core contributor an email containing some Windows code.
The result would be WINE having an advantage to redo the snippet of code in a totally new and different way and MS being forced to show part of its private code, that would expose them also to patent trolls.
Would be a win-win situation for Wine and a lose-lose situation for MS.
Semi-related, the GNU/Linux copypasta is now more familiar to some than the GNU project in general - this is a shame to me as I view the copypasta to be mocking people who worked very hard to achieve what GNU has achieved asking for some credit.
You've extrapolated "some organizations don't allow you to contribute if you've learned from the code of their direct competitor" to "You're not allowed to learn from copyrighted code", which is absurd.
...and the "if" is the important part. This is why Imaginary Property is so absurd.
> Copyright has concluded that reading by robots doesn’t count. Infringement is for humans only; when computers do it, it’s fair use.
Surely there's a limit to this. If I use a machine to produce something that just happens to exactly match a copyrighted work, now it's not infringement because of the method I used to produce it? That seems nonsensical, but maybe there's precedent for this too? (I have no idea what I'm talking about.)
That's the first time I've heard copilot get described as copying little bits of code from the Internet. Copilot aggregates all github source code, removes licences from the code, and regurgitates the code without licenses.
Furthermore, both github and the programmers using copilot know this. Look at any one of these threads written by programmers about copilot. Using copilot is knowingly stealing the source code of others without attribution. Using copilot is literally humans stealing source code from others. Copilot was written for the purpose of taking other's code.
And Github themselves have stated that only 0.1% of the Copilot output contains chunks taken verbatim from the learning set. Of those, the vast majority are likely to be boilerplate so generic it's silly to claim ownership, and maybe sometimes impossible to avoid.
That's simply not true. You might be confusing idealism about software freedom with how both law and society define theft.
Edit: In this comment I refer to the US.
The copyright lobby hedge the term as "copyright theft" (i.e. not actual theft) in order to shift the societal understanding. Whish appears to have worked.
This is not a value judgement on copyright infringement. Just that technically it doesn't meet the legal definition of theft.
cf. The rather amusing satire of the "you wouldn't steal a handbag" campaign in the UK, which ran "you wouldn't download a bear!"
It seems unlikely this distinction would ever matter in a real court though.
That can’t be right.
> Rather than "with the intention of depriving the owner", the US one says "with the intention of converting it to their use", which seems broad enough to cover exploiting a copy
...or rather, on the meaning of "converting". I've always theought of that as "changing", i.e. "it used to be one thing, and now it's something else". But copying IP only adds a use of it, it doesn't fundamentally change it in this sense: it is still available for the original proprietor's use. Is that really "converted"?
At least for the ordinary-English uuage of the word, I think it could be argued that it isn't. But then maybe this isn't just English; maybe the word "converting" also has some term-of-trade definition in that dictionary?
What could be valid is a right to not mimic collections, but that would mean you cannot clone the Copilot, as input is mapped to a non-trivial collection of outputs.
Disclaimer: IANAL, but I do dabble in IT-law.
Until someone trains a DNN to generate Mickey Mouse-like cartoons I assume.
Problem is, the results too closely replicate the source, as in "Suck just as much". And thus there is no demand for this kind of thing.
And sure, nobody cares about your stupid binary tree, but do they care about GNU and the Linux kernel? Imagine someone trained an AI to specifically output Linux code, and used it to reproduce a working OS. Is that fair?
That's a little broad. There's a wide range of licenses for software that explicitly allow precisely this.
But the execution went wrong... They used raw data from repositories for training, without any kind of pre-processing.
Copilot is redistributing it without the license.
So wait, if I write my own AI, lets call it cp, and train it on gnu-gcc.tar.gz with the goal of creating a commercial-compiler.tar.gz then I can license the result any way I want? After all most of the work was done by the computer.
https://twitter.com/mitsuhiko/status/1410886329924194309
this is clearly copyright infringement, and if it isn't: it should be
https://github.com/search?q=float+Q_rsqrt%28+float+number+%2...
The only ways this argument could be less alarming to me is if people were bothered that it was writing the same Hello World as somebody else, or that it was naming variables "foo" and "bar".
Let's wait until we have a bulletproof, egregious, and inexcusable case of it lifting code until we panic.
If those portions of code are not licensed GPL, yes, it is copyright infringement.
Will co-pilot offer royalties for auto suggestions that are committed to code bases? I’m sure our ML can track how similar the commits were.
It’s always fascinating to me how we have the tech to take, but never to give. Pay the motherfucker you stole this shit from.
The proverbial: https://youtu.be/6TLo4Z_LWu4
https://news.ycombinator.com/item?id=27710287
And reading is no infringement but writing maybe is.
But ultimately the human is OK-ing the code and committing it, basically as his own work most of the time. I'm reasonably sure that this may matter to courts.
Plus it is not "minor infringement" but code is being lifted verbatim - e.g. as has been demonstrated by the Quake square root code.
Feel free to test this theory in court ...
Many of the current Machine Learning application try to teach AI to understand the concepts behind their training data and use that to do whatever they are trained to do.
But most (all?) fail to properly reach the goal in any more complicated cases, at least the kinds of models which are used for things like Copilot (GPT-3?).
Instead what this models learn can be described as a combination of some abstract understanding and verbatim snippets of input data of varying size.
As such while they sometimes generate "new" things based on "understanding" they also sometimes just copy things they have seen before!! (Like in the Quake code example where it even copied over some of the not-so "proper" comments expressing programmers frustration).
It's like a human not understanding programming or english or Latin letters but has a photographic memory and tries to somehow create something which seems to make sense by recombining existing verbatim snippets, sometimes while tweaking them.
I.e. if the snippets are small enough and tweaked enough it's covered by fair use and similar, BUT the person doing it doesn't know about this, so if a large remembered snippet matches verbatim it will put it in effectively copying code of a size which likely doesn't fall under fair use.
Also this is a well known problem, at least it was when I covered topics including ML ~5 years ago. I.e. good examples included extracting whole sequences of paragraphs of a book out of such a network or (more brilliantly) extracting thinks like peoples contact data based on their names or credit card information (in case of systems trained on mails).
So that Copilot is basically guaranteed to sometimes copy non super smalls snippets of code and potential comments in a way not-really appropriate wrt. copyright should have been a well know fact for the ML specialist in charge of this project.
Reading by a robot doesn't count. But injecting a robot between copyright material and a product doesn't magically strip the copyright from whatever it produces.
This is silly. Co pilot is not reading by itself, someone pushed buttons telling it to read and write. If I clone the entire github without the licenses I am telling a robot to do it, doesn't make it right.
As a business it is your responsibility to determine if this code-copying is worth a risk to your business.
Based on my experience, I'm pretty sure all corporate lawyers will disallow such code copying, till it has been tested in the court. It's just a matter of who will be the guinea pig.
STOP READING
It should be obvious that if the robot is simply scraping web sites and reproducing their text verbatim (without permission and without giving credit) that would be an infringement.
There are a lot of shades of gray between that and the other extreme, which is where it is scraping millions of sites, learning from them, and producing something that isn't all that similar to any of them. Both ends of the spectrum, and everywhere in between, are things that humans can do, but as machines get more capable this is getting trickier and trickier to sort out.
In this case, it sounds like it might be closer to the first example, since significant parts of the code will be verbatim.
Ultimately, I am hoping that such things cause us to completely rethink copyright law. The blurriness of it all is becoming too much to make laws around. We just need better mechanisms to reward people for creating valuable IP that they allow people to freely use as they please.
Security is a constant issue with humans writing code. Do we really want an "AI" that understands neither code nor security spitting out snippets of code to pasted into network services?
If Copilot ever becomes truly popular it's going to be an absolute security nightmare both from the code it suggests (just bad code, GIGO as you say), but also because adversaries will be gaming it by posting bad code for it to pick up and "learn" from.
It's insanity.
I suspect Copilot won't even be up to that standard.
My objection to that term is only because being at the beginning of a learning curve isn't bad per se. Neither is writing Pseudocode.
Your conclusion is right, however.
This is, imo, unfortunate, as often the legal interpretation is based on a gross misunderstanding of how the tech works, but this is the way.
I don't think copilot should be legal according to my own interpretation but in this (rare) case I feel the "IANAL" tag applies not because I lack (legal) knowledge, but rather because I have (tech) knowledge that is likely absent from actual decision making on legal outcomes (therefore leading to different legal outcomes than how I would see things working).
Maybe nobody cares about that, but the problem is that Github's automated tool is not telling you what code it shows you is actually an exact copy of existing code, or how much of that existing code is being copied, or whether the existing code is licensed, or, if it is licensed, whether your copying is in accordance with the license or not. And without that information you can't possibly know whether what you are doing is legal or ethical. Sure, you could try to guess, but that sort of thing is not supposed to rely on guessing.
This is a very false equivalency. AI and humans are different. First, AI is at best a slave, and likely a slave of a capital. Second - scale makes difference.
At this point, how hard would it be to produce a structurally similar "content-aware continuation/fill" for audio producers, film makers, etc, which suggests audio snippets or film snippets, trained from copyrighted source material?
If prompted by a black screen with some white dots, the video tool could suggest a sequence of frames beginning with text streaming into the distance "A long time ago in a galaxy far far away ..." and continue from there.
Normally we don't try to train models to regurgitate their inputs, but if we actually tried, I'm sure one could be made to reproduce the White Album or Thriller or whatever else.
Saying it is just like user, maybe they start paying taxes like individuals without access to creative accountants pay.
Leeches without morals - Micro$oft
Until then, it's basically "GPL" (and other licences) laundering with one-sided excuses.
Of course people are hurt, namely the original creators who spent years of work and whose work is potentially laundered, depending on how good this IP grabbing AI will get.
If it gets really good, some smug and well connected loser (e.g. the type who posts pictures of himself with a microphone on GitHub) will click a button, steal other people's hard work and start a "new" project that supersedes the old one.
Now suppose it's 10 years from now and it's trivial to build a proprietary competitor.
This is a non-sequitur. Why should it?
> And nobody cares if my ten-line "how to invert a binary tree" snippet is the same as someone else's.
Are you going to make up a rule for every length and type of code? What about twenty line? If ten lines are fine then surely twenty would be? How about pictures? If some code is then surely a picture or two wouldn't hurt? Let's just tweak the AI slightly so it regurgitates more code verbatim -- or do courts have to examine any change made to the AI and okay them?
> Nobody is really being hurt when a new tool makes it easier to copy little bits of code from the internet.
The Windows source code can be found on the internet. As a human you're allowed to read that if you have it. Try making an AI that copies bits of that into your code and release that on the internet.
Then, I'm sorry, but you seem to have done a pretty bad job of summarizing that paper ?
My own take on summarizing it's conclusion would be :
"If the program can pass the Turing Test, then it should be legally liable, just like a human would."
(Yes, emphasis on the should here, but the way that you're presenting that quote might make readers think that that paper's conclusion is the opposite one !)
----
> Nobody is really being hurt when a new tool makes it easier to copy little bits of code from the internet.
Some of the examples from that paper are A&M Records, Inc. v. Napster (no comment) and White Smith Music Publ’g Co. v. Apollo Co. (piano rolls), and in that latter case we pretty much (?) got the whole Copyright Act of 1909 the very next year, where these “mechanical reproductions” were subjected to a statutory compulsory license.
So at the very least there should be concern about how Copilot might be eventually considered by the courts to facilitate copyright infringement and at the very least have to provide the source of its "insights" ?
(Attribution being the bare minimum that most of the software licenses require.)
This is a ridiculous conclusion. The ultimate destination of the robot actions' product is its user, i.e. a human. It is a clear corollary of the transitive law. Therefore, all human-focused legal concepts, including infringement, are applicable in such cases.
The absurdness of the conclusion cited above can be easily illustrated, as follows. Suppose that a person owns or rents an advanced robot (say, like Boston Dynamics' Spot, but better). He/she then programs it to break into someone's house and steal something valuable. All goes by the plan, the robot delivers the stolen goods to the rendezvous point and, if rented, gets returned. Now, according to the conclusion's logic, since a robot has done the actual "work", "it's fair use". Nonsense, right?
Just to clarify: I like the concept of GitHub Copilot (even though I have not yet tried this particular product). It offers various benefits, from pedagogical to adopting software engineering best practices to improving engineering productivity. However, I think that IP and legal aspects of this approach and specific product should be carefully studied and resolved in a consistent manner (e.g., prevent the model or system to output exact source code snippets).
Plenty of examples show that it didn't learn that much and copy literal parts of code. At my school that would have been ground for plagiarism which weren't treated lightly.
You're taking the "learning" metaphor too literally. Machine learning models do not learn. They can and do encode their training material into their weights and biases, too. That's what Copilot was doing, regurgitating parts of its training data line for line.
To me, that is not much different from transforming a copyrighted piece of work with, say, compression, a lossy codec or cropping. There are plenty of people who can learn to play Metallica songs really well, but if they copied specifics aspects of their work it would be copyright infringement, as well.
A human being can literally learn. We can understand abstract principles from one copyrighted work and apply them to another without actually infringing its copyright. A ML model does not understand, it is a statistical model. It is inherently a derivative work, and it often encodes the copyrighted work into was trained on into the model itself.
Glad to hear this. My warez group from here on out will only release binaries written collaboratively with a neural net trained on the best proprietary software available
Like when cp does it?
At first I thought this was a great feature, because easier access to code, but after some reflection, I am also very skeptical.
I am able to make my code open source, because I can make a living out of it, and I have a lot of open source code that I love to share for things like education or private stuff, but if you want to use it for something real, you need to hire me. If you can suck all the code without even I noticing it, that's not fair.
The other thing is code quality. I don't want to sound rude, but there are tons of bad code around. Not necessarily because the author is unskilled, but because the code might not need to be high quality (for example I wrote a script to sort my photos, it was very hastily written and specific to my usage, I used once and was done with it). Also, there are some bad/wrong pattern that are really popular.
I am surprised you are able to DMCA a twitch stream because someone whistle india jones theme but in this case it is considered fair use.
No I do not. Even strictly proprietary code can be copied and used in a for-profit way without approval as long as it qualifies for fair use.
Licenses and copyright are strictly related - the 'terms of use' in the form of a license simply is a grant for someone to reuse something in full granted they uphold their end of the bargain which might mean (for the GPL) relicensing code that touches said licensed code; but, if someone doesn't agree to the license or they don't uphold their end (such as by releasing their own code under the same license), the license states that the use is now invalid and they're not allowed to copy it in the permissive form provided unless the use passes the fair use doctrine.
Now 'fair use' is actually a pretty high bar, and most especially for code - feel free to read 17 USC 107 [0] which lays out how this is determined. What's tough for code is that, in terms of for copying code or using a library, it usually doesn't qualify for the initial requirement "for purposes such as criticism, comment, news reporting, teaching, scholarship, or research, is not an infringement of copyright". Unless you're taking GPL code and writing up how bad it is, or how efficient is is, chances are the use doesn't fit into fair use.
So while I think I was indeed a bit obtuse in saying that code could be used in a fair use way (since that's not what happens in these contexts) it's technically possible.
Co-pilot aside, that's already how it works today. If you make something open source, I can use your code to power my business, and I'm under no obligation to hire you. It's great when companies give back to open source, either by supporting the projects they depend on, or by open sourcing their own internal projects, but it's not obligatory.
If you don't want people to independently profit from your code, don't release it under a license that allows commercial use
"Open source" isn't a license. You're not allowed to just use any open source software that doesn't contain a license by default.
I'm honestly curious now.
"Open source doesn't just mean access to the source code. The distribution terms of open-source software must comply with the following criteria: ... The license must not restrict anyone from making use of the program in a specific field of endeavor. For example, it may not restrict the program from being used in a business, or from being used for genetic research."
Still, if you come across some published source code that does not appear to be licensed and does not specifically define itself as being "Open Source" as defined by the "Open Source Initiative", copyright law applies and you're not allowed to just take it and use it.
GitHub specifically uses the words "source code from publicly available sources" when talking about what they used to train their model on.
As far as I'm aware public code repos aren't by default "Open Source" as defined by the "Open Source Initiative".
I agree that Copilot was probably trained in part on public code that isn't open source -- GitHub's claim (not saying I agree) is that they don't need a license to train on code.
In this post's comment section alone many use the term "open source" but really mean to say "public-source", others use the two terms interchangeably even when they seem to be aware of the distinction, and then there are people who seem to think that by making your GitHub repo public it becomes OSD-spec "open source" and with that free to use.
It's just so confusing and easy to misinterpret each-others' true meaning.
Thanks for making me aware of the existence of that OSD OSS spec btw! Came across the (recovered) blog post where the term was first announced http://www.catb.org/~esr/open-source.html.
(And non-commercial licenses are not open source: https://opensource.org/osd)
> I have a lot of open source code that I love to share for things like education or private stuff, but if you want to use it for something real, you need to hire me.
implies that they have code which they are sharing under that proviso. Do you read it differently?
You are right about the technical distinction of open source from source-available. I think that the GGP (and myself) were both using it colloquially as a shorthand for source-available.
Edit: I was clearly making a distinction between open source (as in covered by an OSI-approved open source license) and only source-available, rather than treating source-available as a superset of open source.
/* web user management */
And copilot comes up with complete user management lifted out of another repo with all pages, db structures and logic but the copyrights stripped then yes. But, as I understand it, that is not what it does. You will need to slowly tell it every tiny part of how user mamagement is to be implemented and for those snippets it copies code. But when you are done, there might be snippets from 100s of different repositories potentially. I think it is hard to show that breaks copyright as many people already come up with roughly the same stuff 1000s times/day all over the world.Potentially, but not necessarily. It's possible also that if there is only one close match for the logic required, it may produce verbatim something it's already seen. GPT is known to do this for sufficiently precise inputs.
Make sure you read those licenses. Just because something is open source doesn't mean that you're free to use it for whatever you want.
There are some projects that are open source and free for non-commercial use only.
That's not open source: https://opensource.org/osd
Regardless: Copilot doesn't only pick up "open source" code... it picks up any code that has been published under any reason including the large amounts of code that is on GitHub without any license at all or which was literally stolen and leaked onto GitHub.
Meanwhile, even open source licenses have restrictions, whether they be "you can't use my work without agreeing to contribute back your work to the collective", various forms of automatic patent grants and associated retaliation clauses, or merely "you have to credit me", a simple limitation almost all open source software comes with which Copilot launders away.
You are stating a rather arbitrary assumption of yours as a fact, unless you have concrete sources or evidence.
There are many OSI approved licenses with restrictions on use, like "your software must also be open source" or "contribute back your changes" or "give me attribution"...
I don't think this is controversial? The OSI defined the term when they introduced it, in the late 90s. When Microsoft came out with "shared source" there was a huge amount of pushback from people saying "don't think that this is open source" (ex: https://www.linuxjournal.com/article/5496)
> Copilot doesn't only pick up "open source" code... it picks up any code that has been published under any reason
I agree. I'm not defending Copilot, and I think the legal questions here are interesting and tricky. My pushback here and throughout this page has been when people say non-commercial licenses are open source -- this thread started with kuon saying "I have a lot of open source code that I love to share for things like education or private stuff, but if you want to use it for something real, you need to hire me"
Just because that organization got their hands on a premium domain name doesn't mean they get to decide what that term means.
You shouldn't assume random strangers online know that you're referring to OSI's definition when you say "open source".
If I defined "open source" to mean that changes the source code must be released publicly, it's going to be pretty hard for me to talk to all the other people who already use "open source" to mean something else. There is already a standard term for the source code being publicly available: "source available": https://en.wikipedia.org/wiki/Source-available_software. You're also welcome to invent and attempt to popularize any alternative term you want, but using idiosyncratic definitions makes discussion less clear.
> Just because that organization got their hands on a premium domain name doesn't mean they get to decide what that term means.
The OSI didn't just claim "opensource.org" -- the folks behind it coined (https://opensource.com/article/18/2/coining-term-open-source...), introduced, and popularized the term "open source" over two decades ago. From the beginning they have used the same definition, which was derived from the Debian project's Free Software Guidelines.
They are not also not the only ones who use the term that way. Wikipedia has "Licenses which only permit non-commercial redistribution or modification of the source code for personal use only are generally not considered as open-source licenses" -- https://en.wikipedia.org/wiki/Open_source
If you license your software such that I can do whatever I want with it, then I can do whatever I want with it. I don't see how you can then go on to claim it isn't fair if I'm using as you allow.
If one of the richest corporations on Earth can't be bothered to share patches for permissively licensed code that they use, I will gladly shame them.
It's a different story for a small shop with no legal department and wariness about being sued over its use of open source code.
Why is it surprising? Indiana Jones is private IP. The code was published with an OSS license explicitly authorizing its use.
> It has been trained on a selection of English language and source code from publicly available sources, including code in public repositories on GitHub.
Meaning GitHub might have also been sourcing from public source / source-available projects that were not OSS licensed at all.Then you should share your code with a license that reflects that.
If your code is open source, they will get it.
That's kinda the point of open source.
I can imagine people being OK with their code being used as-is, and/or being modified, but not used completely out of context to train some corporate AI to inject code into commercial code based.
That’s a good thing.
Copyright is like an ABI: we made it up for convenience. It's just a construction, it isn't some inherently real thing.
If copyright applies, they're already in violation by failing to attribute your MIT contributions and could theoretically be sued for infringement (as they did not abide by the terms of the license).
Githubs use seems very in the spirit of open data and code. Using open source to help others.
In literally no open source license except "do-whatever-you-want-i-dont-give-a-damn-bye"-type licenses like WTFPL are you giving something away for free without limitations. That's not even close to what open source means, at all. Open source and public domain are not synonymous terms. And in those cases where there are literally any terms in a license, use of Copilot obviously violates the terms of the license, unless the user operating Copilot goes back, finds the code Copilot is stealing from, and makes the proper attribution / otherwise ensures that they are meeting the terms of the license of the stolen code.
The OpenAI people seem to grab any bit of data they can get their hands on, regardless of the source. I don't see why they'd limit themselves to Github for something like this.
Copilot is certainly pushing that envelope.
The other similar analogy is of translation: a translated work is still copied by ‘derived from’ copyright laws.
Is this just what copilot is doing in some ways but for smaller components?
There is no defending copyright. It is indefensible from first principles. It makes no logical sense.
Though it sure has proven to be a profitable con.
Because naked men are a shared concept Michelangelo's David is not protect-worthy?
I'm very worried that such opinions are up-voted so highly when Microsoft leeches open source code (but not its own ...).
People have no respect for other people's creations. Perhaps it makes them feel better because they haven't created anything difficult themselves.
Name your very best example that will prove me wrong. It should be so simple. One example, that's all it takes. Take your time, make sure you've got a good one. I'll tell you that not once, not a single time in over 17 years, have I ever seen a single example of this argument hold up under scrutiny.
Oh wait, you already did:
> Because naked men are a shared concept Michelangelo's David is not protect-worthy?
Ah yes, Michelangelo's David. A work free of copyright built under commission! Thank you for again pointing out the futility of the defense of copyright.
JK is a talented and hard working writer, and though I'm not a fan personally of those books I respect that they likely are great pieces of work, but I believe we are getting the scraps of what we could get in the Intellectually Oppressed world compared to an Intellectually free world. I'd rather have a world without cancer, a world with 100x more people able to provide medical care, a world with less pollution, than a world of artificial scarcity where a few who go along with a system of oppression get to be billionaires.
So not imaginary wizards then. What should be regulated? Is it nothing? Does your statement become "Our government should not be in the business of regulating the distribution of a sequence of words"?
Yes. Your lungs is a tree that needs healthy air. Your brain is a tree that needs healthy ideas. When people are not free to clean the ideawaves, they fill with pollution, and that is where we find ourselves.
Maybe the state could grant protection for 10 years after publishing to give the author a chance to recoup their investment. I don't know why the protection extends to the author's grandchildren.
The problem posed with copilot is in fact the opposite. By taking it to its logical conclusion, this might make it possible to disregard this effort and use GPL code on your private project.
But even if only 1% of ideas were copyrighted, that is still a tax on the use of all ideas. In a world without copyright, I can download any dataset at will and analyze and remix it to my heart's content, and share my findings. But in a world with copyright, if there was one "copyrighted" land mine in there I open myself up to financial ruin. So one must tread carefully when working with any external ideas.
If you as an end user want to modify the source code for your own use, then that is fine. If you want to distribute it, you must state the changes, and you must also do so for any code it is linked with. The original maintainers are then also free to incorporate said changes should they chose.
https://en.wikipedia.org/wiki/Debian_Free_Software_Guideline... http://people.debian.org/~bap/dfsg-faq.html
What does that even mean? The intent from the beginning of copyright was to allow people to live off of intellectual works by claiming legal rights over the work.
There are no “first principles” from which basically any societal agreements like these are derived.
Even something as simple as “murder is illegal” isn’t actually derived from any first principles because the government is allowed to murder people, citizens are during self defense, etc.
We know that there was a written intent that it was "To promote the progress of science and useful arts". However, who knows whether or not that was the true intent of all those who signed off on it. We see that lots of written intent, (Exhibit A: Purdue's "Partners Against Pain" Oxycontin promotion), may not match the mathematical reality on the ground. Also, we know that there was plenty of places in the Constitution that were good to amend (the three fifths clause, for instance).
This site (http://www.copyrighthistory.org/cam/index.php) has lots of fascinating old docs where you can come up with your own impressions about the early days of copyright. My general impression was that while it didn't ever actually promote the progress of science and useful arts, it absolutely did in the early days serve as a super smart free hack for the new federal government to build a central intelligence and library of all the latest and greatest inventions from throughout the land.
> What does that even mean?
It means that if you analyze it using logic and put all assumptions on the table (start high up on the tree), you deduce that this is a system of intellectual slavery, not of intellectual "property". You deduce that if there is such a thing as stealing ideas, then all ideas with any value are majority stolen and but a fraction novel.
I disagree. License choice is deliberate, and many open source licenses are chosen for the strict stipulations they put on users and developers, like mandatory attribution and terms of distribution or reproduction.
I release some software under the GPL and AGPL. I don't want anyone to use my software that doesn't intend to abide by the terms it was released under.
If I wanted to release software with less stipulations, then I would, and I have.
That can only be a good thing for society (though perhaps not for rent seekers).
I think a lot more people on this site (and in the FOSS community in general) would be on board with Copilot if it respected viral licenses, e.g. if it had a way of inferring that the code it was copying verbatim were GPL-3 and warned the user that including it in their project would require them to GPL-3 their project as well.
And I notice you didn't say anything about patents?
I am reminded of a line from Terry Pratchett's Going Postal in relation to a hacker-like organization called the Smoking GNU, "...[A]ll property is theft, except mine...", which I thought was rather painfully apt in describing what FOSS evolved into after becoming popular.
Indeed. Everybody is a leet haxors when they're 14, it's 1998, and we're vying for +o in #warez on DALnet. We believed information really did "want to be free".
Unfortunately some of those same kids grew up to create today's data barons and that old saying about getting someone to understand something when their salary depends on not understanding comes into play.
"No restrictions" has never been the goal and to claim that they're egoistic hypocrites who are just scared for their own livelihood because of this is just an absurd strawman.
Symbolics had a license for MIT's Lisp system.
Anyway, that still wouldn't change that the FSF and Copyleft are explicitly anti-proprietary, not intending to be 'no restrictions'.
Are you sure you're in a position to be saying things like that? The closed source Xerox printer driver incident is generally viewed as the origin of RMS's thinking on Free Software, not the Symbolics incident. And, as others have pointed out, you were mistaken even on the particulars of that.
As for Free Software not being about no restrictions, may I remind you of the four freedoms that are at the heart of the Free Software philosophy? Copilot runs afoul of none of them and I would go so far as to say that Copilot is an embodiment of 1 through 3.
"A program is free software if the program's users have the four essential freedoms:
- The freedom to run the program as you wish, for any purpose (freedom 0).
- The freedom to study how the program works, and change it so it does your computing as you wish (freedom 1). Access to the source code is a precondition for this.
- The freedom to redistribute copies so you can help others (freedom 2).
- The freedom to distribute copies of your modified versions to others (freedom 3). By doing this you can give the whole community a chance to benefit from your changes. Access to the source code is a precondition for this."
Again, the GPL is a tool to achieve the Free Software philosophy, not the end goal.
Your recollection seems to be completely off. That wasn't the goal of the Free Software movement.
Also, the code they champion comes with restrictions and, optionally, cost. So again, you're off.
As such I really don't understand why so many people are saying things like you have said here. These people believing in free software, including the original intent of the GPL to be a tool against copyright, does not mean they should, or would, be all for co-pilot, a non-free product by a company that once spat in the face of free software (and likely still does behind closed doors), who is using their free software, including the GPLed works, to create a non-free software.
As such, even if it ends up being legal from a license standpoint, it really feels like copyleft is being taken advantage of and exploited because nothing is given back in return (even if copilot remained free of charge, that's hardly giving back in foss lingo). So I suppose it's more a moral and intent thing rather than a legal thing at this point, unless copyright law decides copilot is a derivative work (doubtful, personally).
> Unless my recollection is off, the GPL was never the goal of the original Free Software movement; it was merely a tool to get to the end state where all code becomes available for use by anyone for any reason without cost or restriction.
Yes, but the GPL was created for a world with copyright to step towards a world without it. However we are extremely far away from such a world, and I don't see how copilot helps step towards it at all. I just don't see the argument. It just results in code that will be used potentially wrongly in copyrighted works, proprietary or not. And all of this is enabled by, as I mentioned, a non-free software created by Microsoft, who have used their huge capital to gain access to a proprietary AI by throwing money at it. Nothing about it seems fair even if it ends up all being legal in copyright terms.
[0] I'm seeing a lot of people in these threads who don't know the definition of free software or open source. Free software referes to free as in freedom, not free as in 'free beer' (gratis). Of course, free software is still often gratis, but you can still monetize it in various ways.
Now 'the edge' is already mostly open source. All the lock-in and value has moved into either infrastructure or in software you don't even get to touch since it runs in the Cloud and you just provide IO to it.
Endless security breaches will encourage firms to do “less IT” themselves and accelerate the adoption of SaaS solutions (and PaaS, with no/low-code etc.)
Also, perhaps not a massive driver but still, not for nothing: M1-style processing innovation (ARM) will see more developers creating for ARM servers, because they can, which will almost exclusively be run by the hyper scale cloud providers.
As it stands, the combined inputs leaves the model in the most murky of gray areas.
Since the Microsoft acquisition it’s becoming painfully obvious how unhealthily centralized the dev world has become, and they seem to strive to become ever more entrenched in the name of maximizing shareholder value.
I only have a small amount of open source projects on GH but I intend to vote with my feet and abandon the platform by self-hosting Gitea. By itself it won’t be a big splash but I’m inspired by posts such as this and I hope to inspire someone else in turn. Of all people we devs should be able to find good ways to decentralize.
It’s interesting how incumbent companies such as Google and GitHub try to capitalize on their user data with machine learning in any way they can to maximize shareholder value.
GH and MS spend a lot of time talking about how important open source is to them. They didn’t exactly prove this by building Copilot to be so oblivious about licensing. Either they took a gamble and hope they’d get away with it or it didn’t occur to them that this would be a problem at all. Either way, I’ve lost faith in GH's ability to act in the best interest of its users and the larger open source community.
It’s a free market and I hope to see more competition in this space: Both from GitHub-alternatives that respect code-licenses and from self-hosting alternatives.
It appears that GitHub wishes to address this issue via UI changes to Copilot. A quote from a recent post on GitHub[0]:
> When a suggestion contains snippets copied from the training set, the UI should simply tell you where it’s quoted from. You can then either include proper attribution or decide against using that code altogether.
> This duplication search is not yet integrated into the technical preview, but we plan to do so. And we will both continue to work on decreasing rates of recitation, and on making its detection more precise.
That post is also on the Hacker News front page right now[1], but has 10% of the upvotes as this post so it's less visible.
I'm hoping all the criticism will encourage GitHub to make a better product.
[0]: https://docs.github.com/en/github/copilot/research-recitatio...
on edit: fixed typo
I would also be extremely surprised if most open source copyright holders didn't already expect their licensing terms to protect against this kind of code/authorship laundering. Speaking individually, I know that it certainly surprised me to hear that GitHub thinks that it's probably okay to regurgitate entire fragments of the training set without preserving the license.
It should be considered as fair-use of USA except we don't use Common Law system so we explicitly state what exempt from the copyright protection.
That said - so they would be able to sell some things in Japan that they couldn't other places.
It's still copyright laundering, if you ask me.
My understanding as a two-year student of ML is that you are allowed in the US to go download any old image, train on it, and then release the model as long as the outputs are "sufficiently transformative."
That last phrase is the key part, and has never been tested in court. It's entirely possible that either I'm mistaken here, or that the courts will soon say that I am mistaken here. https://www.youtube.com/watch?v=4FA_gt9w28o&ab_channel=guava...
Copilot seems ... well, less transformative. I'm still not sure how to feel.
https://twitter.com/luis_in_brief/status/1410985742268911631...
The irony is that copilot won't suggest its own source code, just everyone else's. It is open source without the benefits.
Yeah, don't post your code in the public for everyone to read. If I am musician, and I play my song publicly, people will hum it if they like it. I can't do anything to stop that, except not play the song to anyone.
Plus, to be honest I’m not even sure whether I’m for or against that, I am just wondering if one was against it but still wants to do open source work do they have any recourse?
Copilot is no different than stackoverflow with the exception that your selection mechanism is done by an artifical neural net rather than a bunch of people with up and downvotes in front of their screen. The reason people are uncomfortable here isn't some vague PR speech notion about the future, it's that copilot appears to be quite literally ignoring software licenses. Can we not devolve into this Silicon Valley corporate speak of rebranding a company ignoring intellectual property as innovation?
Why does it get a lot of upvotes? Because it's an AI product that makes things appear on your screen and that's the threshold for hype in our current age. Of course the number of upvotes on HN regardless doesn't speak to the innovation of anything. As best as I can tell all pre-covid posts on HN concerning mrna vaccines have a total of one upvote. Useless tech right?
Weakening copyright also weakens copyleft - for example, it seems reasonable to me that the producer of an open-source work should be entitled to require reciprocal openness from people who build upon it. If I can legitimately launder some GPL source code (say, a Linux kernel driver) through an ML model without being obliged to release the resulting code, I think everyone loses.
only people who have released their code publicly under a (mostly) open license
so, not Microsoft
It's a bad idea.
The thing about copyright law that needs reform is its bias toward the benefit of large corporate entities. Platforms' implementations of DMCA compliance allow "rights holders" to spam perjurious takedown requests en masse, garnishing the earnings of creators and legitimate rights holders in what can only be called (in addition to perjury) outright fraud. Companies like Github scrape the web for content, most of it copyrighted, and use it to construct new products for their own profit. Rare recitation events aside, I think their use case is legitimate fair use in the eyes of the law (and if you look at my comment history you'll see me vehemently arguing to that effect), but should it be? We don't seem to be asking that question, which is really disappointing--we're either complaining loudly and without substance, or blithely accepting the might-makes-right ethic as the central pillar of our IP law.
That doesn't look like it's the point to me.
""[the United States Congress shall have power] To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right , to their respective Writings and Discoveries." "
As I read that, copyright is there to 'promote progress', not to maximize gains.
No doubt there is a million linear feet of case law that got us where we are.
Honestly, I rather like this whole question of copilot. I solidly appreciate the brilliance of github as a honeypot.
What better way to promote said Progress than by making sure said Authors and Inventors can make enough money off their work to keep doing it? As written, it's a roundabout way to get at the instrumentality of capital, but if that's not what they had in mind then I'm not sure what they were getting at. Without copyright, a creator's rights to their own work aren't diminished; it's just that everyone else's are expanded to the same level.
(I'd love to know if I'm way off base about this. I'm not a lawyer, and I'm sure it's been discussed to death.)
> Honestly, I rather like this whole question of copilot. I solidly appreciate the brilliance of github as a honeypot.
I think it's really cool, and I'd probably use it myself. As much as my favorite kinds of programming (e.g. writing experimental text editors) might not benefit from it, in my day job I sure would love to spend less time filling in boilerplate and looking up mundane API details.
I don't mean to single Github out in my mention of big corporations benefiting from copyright law. Scraping vast quantities of copyrighted data to build new products is a common business model at this point, and--like other new IP-related paradigms enabled by modern information technology--I think it deserves a fresh look, being mindful of just what it is we're trying to accomplish with copyright law. As you say, it's not always obvious, even in written law.
There are a lot of better ways. Having more information being public and free, and usable by tools like this sounds like an excellent way of promoting progress.
If you want to strip creators of (default) exclusive rights to their work, that's a different conversation than the one around whether Copilot and similar applications fall under fair use. Both have been touched on in this subthread, but your comment seems to be conflating them in a way that doesn't follow directly from the discussion above it.
Ok great. So then new tools like this are good, even though they weaken copyright, and the concerns that people have about it (that it allows easier copying of code), are actually a benefit.
And the fact that it might hurt people's ability to profit from their code, is overruled by the benefit that this stuff provides.
> doesn't follow directly from the discussion above it.
You suggested protecting profits as if it is the only or best way of promoting progress.
When, in reality, stuff like these tools are actually a much better way of doing so. And it does so in a way that undermines copyright law, in a beneficial way.
Does it weaken copyright? Like I mentioned, it seems like it's probably allowed under existing law.
> it might hurt people's ability to profit from their code
I don't really buy this. The outputs of Copilot seem transformative enough that they won't by themselves meaningfully compete with the applications built from the sources in the training set.
It seems to me that people are objecting to it more as "theft" on ethical grounds alone. I don't really have a strong opinion either way on that front, but if I did it would be based on principle and not some theoretical material harm, because I think the latter is marginal at best.
> You suggested protecting profits as if it is the only or best way of promoting progress.
In fields where creators make money by selling access to copies of their work, what is a better way of promoting progress? People need places to live, and things to eat, and other things, and all of that costs money. If working in these creative fields becomes even less lucrative than it already is, fewer people will be able to do it, and for less time, because they will have to spend more of their time making money in other ways.
In tech, many of us are privileged to have a fair amount of spare money and time. Don't forget that not everybody enjoys that privilege, and please try not to attach a negative-valued concept of "profit" to the necessities of survival.
> When, in reality, stuff like these tools are actually a much better way of doing so. And it does so in a way that undermines copyright law, in a beneficial way.
Again, this is a very tech-centric view. I can't imagine (for instance) the average novelist being particularly happy to have the exclusivity of their rights to their own work curtailed to enable the creation of some tool, using their work as an input, for generating prose. And such objections would be absolutely correct, if anybody was actually talking about doing that.
Fortunately, nobody is talking about doing so--not for code with Copilot, not for fiction prose with the new GPT-3 tools that are popping up, and not for any other medium I'm aware of. These applications are covered under existing fair use law, and their existence does not depend on weakening the exclusivity of creators' rights to their work.
If you were to tell me that such rights should be curtailed to enable tools like Copilot to exist, I would strongly disagree with you. But--again--such curtailment is not necessary. The only reason I'm talking about it here is that there are people who think copyright should be abolished or strongly weakened. Almost universally, I've found, they're people who don't make money off distributing authorized copies of their work. So, if you ask me, they really have no clue what they're talking about, and shouldn't be running their mouths before seriously listening to (at least) the independent creators who would be impacted by such a change.
Yes it does. It is legally allowed, but in the past it was much more difficult to code launder, or copy things, in the way that an AI would do it.
Something becoming easier to do, has an effect, even if it was legal in the past.
> what is a better way of promoting progress
More technologies like this, that allow better sharing of code and information. It reduces the barrier to entry to creating content, thus causing more of it to be made.
> If you were to tell me that such rights should be curtailed
You have it reversed. Rights do not need to be curtailed, to enable these tools. Instead I am advocated for the production of these tools to be done for the purpose of curtailing these rights. The causation is reversed.
The rights should be "curtailed" through the process of tools that allow people to easily get around the law, and to make the law unenforceable. Changing laws is much harder than making the law irrelevant.
We don't need to change any laws, if we just make it impossible for laws to be enforced.
It is kind of like how bittorrent undermined copyright laws. No laws needed to be changed, for piracy to become rampant and unpunishable. (And don't even try to challenge me on this point, that piracy is effectively unenforceable these days. If you do, I'll just go watch the lastest episode of some marvel show, for free, right now, lul)
Like I said, this seems like a highly tech-centric viewpoint. Keep in mind that source code is far from the only thing covered under copyright law. Personally, if I was stranded on a desert island, I'd rather have a single original novel written by a human than a hundred novels' worth of GPT-3 output.
Beyond that, your perspective is pretty interesting--I guess you support the existence of tools like this because you see it as an opportunity to erode existing copyright law. Personally, I may not support the full extent and implementation of copyright law in America, but I do support the fundamental principle that a creator should have exclusive rights to their work. So we disagree pretty strongly on that, and I doubt we'll find common ground.
I guess I would just urge you, if you value art at all, to consider how independent artists like writers and musicians would be affected by the elimination of copyright. I don't really give a shit about the IP rights of programmers (even though I'm one myself, with public FOSS contributions), but you seem willing to throw out the baby with the bathwater.
Gaming things out, I don't think copyright is really helping any of those artists or society as a whole. If it didn't exist, you'd still have breakout artists who make money through endorsements, live shows, and selling original copies of their work.
I agree that there needs to be talk about licensing and copyright but with so "less/no content" there can be no meaningful discussion other than aimless banter.
And their only huge project on GitHub is dbxfs, a userspace dropbox filesystem with 687 stars https://github.com/rianhunter?tab=repositories&q=&type=&lang...
I think this is just a post meant to continue the discussion of CoPilot past the first 2 days of news.
One of the beautiful things about HN is that you don't need to be anything, you just have to have something interesting to say.
It's just posted(not by the guy that made the page, mind you) to farm karma, exploit the news cycle and carve out some more space for discussion of this tired topic.
> It's just posted(not by the guy that made the page, mind you)
Others would complain if the author himself had posted this.
Sometimes we've seen it but it is a new angle.
If he had a good argument for that, fine. But without that he really needs to be someone whose opinion I care about.
> What is there that is not solid, CoPilot is using community code that is under GPL licence therefore Microsoft should not be able to charge for CoPilot but give it for free, or not create another revenue stream.
You're doing the same thing as OP by assuming that this is illegal. That had yet to be determined. It could easily be the case that this falls under some fair use law or isn't even covered by copyright. It isn't for humans!
Fair use can be drill down to following, a very simple exploit question.
Would you work for free, if someone would earn million on your own work?
if (answer===yes) {
"then let me employ you I have few ideas and I need free work that could make me wealthy."
} else {
"if no then your entire argument is pure hypocrisy, trying to justify wealth built on exploit of others" }
The topic is a current one [1], which makes it even more valuable.
Why is this comment noteworthy? Who is this person? Am I missing something?
And basically every software company avoids GPL like the plague, due to its strong copyleft conditions.
Other popular choices are gitweb and cgit (both dynamic).
https://codemadness.org/stagit.html
Contrast these two pages, and you'll see it's a match:
Thank you. :)
Love the minimal style and monospace font!
on edit: just saw there was a description of who he is https://news.ycombinator.com/item?id=27724247 as noted I don't know but not sure if it's enough to imply a bad motive of him wanting to get some sort of attention for opposing copilot.
on second edit: huh, seems to be one of those occasions when I have mysteriously offended some people on HN without swearing, joking or being rude.
Has this ever been used as an argument in a legal case?
So hosting elsewhere might not safe your code from ending up deep in the bowels of some corporate black-box ML model that occasionally regurgitates your IP if accidentally given the wrong (right?) prompt.
If you make your code public, you basically accept that someone will copy it verbatim. Other companies still might have it in their closed source product somewhere, even if it's just accidental copypasta from SO.
Amongst other things, hosting something on github is a public ledger of authorship.
I seriously doubt that Copilot scans my (very few) private repos. Even then, I don't think I do anything particularly noteworthy.
But that is just me.
[1] I also MIT license my public code on Github, and also wouldn’t care that much.
The only reason I use MIT, is so some knucklehead doesn't try to sue me, because they cheezed up my code.
What if you published a subtle proof of concept that takes out nuclear plants, and then some knucklehead deployed it because Copilot suggested it?
I could see certain...agencies doing something like seeding the tech scene with insecure hashing algorithms, and them becoming a part of the Canon, due to consumption by uncritical ML training algorithms.
We get back to the old "data quality" conundrum. We need to have a way to rate the quality of our data, which then opens the door to corruption and gaming.
The circle of life, I guess...
Unlike other licenses, I don't think I'll have a bunch of self-appointed juntas going after people that use my code without attribution.
I suppose one issue is that you (presumably) can't request deletion from it (which may even be a GDPR violation).
Edit: I looked up the relevant GDPR stuff, apparently there's an exemption for when "erasing your data would prejudice scientific or historical research, or archiving that is in the public interest.", which it arguably includes the Arctic Vault.
It's nothing more than a publicity stunt whose one and only purpose is to advertise GitHub.
What if I were to tell you, that in order to publish any code, on the internet, that code has to be "reprinted" to many different computers and places?
In fact, whenever you yourself need to even access that code, that code is copied over to many different computers along to way, as is necessary to send it to you.
Github are the ones doing all the archiving. So, in essence, they do own that. Piql are just the ones providing the storage: it's a commercial for-profit entity employed for backup by another commercial for-profit entity.
By the way, my initial statement that it may qualify for copyright exemptions turned out to be false for a different reason. They only apply when the library and/or archive in question is open to the public, and the Github Arctic Vault isn't. Thus I think it's actually a Github's generic usage grant in the ToS [2] that allows for the Vault. The Copilot is, of course, very different to anything described in the ToS.
[1] https://arcticworldarchive.org/contribute/
[2] https://docs.github.com/en/github/site-policy/github-terms-o...
...provides prime-rate marketing bullshit in its marketing materials
> Thus I think it's actually a Github's generic usage grant in the ToS
If you refer to Section D.4, then:
- Arctic Vault is not "for future generations", but for GitHub only, since that section doesn't permit GitHum to just make copies willy-nilly for anything other than "as necessary to provide the Service, including improving the Service over time" and "make backups"
- This specifically makes GitHub "the owner" of that data, and not "some third-party" as you originally suggested
> This specifically makes GitHub "the owner" of that data, and not "some third-party" as you originally suggested
This one is my fault though, I've used the "Arctic Vault" as an archival site, but as I later realized it is a Github's archive stored in the Arctic World Archive. So yeah, it's (only) Github that can retrieve the data.
This is a commercial for-profit company, GitHub, taking some code and storing it in cold storage of some other commercial for-profit company, with no one, except these two parties have access to this code. And it doesn't look like GitHub even has the right to do it because it stores it for some purpose other than whatever is stated in their ToS.
I wonder if the whole kerfuffle around Copilot will end up spilling some light on this, too.
If they had taken one of their existing DB/disk backups and called it a vault, would that have been an issue?
As a senior developer, I am strongly biased against the SO+c/p programming approach that I’ve seen many Junior and mid level developers use. There’s certainly a time and place for it when you become really stuck but at least having to go out and find the code yourself requires thought which helps you grow.
My gut reaction to Copilot is that adding this automation into IDEs is going to have a net-negative effect on growing developers as it lowers the level of thought and effort necessary to write even trivial applications. This is a huge detriment to learning. You don’t even get the chance to try to solve the problems yourself because the AI is going to be proactively getting in the way of your learning.
All that being said, I think a tool like this could be of great use with boilerplate within a project — but only suggesting things from that project. For example, setting up a new api route, dependency injection, error propagation, etc. Help with all of these mechanical things would be awesome.
I sometimes code from devices which are not my own and on which key management is a major impediment and accessibility issue for me.
Does anyone know how those listings were generated? I like their simplicity, and would like to do something similar.
Then they came for the warehouse and shipping jobs, but I did not care because I wasn't in a warehouse.
Then they came for the call center jobs, but I did not care, because I did not work in a call center.
I was safe, because they could never automate programming.
Interesting times ahead. For example, if you believe these kinds of tools will become a huge competitive advantage, and that the inclusion of GPL code is a meaningful force multiplier, it kind of implies the fusion of AI code generation and the GPL will eat the world.
If anything, I think the real test of this tech is going to be audio, as it has the right overlap of "big copyright is going to get pissed", "there already exist tools that attempt to automatically detect even small bits of infringement", "people actually litigate even small bits of infringement", and "it feels feasible in the near future": you whistle a tune, and the result is a fully produced backing track that sometimes happens to exactly sound like the band backing Taylor Swift on a recognizable song and generates Taylor Swift's voice, almost verbatim, singing some of her lyrics to go along with it.
https://en.wikipedia.org/wiki/Clean_room_design
E.g.
1. Train one ML implementation to produce "specification text" in a way that they're agreed to be free from copyright claims. E.g. train to avoid any direct quoting, possibly via a different human, programming or custom specification language.
2. Train a separate ML implementation to produce code from the specifications.
3. Hook them together and you've got a pipeline for generating learned, but copyright-free, code.
Kind of reminds me a bit of some of the machine translation work with human languages.
Note: this is how the GNU project itself sometimes clones the functionality of copyright-free way, so I'm pretty sure it would be safe to use this on GPL-licensed code.
I expect the best lawyers from Microsoft have had a look into this and maybe there a weaknesses in GPLv3 ready to exploit for corporate AIs. What is the response from the FSF?
This just seems like a massive lawsuit waiting to happen.
What happens when you discover that you're using '20 lines of code from some GPL'd thing'?
What will your lawyers say? Judges?
It seems to me that if you use Copilot there's a straight up real world chance you could end up with GPL'd code in your project. It doesn't matter 'how' it got there.
I don't understand therefore how any commercial entity could allow this to be used without absolute guarantees they won't end up with GPL'd code. Or worse.
The only reason open source software is possible is because of the protections offered by copyright. Copilot basically shits all over that.
I also don't see how any of this follows - they could've just crawled GitLab or any other OSS repository. They didn't even need Github for this.
Heck, is OpenAI doing embrace, extend, and extinguish on the entire web now, because they use Common Crawl [0] to train GTP-3, which forms the basis of CoPilot?
The GitHub CEO is no where to be found to answer the important questions on software licenses, copyright and the legal implications on scraping the source code of tons of projects with those licenses for Copilot.
The fact you can only use it in VSCode and with Microsoft having an exclusive deal with OpenAI screams an obvious 'embrace and extend',
As for 'Extinguish', they will need to be very creative on that.
It could maybe at some point send rewards in form of donations etc. from Copilot users, similar to Sponsored repos.
This is ultimately a Microsoft project, and they have Microsoft money and Microsoft lawyers to defend their position.
Many of its features are available in the self-hostable free and open-source version, GitLab CE.
But what did we expect from GitHub....MS or not? This was an obvious survival mechanism for GitHub sans MS. All that coding data there? Let's turn machine learning or AI onto that and make something.
And if MS were treating GitHub as an "at arms length" corporate entity so as to NOT to upset the opensource/free software community (because MS) then the blame lies fair and square with the management of GitHub.
MIT? GPL 2? GPL 3? BSD? Apache?
They're not interchangeable, and not all are directly compatible.
Say what you want about GitHub's almost monopoly position, but the UX is really great and accessible even to non-technical people. Maybe you don't need that, maybe you don't want the issue-trackers, but it's worth thinking about who you're excluding with these kind front-ends.
As opposed to an anxiety that a machine might be able to do some of our jobs better than we can?
As it is now, it works towards weakening the copyright of free software while doing nothing (or very little) to closed software.
But what did we expect from GitHub....MS or not? This was an obvious survival mechanism for GitHub sans MS. All that coding data there? Let's turn machine learning or AI onto that and make something.
And of MS were treating GitHub as an "at arms length" corporate entity with their own lawyers etc such as to NOT
Won't be long until we see an infringement case. /me grabs popcorn
Unfortunately, GitHub enjoys the effect of most people being on it (correct me if I'm wrong), and leaving it is costly, regardless of whether the alternative is a reasonable service or not.
They send the request off to a sweatshop in Bangladesh where a bunch of mechanical turk workers scroll through licensed codebases, find an appropriate snippet of code, agree on the best one, strip all of the licenses and attributions out, and send it back to you. (turns out they're very quick at their job)
How is this any different than what Copilot does with the purely technical difference that you're pushing code through an artifical neural net instead of a bunch of human ones. Why is that supposed to matter
also, someone could deliberately obfuscate the license text to fool it but still be clear enough for humans. something like "License: if you use this source code to train a bot then you must obtain a commercial license, otherwise MIT license applies". bot searches for "MIT" and thinks it's safe.
Whether training an ML model on code is fair use is still an open question, but I don't think GitHub is a greater villain here than anyone else doing the same thing (at least until they start using private repos).
I don't ask for opt-in when I copy a snippet from Stack Overflow or use some code from an MIT licensed repo.
The only question, I think, is if what Copilot does is actually in compliance with the licenses.
Not everyone has the patience and ability to discuss their objections in a public forum while their rights are being violated (in their view).
Some people have a passion and a very strong belief in their ideals and I applaud them for following through with it, even if I don't necessarily share their opinion on the matter.
It's not the technology, but the fact that any code you pushed to GitHub in the past 13 years is now 'accessible' to anyone.
Private repo? Paid account? Deleted repo five years ago? Deleted repo today? Proprietary code? Embarassing commits? Accidental API keys or passwords in commits?
All 'available'.
It feels like the entirety of GitHub was just 'leaked', and converted into a marketable product.
Would you push your code to a service if you knew it could be read by anyone one to ten years from now? Even if you paid to keep it a secret?
> GitHub Copilot is powered by OpenAI Codex, a new AI system created by OpenAI. It has been trained on a selection of English language and source code from publicly available sources, including code in public repositories on GitHub.
I wonder what Microsoft will do when snippets from that code start appearing in your code because of copilot. I'm guessing their lawyers wouldn't accept "the robot did it" as an excuse in that case.
I'm tempted to just throwing stuff like "AWS_KEY=" at the algorithm and see how many working credentials I can steal from private repos.
Anybody tried? What does actually happen if you do this kind of thing? I can think of a few more obvious "script kiddie" ideas, but I won't post them here lest a copilot developer sees it and closes all the elementary stuff.
I'm old enough to remember when "assume anything you put in cleartext online is public" was received wisdom. We were taught that if you want to keep something private, keep it encrypted on your own local media. Or, failing that, at least on a server you control.
50% is anti-big-tech and 20% is our fear of being made redundant.
Will it be fair use then?
Fun.
"You know what you know."
My sense is that this is either:
a storm in a teacup;
a blackhole that swallows everything around it;
a massive copyright mess that piles up without anyone noticing then explodes all over everything;
or something else entirely.
The next few years will be interesting then. I'm wondering what happens if/when a significant chunk of GPL code gets included into a commercial product. That will get lively.
Popcorn with butter please.
[0] Bakos, Y., Marotta-Wurgler, F. and Trossen, D. R. (2014) ‘Does Anyone Read the Fine Print? Consumer Attention to Standard-Form Contracts’, The Journal of Legal Studies, 43(1), pp. 1–35. doi: 10.1086/674424.
[1] McDonald, A. M. and Cranor, L. F. (2008) ‘The Cost of Reading Privacy Policies’, A Journal of Law and Policy for the Information Society, 4(3), pp. 543–568.
There's plenty of skepticism there, even in the early comments.
I never considered the copyright and related ethical implications of ML at all, or thought about the impact it may or may not have on programmers. Your first thoughts on something can be wrong (and actually, often are) and it takes a bit to really think things though – or at least, it does for me.
You can spin your wheels all you want but going from simple first principles it is fundamentally flawed. If you believe ideas can be property, then you believe people can be property.
Can you defend that? I generally think copyright isn't a great idea as it exists, but this statement feels extremely dubious at best.
Yes. Though I'm not the sharpest tool in the shed, so the delivery may be less than ideal.
Copyright gives PersonA legal control over a subset of PersonB's behavior, even when PersonA is not involved. This is hard to defend, unless you are fine with people being property. Slavery gives PersonA legal control over PersonB's behavior. Under Slavery, PersonB has one master with lots of control. Under Copyright, PersonB has lots of masters with small controls.
Is there a way to believe that ideas can be property without it being a system of slavery? There's no logical way to make that work. Think about the moment that copyright "expires". At that instant, does matter disappear? Did property vanish? What changed? The only thing that changed was each person suddenly gained a little more freedom—the ability to share a new sequence that they couldn't share before. The property rights that the copyright holder had over other people went away. People became more free.
Anyone who thinks for themselves should be able to quickly deduce that these laws are shades of slavery laws, and not about property rights. I figured that out before I could legally drink. It's not that complicated. The question is why are so many duped? I think it's probably a question of priorities (I would say Freedom of Speech and the Press are more fundamental, and then Freedom to Remix and Distribute would be next) or perhaps it's because before the Internet there wasn't enough uncontrolled bandwidth for the truth to get out, or perhaps it's that the people are bombarded over and over again by the big lie from the moment of childhood—look at the FBI Warnings at the beginning of Disney Movies, or the dozens of times per day that you see the phrase "All rights reserved".
The Internet isn't really a place to exercise an inflexible moral code. His new repository probably can be traced back to slave labor somehow if someone digs deep enough. Probably won't even take 6 degrees of separation.
If it makes it easier for me to code and gives me more time to do something other than work without doing irreprovable harm to some sentient entity, I'm firmly in the who gives a shit camp.
And you're ok with that? It doesn't HAVE to be like this. Just because you've chosen nihilism, doesn't mean that's the only choice, and it certainly doesn't help anything.
It doesn't HAVE to be like this, it just is, and all of the alternatives suck. If you want to choose to inconvenience yourself in order to pass a morality test that doesn't exist, go ahead I suppose.