GitHub Copilot investigation
githubcopilotinvestigation.com
githubcopilotinvestigation.com
It seems clear enough to me that training AIs on copyrighted works is typically or commonly a fair use under existing law, because the AIs can and commonly do learn non-copyrightable elements and aspects of those works. It's very obvious from enormous numbers of examples that current AI systems are capable of learning much more abstract features of human culture (grammar, concepts, facts, cultural tropes, and many others).
A human being doesn't violate copyright in learning from a copyrighted work, including when that human being is later more able to produce other works based on that learning (e.g. reading fantasy novels and learning concepts, tropes, or vocabulary that one uses to produce other fantasy novels; reading a newspaper and learning facts that one incorporates into an essay; learning artistic techniques or stylistic conventions from studying existing artworks and using them when producing new artworks). Current AI systems are (amazingly) becoming capable of all of these things and may do them in ways that are somewhat akin to how human beings do them. (although I guess Jaron Lanier would object "that's what they want you to think")
But there are also examples in existing copyright doctrine where people accidentally repeat enough of a prior work to get in trouble for infringement -- most often with song composition (like George Harrison's "My Sweet Lord") because relatively small pieces of melody (which a person might easily memorize) may be considered copyrightable.
If human beings had much more accurate memories, copyright would be quite a bit more intrusive (and/or quite a bit less effective) because, following any exposure to some kinds of works, we could use our own memories to reproduce those entire works from scratch for our own use or pleasure without obtaining authorized copies from elsewhere.
Computers do have such accurate memories, and machine learning systems, which are optimized for things like maximum likelihood estimation, can and do reproduce both copyrightable and non-copyrightable elements of works that they've been trained on. After all, the maximum likelihood continuation of a fragment of a text or a song is ... the complete original work. And the ability to reproduce the complete original work would, other things being equal, reduce loss in training. After all, that's something someone might specifically ask for, and if the system could oblige, it would be doing a better job of providing what the user wanted.
It's relatively foreseeable that machine learning systems would potentially be able to reproduce both copyrightable and non-copyrightable elements of various works, because the distinction between the two isn't especially clear from an algorithmic or mechanical point of view. (For instance, facts aren't copyrightable, but the notion of what constitutes a "fact" for this purpose is a culturally-bound legal notion and not at all straightforward to make precise.)
But if you had a human author or artist or scholar or programmer who was "trained on" exposure to an enormous body of works, and that person had an exceptional eidetic memory, you could imagine that he or she would be perfectly capable of recreating many of those works from memory (and that other people might request such recreations). (Again, in music in particular, it's already routine that someone could have unambiguously copyrightable material memorized and be subject to copyright restrictions on performing songs. Like if a singer or band performs a cover from memory.)
If you wanted to avoid this ability then you might need to build in an explicit notion of copyright that limits the accuracy or level of detail inside of the model in some way. This is tricky because (1) I don't think people have really tried to do this much so far, (2) copyright applies very differently to different categories of work, (3) it obviously wouldn't satisfy critics even if it mitigated the most extreme examples of "regurgitation", and (4) it would be kind of weird because you would be intentionally limiting the quality and extent of learning that the system was allowed to do. (I imagine Jaron Lanier getting mad again about my repeated comparison between human learning and machine learning, and between human memory and machine memory)
Some of the weirdness in point (4) is that accurate prediction is usually cool / great / impressive / accepted as an appropriate goal or capability, but if it's too accurate in certain contexts, it may be deemed a copyright infringement. Like if you said "what word comes next? FOUR SCORE AND SEVEN YEARS AGO OUR FATHERS", there's a clear correct answer and knowing it requires having a certain text memorized. OK, if you said "what word comes next? MR. AND MRS. DURSLEY OF NUMBER FOUR PRIVET DRIVE WERE PROUD TO SAY" ... same thing, but Bloomsbury Publishing may be unhappy if you have a system that can get all such questions right.
I don't know the name, but I remember some sci-fi story about some academy where humans were trained from birth without exposure to music others had written and had to reinvent it on their own. Some would cheat and access the outside world's music, but they would always be caught by their later compositions all having obvious influence from conventional music.
<Insert obvious joke about how I'd remember the name if I had computer-like memory.>
You can read it here: https://b-ok.cc/book/4395497/b2fb2e
It turns out that humans can extrapolate generalisms to a degree we are currently unable to explain clearly enough as a model to imitate.
It turns out that much ML is merely referenced regurgitation.
Marketing and hype are rather advanced skills in 2022, however…
But this can't be right, it is inconsistent with how copyright has worked so far. Artists and musicians and engineers all learn from each other and have seen and learned from, "trained on" many other examples of works from their field. Even when works are clearly inspired by other works we tend not to give them the legal status of derivative work.
You're suggesting we treat models with a much stricter copyright regime than has previously existed.
If we get to the point that capitalist maximalist-utilitarianism insists upon hijacking the very concept of what a living organism is, I can only compare it to a teddy bear vs an actual bear.
It’s not enough to merely put fur on it and an internal rom for it to regurgitate prefabricated roars upon contextual prodding.
Respect the life you are only one instance of, for hubris has always brought suffering and pain in its wake.
You have a mechanism that can regurgitate (digest, remix, emit) without attribution all of the world's code and all of the world's art.
With these systems, you're giving everyone the ability to plagiarize everything, effortlessly and unknowingly. No skill, no effort, no time required. No awareness of the sources of the derivative work.
My work is now your work. Everyone can "write" my code, without ever knowing I wrote it, without ever knowing I existed. Everyone can use my hard work, regurgitated anonymously, stripped of all credit, stripped of all attribution, stripped of all identity and ancestry and citation.
It's a new kind of use not known (or imagined?) when the copyright laws were written.
Training must be opt in, not opt out.
Every artist, every creative individual, must EXPLICITLY OPT IN to having their hard work regurgitated anonymously by Copilot or Dall-E or whatever.
If you want to donate your code or your painting or your music so it can easily be "written" or "painted", in whole or in part, by everyone else, without attribution, then go ahead and opt in.
But if an author or artist does not EXPLICITLY OPT IN, you can't use their creative work to train these systems.
All these code/art washing systems, that absorb and mix and regurgitate the hard work of creative people must be strictly opt in.
That's how the law needs to be.
This would prevent anybody from accidentally infringing when using these tools. Does that seem like a reasonable solution, or is your concern greater than accidental infringement?
A world without copyright would however save more than just a few millions of man hours that copilot might do. Allowing people and companies to freely use the best software available, view the best art, enjoy the most relaxing music, have the best recreational time with the best films. The only harm is the highly speculative claim that people won't be creating the best software, the best art, the best music or the best films.
In those cases it seems that humans are already copying code without also propagating licenses appropriately. LLMs are more likely to memorize things which occur a lot (and I'd bet rare things that are representative of some conceptual axis).
The main examples presented so far, Davis and Carmack, have the property of having been copied a lot. The generative model is only surfacing an existing pattern of ignoring attribution. Sort of like the code-gen version of generating bigotry if appropriately prompted.
I'll also note that this pattern of retrievable memorizing of copyrighted and sensitive material is present in GPT-3 too and not just for code. As the situation is equivalent, a lawsuit should address the concerns of non-programmers too.
An open equivalent might be wary of being accused of contributing to copyright violations since in that scenario, there is no way to force people to respect it.
Kinda the same problem as YouTube. Lots of people copy movies on the high seas, but if you are as big as yt you cannot easily get away with it.
This is not like youtube because Github is already hosting those violations and people are already inappropriately copying or including such code. It matters not whether the local inclusion was fetched by copilot or a human fetched it using more manual steps through search.
But in defense of copilot, code regurgitation is uncommon in routine use. An editor extension allowing search of github would be at least as easy to use to violate licenses but I do not think it'd be taken down since that would not be its core offering.
Copilot goes far beyond mere search and provides a useful service. GPT-3 can also be prompted into generating copyrighted works of writers but I do not see people talking as if that is its primary utility nor as much clamoring in these forums to end that service.
Well, try doing that for music or movies or proprietary leaked codebase.
If you think copilot is uniquely producing things that are not that different from humans then surely no one would have any problem with feeding it massive amounts of corporate programs?
I am not aware what writers are doing but there have been plenty of uproar regarding stable diffusion. I have a feeling that if any tools like this get built for musicians/film-makers, it will look vastly different from the current situation.
> surely no one would have any problem with feeding it massive amounts of corporate programs?
There is a similar gymnastics done by human engineers today due to the issue of patents. I don't think this is a good trend to uphold.
> I am not aware what writers are doing but there have been plenty of uproar regarding stable diffusion
Yes but mostly in the art community. On HN there were plenty of arguments just the other day how it is not the same for art and programmers have a stronger case. I disagree but regardless, the case is exactly equivalent for GPT-3 and writers but it wasn't an issue generating about a thousand comments on respecting IP and ceasing deployment of LLMs until copilot.
Are you arguing that copyright laws should be abolished? I have no problem with that as long as it's clearly defined, you can't not respect copyright of open source code but enforce it for proprietary code, just because the value is arguably non-monetary.
This is quite important, actually, and I don't think enough people realize this. I am a photographer sometimes and it would be really cool if I could share my photos online under a copyright license that forbids their use in training AI.
Or they could integrate a way to find the produced output back in the corpus if it's sufficiently close and provide a reference/attribution. Basically whatever tool a copyright lawyer would use to track down original work.
And that's just the engineering solution. The AI researcher solution would be to extend AI learning algorithms to attach attribution metadata to the learned data so that the output could already come annotated with information about the source.
But the latter is much harder to do, so maybe the engineering solution would suffice.
Edit: here's the related tweet: https://twitter.com/DocSparse/status/1581632706693079042
That assumes that the licenses of your code and the original code are compatible which often isn't the case.
A user once replied to one of my comment[0] about this with the following:
> It's not really an issue when you're a large software corporation; you already have mechanisms in place to check for license compliance in everything that ships, including F/OSS plagiarism checks [1].
IOW, from my understanding, they don't care. Big players do their own checks anyway, and small fish won't be creating problems because it's too convenient for them. Classic Microsoft (as I know from 90s).
The bigger thread can be seen in [2].
[0]: https://news.ycombinator.com/item?id=32534697
Copyright (unlike patents in general) allows for independent creation. If I sit down to write a quicksort routine, it is going to look extremely similar to a zillion other quicksort routines out there.
The other question (IANAL) is whether writing a quicksort routine is even a creative act at this point.
Which amounts to saying, humans are trained with a model that they can use to recognize when something they are thinking of producing is 'too similar' to something they have seen before.
And, of course, some humans choose not to apply that filter and go ahead and plagiarize anyway; some humans try to apply that model but get back a false negative, thinking they're producing something original when they aren't. And we have ways of dealing with humans who do that.
In the case where an AI is coming up with the work, perhaps the mistake is in relying on humans to try and apply their own trained judgement to figuring out if the result is unoriginal. We need an AI that scores work for how likely it is to be infringing on a prior copyright.
Then you use that AI to train the creator AI, and teach it 'originality'.
> In the case where an AI is coming up with the work, perhaps the mistake is in relying on humans to try and apply their own trained judgement to figuring out if the result is unoriginal. We need an AI that scores work for how likely it is to be infringing on a prior copyright.
Isn't that latter AI going to be more likely to need to contain verbatim copies of original works? Or maybe not?
This in turn (and the SFC's and the law firm's concern about the GPL) makes me think that there are several different things that people may be concerned about machine learning systems doing:
* they could allow you to access verbatim copies for "consumptive" use (like if you asked an AI a question about what the text of a chapter of a Harry Potter novel was, and it answered you correctly)
* they could facilitate intentional or unintentional plagiarism, and, in the case of publicly-available works that are published under a license, intentional or unintentional reproduction or creation of derivative works contrary to that license
* they could contain something like a verbatim representation and allow you to use that in various ways that themselves are not extracting or literally copying that representation, but where the original copyright holder might complain that the existence of an unlicensed copy inside the model is already objectionable
* they could contain representations of uncopyrightable subject matter which was learned through training on copyrighted works, which can then be used to compete with the original creators for jobs, prestige, or attention, or can be used to produce works that the original creators would have found offensive or objectionable (this case isn't supposed to be restricted by copyright at all, but that doesn't necessarily stop people from caring!)
Not only will the same measures not prevent or avoid these cases, but if you wanted to prevent the first and second situations, one of the easiest ways to do it might be to literally include verbatim copies of lots of works inside a machine learning model! (along with software specifically trained or programmed to warn you against unintentionally making uses the user or copyright holder finds objectionable ... to facilitate the exercise of "essentially, their ethical judgement", as you put it)
I'd be wary of this one. There doesn't seem much distance between this and the same claim against a human's memory of a work in their own brain. Yes, that sounds like dystopian fiction. So do some things that have already happened.
The big question is if we think that Google made a poor work of that AI and if more money and more data rich company can make a better AI that teach originality.
> It seems clear enough to me that training AIs on copyrighted works is typically or commonly a fair use under existing law, because the AIs can and commonly do learn non-copyrightable elements and aspects of those works. It's very obvious from enormous numbers of examples that current AI systems are capable of learning much more abstract features of human culture (grammar, concepts, facts, cultural tropes, and many others).
I said it already in a previous discussion, I would be very careful with comparing ML with how humans learn. To me there are still a lot of examples that show that AIs don't understand prompts (see e.g. the discussions around the "horse riding astronaut" prompts för stable diffusion et al.) and it seems like they really are just doing sophisticated pattern matching. If that is what they do aren't they themselves covered by the licenses/ restrictions placed on the "patterns" they "choose" from?
Capitalism is still very alive, and will continue to be. It's in conflict with the general welfare of the people...
Something will have to change.
I mean sure, maybe humans aren't "just" doing sophisticated pattern matching, but there are good reasons to suspect this is some part of what we are doing. (Even if its not implemented with back-prop).
e.g) consider the work of people like Anil Seth, who propose that our brain is basically a generative model of the world, which aims to minimize the likelihood of perceptual data. (see also: Karl Friston's free energy principle). What's up for debate is how it is structured, what priors are built in, what is the learning algorithm etc.
Anyway, for all their limitations, it seems clear that current artificial generative models can: 1. learn hierarchies of abstractions, which 2. explain the observed data in the fewest possible number of bits, and 3. generate new, novel data based on the patterns that have been learned
If you want to describe this as "just sophisticated pattern matching".. then sure I guess? But I think there's a clear qualitative difference between this and searching for code in a discrete database (which imo would not be okay).
The paper suggests that transformer work on a set of select, aggregate and element-wise operations. Which seems pretty close to the SQL statements i write from day to day.
And at this point, you are getting both attacked and supported by AI, and not really better off in a meaningful way.
There's no good way of solving this issue, without general intelligence, and the problems that will bring (what reason does a generally intelligent AI have for supporting us or not enslaving us).
This is why all AI research IMO is unethical. Point me to a single AI use that has not already been abused, maybe I will change my mind, as it stands though we should be prosecuting the people misusing this technology, or at least irresponsibly releasing it, as fast and as quickly as possible before we get the point we are no longer fighting bad actors but the machines themselves.
I don’t think AI research in a vacuum is as deeply unethical as you suggest. It’s about current societal context- people won’t like being worse at things; resources will be hoarded rather than distributed.
I think OP explains clearly, in many paragraphs, why it's fair use. That's literally what their whole post is about.
Clearly, training on large volumes of data is not small volumes in any sense of the word. The argument that it is fair use is itself flawed.
Or here is the real analagous question:
Fair use is about more than just the size of the excerpt.
If you write an article about good writing, and quote a choice paragraph from someone else's work to show an example, and credit that quote, that is fair use.
Is it fair use if you read an awesome paragraph, something that really is the result of the authors unique intellect and effort and craftsmanship, and makes you think "damn", and then drop that same jewel into your book?
The difference is, the paragraph isn't being included for examination or comment or transformation, it's being included to directly copy and perform it's original function as part of what makes a work a great work, and, it's not being credited in any bibliography or footnotes or directly.
The reader reads the paragraph and is impressed by your deep insight, which you never had, and the original author did.
I think all in all, this sort of copying & re-use should be allowed to happen somehow, because software is more like a machine than a novel, and humanity benefits when machines work well. There just needs to be some sort of rules around it about what gets included in the training sets and how both the input and the output are credited and acknowleged.
Right now, I think Github are simply outlaws. 100% of the output is violating the copyright of the code in the training set, because 100% of the input is copyrighted one way or another and none of it is being declared on the output. And it's allowing incompatiple sources to mix and the origonal terms to be stripped. The training set includes both proprietary and open source, and the output is being used in both proprietary and open source.
And there is no way that Github does not have this same understanding that I just described. I refuse to believe I am that special that I can see this and no one at Github did.
So they are not merely possibly inadvertant outlaws, they are deliberate knowing intentional outlaws.
I think in programming therms a useful parallel might be copying at the module rather than the statement or function level. For example, if I write some code prompts to do the following:
- validate my API key with Twitter
- solicit the input of a Twitter username
- download the up to 500 of that user's tweets
- convert the json to a dataframe
- plot the derivative of the intervals between tweets
...many of those tasks can be fairly described as helper functions, either taken directly from documentation (like interfacing with an API) or being so elementary as to be generic. If any one of these tasks happened to come from your code or mine, and the rest from other programs, it wouldn't feel like much of an infringement. If all of them came from the same body of code, it would.Actually, what the OP said is, "is typically or commonly a fair use under existing law, because the AIs can and commonly do learn non-copyrightable elements and aspects of those works". The rest of the eight paragraphs had nothing to do with fair use.
It's honestly a ridiculous argument to say that learning one non-copyrighted thing means that the regurgitation of another copyrighted thing, after stripping the license, will magically be fair use.
FTA: "On the other hand, maybe you’re a fan of Copilot who thinks that AI is the future and I’m just yelling at clouds. First, the objection here is not to AI-assisted coding tools generally, but to Microsoft’s specific choices with Copilot. We can easily imagine a version of Copilot that’s friendlier to open-source developers—for instance, where participation is voluntary, or where coders are paid to contribute to the training corpus."
This is the same argument that people use about Stable Diffusion, and it's kinda meh to me...I guess it'd be nice to allow people to opt-out, like Stable Diffusion is doing with their next versions, especially since a negligible percentage of people will do so and it won't affect the models at all. But yes, it basically is yelling at clouds. Opt-in would cripple models, and some people would make them anyways and just keep them secret, which is worse for the world. And at the end of the day, this really does just seem to me like a fair use of stuff that you've published on the Internet for anyone with a browser to look at. The AI models of the future are going to gobble the whole net up, and if you don't want them ingesting your stuff and learning from it, then you just shouldn't make it freely available.
If OpenAI/GitHub/MS really wanted to get ahead of this and head off any potential legal conflict, they could always just open source the models and weights, which would be in line with the name "OpenAI"...it would be a minor project to scrape all the correct headers to add to a license file(s), but negligible compared to the many millions of dollars spent on training.
Also, it's pointless to say "But X does Y" in copyright discussions. You never know if they license the content properly or if they infringe the rights. In the Cliffs Notes case, they might not need fair use at all, because the old works are already in public domain.
If the AI is learning to repeat text (e.g. Copilot) or images (e.g. Dall-E), then that makes it possible to reproduce the copyrighted works, so I would agree that that case is not fair use. -- It would be akin to compressing and distributing those works.
If the AI is learning patterns -- such as "muggle" being a noun that relates to Harry Potter, or that the lemma for "muggles" is "muggle" -- then that is less clear. You can avoid the situation by creating your own sentences with those terms in them, and annotating those sentences instead of the copyrighted ones. That way, the AI is still learning the same information.
Because copilots "use" of the works _was_ the learning.
So it would seem to me that Microsoft needs to apply "fair use" to copy and redistribute _the entire works_ they used for training.
In which case lack of fair use my well be the least of their problems, they are really crossing into Computer Fraud and Abuse Act territory similar to when Aaron Swartz "borrowed" MITs data.
This issue has been claimed many times and I've heard that DALL·E 1 & 2, Stable Diffusion & Midjourney all can create images that are exact copies of the training material.
This doesn't make sense considering the compression ratio of training images to model is about 1:25,000.
Further investigations I have made show that all these cases can be explained via the following:
1) The prompt included an image, so some form of image2image was used. Of course if you use an image as a base, and tell the model to stick closely to that image, the output will largely resemble that image.
2) The example was completely made up.
So far I have seen no evidence, given a text prompt, the output of an image containing some portion of any image from the training set.
Well, I don't really think you are violating my copyright. But by focusing on parts, you go down a rabbit hole of equating an element with the whole thing. This would render all collage art illegal. Lawyers and art pundits love ruminating on the uncertain legality of collage art (because it's not a binary question, so they can churn out endless articles that boil down to 'it depends'), but this glosses over 2 important realities:
1. Nobody gets sued over collage art largely because any case is doomed to end up with lawyers measuring the size of collage elements with rulers and then arguing about what small percentage is too much, and uncertain exercise few law firms wish to gamble their reputation on, and
2. nobody gets sued because collage art isn't worth very much to begin with; collages aren't valued very highly because they aren't as hard to make as painting or other art forms. 'Appropriation artists' like Richard Prince get rich and famous partly because their art is less about the image than the cultivation of notoriety for artistic effect; they are artists of scandal rather than pictures.
In general, bits of things are just not that important, and I'd argue that the same applies to code. If part of your code matches a prompt (excluding highly specific prompts like '# insert Woodson's unique XYZ algorithm here') and is then deployed in another program without alteration, isn't that most likely to be because it performs some generic function?
The first assumption is highly flawed though. Humans routinely do violate copyright law. Plagiarism is a huge problem in many sectors; un-cited direct copies of people's work in violation of fair use is a regular every day occurrence in the human world. It doesn't matter if you memorized the source material or if you transcribed it or if you copy and pasted it, if it isn't your source material and is someone else's, you've committed a violation of the law. Learning to produce original work and reproducing someone else's work is not the same thing. If an AI is ingesting and perfectly reproducing someone else's copyrighted works, it is in violation of copyright law in the same way a human would be if they reproduce someone else's copyrighted works.
If a corporation was to directly publish some copy that appears plagiarized, we'd call that plagiarism. I don't see how adding a piece of code—one that's fully created, owned, and wielded by the corporation—as an intermediary changes anything. If anything, it looks like plagiarism-as-a-service, which seems worse (at least to my eyes).
Of course, this matter is a bit confusing. Because, for example, (1) it's not always plagiarism, (2) defining what exactly is plagiarism even in the purely non-technological realm is difficult (and likely somewhat subjective), and (3) there is a lot of corporate marketing which suggests this "AI" is "autonomous" (presumably to distract from who exactly is autonomous in this picture). And of course ML art is quite useful for many things. But I mean, so are artists.
Not long ago, a lot of Silicon Valley rhetoric was that the purpose of "technology" was to free up time so that people could be more incentivized to "do what people love to do" like, for example, artistic creation. But now it seems that rhetoric was just that: rhetoric, or what was needed to be believed/said at the time.
And now at our present time, when technological "progress" has been followed a bit further (that is, when we've developed our machinery a bit further under the incentives of our present economic system), much rhetoric has conveniently shifted to something else, something largely contradictory, but again precisely to what is needed to be believed/said to continue following the same incentive structure.
Copilot is capable of going beyond retrieval and is competent at using variables, comments, types and local context to infer intention and generate appropriate code and even comment on it. Whenever copilot correctly predicts code of yours that's a novel combination of concepts, copilot has originated novel code.
For esoteric concepts, you usually already have to know how to prime it but Copilot is especially useful when it helps you bump into things you didn't know you didn't know (one way to increase the odds of this happening is to write out your thinking so far in markdown or comments. You'd be surprised how helpful and clever Copilot can be in some instances). My point here is Github isn't charging $10/month for run of the mill retrieval. My opinion is code-gen LLMs contribute value and more open versions are worth building.
That said I think the results are often useful and sometimes fascinating. We should not fool ourselves about the learning that these large neural nets do, though.
>Sure, we use the word "learn" to describe what they do, which is one word that we also use to describe what people do. But ML models are always wielded by people or corporations for particular purposes.
This is extremely important. "Learning" in machine learning is an aspirational label, not a descriptive one. People who claim otherwise either drank too much of their own Kool-Aid or are simply dishonest. This isn't just "wrong" in some taxonomical sense, this is dangerous in a very practical way. Conflating machine "learning" and human learning will inevitably lead to various kinds of sabotage of human learning.
Programmers tend to think of copyright as a Boolean valued function. Either something is infringement or it isn’t.
Judges think of copyright infringement as a real-valued function of many arguments corresponding to the circumstances of the parties (e.g. what actual damage was done?).
A human quoting a human without attribution, without any profit made or identifiable damage, returns an infringement value very close to zero. Such cases, if anyone is petty enough to bring them, are likely to get dismissed.
I absolutely agree with you that copyright is not a boolean, but I don't buy the idea that a judge will just shrug and allow infringement to continue just because there was no commercial harm.
I also think your example is just irrelevant to the case at hand. Sure, someone "performing" someone else's copyrighted words once may not be a big thing. But if Copilot is actually found to be infringing, these infringements will keep happening, over and over and over.
Bottom line is that none of this has been tested in court. I think it's great that someone is working on doing just that. Maybe the end result will be that Microsoft's use is indeed fair use, and that Copilot users have no further obligations. But I'd like to hear a court decide that, not a bunch of armchair non-lawyers (myself included) on a random web forum.
(sorry i wrecked the citation and lost where i found it)
> CONCLUSION:
> The court of appeals affirmed the district court's judgment. The court held that the Performers' use of the composition, as distinct from the use of the composer's performance, was de minimis and therefore not actionable. Considering only the compositional elements, the brief and relatively simple segment of the composition used by the Performers was neither quantitatively nor qualitatively significant when viewed in relation to the composition as a whole. Thus, despite the high degree of similarity from the actual use of the recorded composition, the scope of the similarity was not sufficiently substantial to support Newton's infringement claim.
Also I'm looking for a case where the judge acknowledged that copyright infringement was occurring, but decided not to do anything about it. From the bit you quoted, it sounds to me like it's implying that the judge believed there was a valid fair-use defense? Or even stronger, that the judge just did not believe the use was infringing at all?
... as opposed to AI. At the heart of the matter lurks a debate whether AI is an independent phenomenon which behaves in its own right, or a just a tool that's created and wielded by humans against a backdrop of clear incentives and motivations.
The argument isn't about whether or not the law deals in a absolutes - it's a basic principle that law is tested in courts through interpretation - the argument is that co-pilot can be perceived as merely a means to an end and that GitHub / Microsoft have created a massive mountain of liabilities for themselves.
Suppose we add a button to a visual studio plugin called 'Copy me a function' and when you click it, it 100% grabs some random code from github and plops it as-is into your code base.
I don't have to argue the ethics of if the button is 'thinking for itself'
> Suppose we add a button to a visual studio plugin called 'Copy me a function' and when you click it, it 100% grabs some random code from github and plops it as-is into your code base.
Personally, that's exactly how I see co-pilot. To my mind, it's a tool that sits in the same category as p2p platforms, copying machines or video recorders. They are just tools.
How, for what purposes and by whom they are leveraged makes all the difference here.
P2P platforms and those who violate copyright are routinely shut down (or attempt to be shut down) in the US. If co-pilot sits in that same space, it seems the books been written already... we know how it ends.
"In computer programs, concerns for efficiency may limit the possible ways to achieve a particular function, making a particular expression necessary to achieving the idea. In this case, the expression is not protected by copyright."
https://en.wikipedia.org/wiki/Abstraction-Filtration-Compari...
Consider the impact on innovation if Microsoft or Oracle were allowed to claim a copyright over utilitarian aspects of their works such as the Java or Windows API!
BTW, Copilot seems to be reproducing copyrightable material when the tool reproduces comments verbatim!
Specifically within AFC, note the "idea/expression dichotomy" [1] which clearly states:
"copyright law protects an author's expression, but not the idea behind that expression"
Thus, if this tools spits out someone else's code verbatim it is a definite copyright infringement. If it outputs code that is similar but not verbatim then it "could" be an infringement, at your own risk, and to be determined by the courts. Simply expressing the same idea in a different way is not a definite infringement.
Program code is naturally copyright-risky because the keyword/grammar space is constrained. It is far more difficult to accidentally duplicate verbatim the expression of one's ideas in a full language, such as English, than in C. And what of two separate programs (or constituent sub-parts such as functions) that by chance emit the same compiled binary?
Personally, for now I won't use this tool due to the risk of accidental plagiarism, and because it is a black box: I can't examine any lineage or attribution metadata to understand the source(s) of what I would then be incorporating into my own body of work. Of course I doubt I could get that type of traceability information for any other trained ML model I might use, so perhaps I need to re-examine my policies heading into the future.
[1] https://en.wikipedia.org/wiki/Idea%E2%80%93expression_distin...
That is not true. It can be verbatim and not a copyright violation if it can be shown that the expression in question is strictly utilitarian! I literally provided a quote from that AFC article that says this!
There's even precedent that prior art nullifies a copyright claim, as seen in Johannsongs-Publishing, Ltd. v. Rolf Lovland:
"Johannsongs failed to offer admissible evidence to rebut Ferrara’s analysis, so there is no genuine dispute of material fact as to his conclusions that Söknuður and You Raise Me Up are not substantially similar and most of their similarities are attributable to prior art."
And this was about music, not software, which has always sat uncomfortably between utility and expression, if only because it is some kind of writing. No one is claiming copyrights over Photoshop filter settings or other inputs manipulated by sliders or buttons!
First, let's preface this with the substantial similarity of the structure, sequence and organization as established in Whelan v. Jaslow which amongst other things says that you cannot merely change the variable names if the expressive structure of the code remains the same.
Now let's imagine 10,000 software developers who all implement Dijkstra's algorithm in C and then run it through clang-format. Aside from variable names, isn't it safe to assume that many of the implementations are going to be exactly the same?
So if it's identical it might or might not be a copyright violation and even if is totally different it still might or might not be a copyright violation... and only after spending insane amounts of time and money in the court system can you ever be sure if what you (or your AI) created has made you a criminal. This is increasingly sounding like a very very broken system.
However, the statement was only for cases of verbatim copies produced by Copilot. The AFC Wikipedia article states that "Proving copyright infringement requires proving both ownership of the copyright and that copying took place." The 3 detailed tests developed in that case appear to be "expand" the determination of infringement to close potential loopholes where, while there is not a verbatim copy, infringement is still deemed to have occurred because of "substantial similarity". e.g. someone copies a program but changes the variable names.
So where is the line between infringement and not, in cases where there is an exact copy of a code fragment? Can we still use the utilitarian defense or is that only used by the court to exclude portions of the code in the tests for "substantial similarity"?
At this point Copilot is awful at this higher-order level of abstraction but I can see a time where this is not the case!
Microsoft will have to put more work into filtering out responses that are indeed copyright violations if they want people to use their tools.
I doubt that MS will ever be held liable for the violations themselves as there is precedent in themselves and their legal department has plenty of cash to burn.
And we need to accept that and get over it, not get better at outlawing it.
Likewise, you can tell Copilot to crank out code specific algorithms written by specific people, but if you do so, you're still creating infringing code, same as if you'd taken the more direct route of ctrl-c+ctrl-v. The fact that you /can/ make the algorithm misbehave through adversarial input is irrelevant to the primary use cases which lead to boring non-infringing code completions.
Your argument just disallows discussing the problem while doing absolutely nothing about it.
If you train your dog to NOT attack random passersby and it still does, that dog is euthanized no matter your intentions.
There is plenty of space here to discuss developing tools to check for unintentional infringement. I would guess, though, that such tools would sweep up a whoooole lot of non-copilot human usage and make it much harder to deploy anything new.
So, maybe a better discussion to have here is how to make the animal safer, not the total outlawing of the animal. Single-line completions (the majority of co-pilot usage) aren't infringing. Probably true for almost-any few line completion. So, capping the amount of consecutive auto-completed code might be a reasonable 'muzzle' on the model to keep it reasonably safe.
Now my software has been assimilated into a proprietary blob. Had that blob been free, like my software within it, I would have accepted it, but it's not. It's controlled exclusively by Microsoft and OpenAI, two entities which I place no trust in.
For me the dog has already bitten. The free software I extended to an audience I believe would show the same generosity has instead been made into a proprietary product.
The "copyright" question for me is not a question of "fairness" or ability of Microsoft or anyone else to make a product. For me it's a tool to protect my contribution from proprietary business.
Basically. I dont want the animal safer, I want it free (according to the FSF freedoms).
That's exactly my complaint about Copilot. And since all code hosted on GitHub is now subject to this land-grab, my only recourse is not to use GitHub any more if I want to publish a project of mine.
Your code isn't being run by Copilot, as such; it's been categorized in a way that allows partial retrieval without the license or attribution. This might seem like a distinction without a difference, but it's kind for a ship-of-theseus problem; probably nobody is running any of your programs in their entirety, but it's very possible that bits of your code have found your way into other people's programs. How do you distinguish between contributions that are uniquely yours, and those which are just helper functions or cobbled together from other example code, eg in documentation or from a book or Q&A website?
With copilot, you are not acting with knowledge of the source and not with intent to copy, in fact I'm sure the users would have a reasonable expectation of the tool not copy-pasting existing code verbatim.
IANAL, but I'm pretty sure that intent matters a lot.
Popcorn time was also just a tool to allow you to stream data from torrents. That didn't seem to help them put up a legal defence (nor should it have, because the intent was pretty clear on that one).
And seriously, if cases exist, where the only thing a tool does (albeit via a VERY complex implementation path) is to strip a license from a piece of code and serve that code up via an API, then that really does sound like the creators of the tool are at fault.
I think the important thing to note in both dog attack scenarios presented is that the owner is responsible in both cases. Either they purposefully created an unsafe situation or they were negligent in protecting the public from their property. Whether the dog is euthanized is about preventing it from happening again. Preventing it from happening in the first place is done by making the owner liable to disincentivize it.
Personally, i don't care about the end users. If you want to read my source i welcome that. I just want the CoPilot model and system open, since it was based (in part) on my work. Otherwise they are free to remove my work.
What was your plan when someone eventually infringed on your work?
If you want people to abide by your license, you have to enforce it yourself.
Of course, but you will not face manslaughter charges in that case.
So, following the same logic, if you train your copilot NOT to infringe on other people's copyright and it still does, it should be destroyed no matter your intentions. But at least you won't be charged with copyright violation yourself.
That said, I don't believe Microsoft's actions to be benign. I think this copyright whitewashing scheme is fully in line with their old MO, purposefully creating a legal quagmire surrounding all open source code.
- whether there are any damages - what the big deal is
In my mind, the entire thread identified a lot of those, but I think someone already said that it will likely be tested in court ( and I have zero idea, which way it will turn ).
For the record, I personally think Copilot is a cool tool ( frankly, it is not that different from automated stack exchange in terms of results ). If I worry about anything, it is that the overall standards will decline even further.
If you took a bunch of copyrighted and non-copyrighted books, cut them into pieces, shuffled them all together, then picked a passage at random from a hat; "not knowing" what you are going to get doesn't mean you aren't violating copyright.
That's essentially what copilot is doing: it's taking a bunch of code - some of it copyrighted without license - and using it as a dataset. The ML algorithm then tries to pattern match against that data to provide the user with something they want. That's just copyright violation lottery with extra steps.
All the problems and confusions mentioned above is due to this concrete inherent property of the Copilot.
If Copilot was made to help rearrange [1] existing code to satisfy new or changed needs, there would be no need in such a deep and explanative analysis of yours.
[1] https://www.folklore.org/StoryView.py?story=Negative_2000_Li...
Code is a liability. Less code is less liability. New code is a new liability.
Even tools to create new code is a new and unknown liability, it seems.
Which laws are considered in this case? I understand that fair use is a US concept. For example how does that apply to my projects, published and licensed by a European living in a European country? I would expect the majority of GitHub contributors to not be based in the US, so what laws should be considered?
> Except to the extent applicable law provides otherwise, this Agreement between you and GitHub and any access to or use of the Website or the Service are governed by the federal laws of the United States of America and the laws of the State of California, without regard to conflict of law provisions. You and GitHub agree to submit to the exclusive jurisdiction and venue of the courts located in the City and County of San Francisco, California.
[0] https://docs.github.com/en/site-policy/github-terms/github-t...
The mechanisms limiting this are mostly about privacy. Not whether you can agree to adjudicate copyright or TOU in California
Some seem to assume there's some general "If it's American it's invalid" law in Europe. This is not the case. With the exception of specific laws, such as GDPR regarding privacy, this is a perfectly valid clause.
Which is what everyone is talking about here.
For criminal, there is only jurisdiction in Sweden if the crime happened in Sweden. I would need you to link me a case where Sweden criminally convicted someone for copyright infringement who wasn't Swedish and wasn't in Sweden.
For example, I am not Swedish and do not travel there. Sweden has no power to enforce its laws against me. No matter what I do, I shouldn't be able to be convicted of criminal copyright infringent in Sweden.
Additionally copyright infringement can be a criminal matter in my country and the Swedish prosecutors have certainly not signed these agreements.
CoPilot doesn't seem to be a terrible implementation, instead it seems to be relying on it operating in a grey area. So they're going for broke, to try and get wide enough adoption that it becomes a fait accompli.
Not really.
A human being learns by doing, it takes a lot of time, their knowledge is not transferable, and, above all, they buy the material they learn from (most of the time) It's not fair use, it's "I paid for the entire opera" , sometimes multiple times: different editions, movies, tv shows, etc.
Secondly, it's not true that derivative material is automatically copyright free.
It is in all honesty the contrary, most derivative work that reached popularity is plagued with plagiarism, lack of attribution, undisclosed ghost authors etc. all things that get settled with a contract or in court if the publisher thinks it's worth it.
Otherwise the publication simply disappears.
In other cases the work is licensed, so that the publisher can use someone else's IP and literally resell other people's ideas and/or change them the way they like (or the license permits), without having to create new material and take the risk that nobody will notice it.
Case in point (among too many)
https://en.m.wikipedia.org/wiki/Legal_disputes_over_the_Harr...
This actually depends on how you train the model. Techniques such as using unlikelihood to penalize plagiarizing models exists. Microsoft/OpenAI are of course aware of those techniques but have chosen not to use them. The reason why is not difficult to figure out. Because the model hasn't learned how to implement sparse matrix multiplication in C, it has learned how to spit out someone else's code with a few variable names changed. Not unlike how many CS students not cut out for software development try to pass their entry-level programming courses. Professors use anti-cheating software to catch cheating students. Such software would catch Codex too and expose it as incompetent. Hence why it is not used.
Imo it is very simple: IP law is intended to incentivize creative work, so that it remains possible to profit from one's creation in an environment where it might be easier to copy than it is to create. We just need to figure out what outcome we want to create: one which incentivizes human creations, or AI "creations" - and build a legal framework to support it.
"Old systems suck and our new system is great and it is new technology. Therefore old rules do not apply to it."
Not surprisingly, the moment crypto started gaining traction, everyone was quickly made to understand that rules do indeed apply even if it is a new a facet of finance regulations ( or in the case of Copilot copyright law ).
For the record, I am sympathetic to your sentiment, but you can't really expect existing interests to accept a major change if it happens to undermine someone's way of life and, possibly, alter the current legal landscape. And this may end up a much bigger change than expected and may finally usher in the era management always wanted.
This doesn't seem to translate well to code. You can't copyright 'rembrandt's style', which is what dall-e and co learn from analysing those paintings. But what copilot gives you is a sizable chunk of code. That code is (mostly) exactingly precise: It's not like the AI learned a style and recreates the style. It learns what you intended to do and then verbatim copies a code chunk in. I'm pretty sure the AI part comes in to determine what it is you were likely trying to do, so that it knows which code to copy. Not to generate the code out of the AI model. At least, that's how I understand it works.
That is the fundamental difference.
> OK, if you said "what word comes next? MR. AND MRS. DURSLEY OF NUMBER FOUR PRIVET DRIVE WERE PROUD TO SAY" ... same thing
Let the AI fill in the rest of that sentence and we can debate whether that is copyright infringement or not. However, if the AI system is capable of finishing that sentence, it can presumably fairly trivially be asked a slightly different question. Instead of 'finish the sentence', how about: "Suggest the next likely sentence". That system would presumably generate the next exact sentence straight from the book, and keep going and - voila you recreated the entire book.
Which is clearly copyright infringement.
AI systems have a 'volume dial' to configure how much they mix and match. Turn it down low enough and asking DALL-E for 'a girl with a blue bandana and an earring in the style of Vermeer' will just give you a copy of Girl with the Pearl Earring, reproduced sufficiently accurately that it's trivially a copyright infringement (let's leave out the notion that the painting has moved past its copyright date of course).
Point is, for copilot, the volume dial has to be kept extremely low, because you can't just mash 5000 snippets together, unless those snippets are identical. Which is its own intriguing copyright infringement conundrum (500 artists each individually paint the same thing, and they can each prove they weren't influenced by the others, thus, not copyright infringement. Then, you reproduce the averaging of the 500, which results in yet another painting. Did you just infringe copyright? Surely the answer is 'yes', but whose copyright did you infringe? All of them? Is it 'yes'?)
You have to be very careful with this line of thinking. I remember SCO versus RedHat began on much smaller premises. I also remember it took years for ReactOS to audit their code after mere suspicions arose about code that seemed to be inspired by something like asm to C translation.
The GPL is clear: derived works must be under the GPL too. The license must be respected, it doesn't matter if it was copied "from inspiration because it was learned" by an algorithm or a person.
Many of humanity's best works (paintings, classical music, golden age of physics) have been created before humans voluntarily reduced themselves to automata.
AI hasn't produced anything apart from mashing together other people's creations, usually with a somewhat creepy result.
In programming this may work because quality does not matter, only LOC and social capital with the rest of the brogrammers. The objection that therefore real programmers do not have to be afraid is false. They either have to join the mediocrity or clean up the mess that the brogrammers make (while being disrespected by them of course).
The problem here is the scale and the power that enables that scale. This is industrial level mining of non-consenting humans, exploiting their life's work in many cases.
Yes, but a human being isn't allowed to copy that work before learning from it, even if they destroy the copy afterward. AIs don't watch Youtube or browse Github. People download copies of content stored there, analyze and categorize it, and then feed it into AIs. Copyright is broken at step 1.
Luckily, someone will probably come out with a "renegade" version trained on whatever makes it a useful assistant to my coding. I won't be afraid of accidently violating copyright myself, because I won't be trying to bait it into reproducing heavily copy&pasted cherrypicked examples, and I won't use 20 lines of its output with zero modification.
What I have more of a problem with is Microsoft charging for copilot which was trained on copyrighted code without any permission whatsoever which they really have no right to utilize/charge for.
What if instead of Copilot, it was a bunch of humans who were searching all the source code they could access and then copying/autocompleting that code, regardless of the license.
Is that still OK? If yes, why?
Copying code, even if its from a mix of many different places, and the results look like a mosaic, would still be copying.
If the Mturk worker just suggested an implementation they came up with that be fine.
You’ve been authorised to see this code.
Nice you got something out of it, I'm not judging you either, but it was probably not the correct way to operate. What you should've done was notify the owners of the incorrectly configured system and left it at that.
You're also not a massive international conglomerate who should know better than to read every ones code and use it to turn a profit without first asking for permission.
I use Github like a bank, not a public library (unless I'm working on open source). I never would've allowed them to read through all my code and use it for profits without at least asking.
Illegal would imply some kind of intent or malice. I was legitimately trying to access the executed result, which I would have been authorized to do if the service was operating normally.
> What you should've done was notify the owners of the incorrectly configured system and left it at that.
Seems unrealistic to "leave it at that". I had to read the output to understand it wasn't what I expected, and once I read it I knew how to code, at least to a cursory degree. The code was simple and it was a service I used frequently, so it was immediately clear how the code translated to the results I was accustomed to. Maybe that would be harder to do that now in my old age, but I was just a kid so I had neural plasticity on my side.
> I use Github like a bank, not a public library (unless I'm working on open source). I never would've allowed them to read through all my code and use it for profits without at least asking.
I don't know what kind of banks you deal with, but banks normally do read through your banking records and use that information to sell services to their clients – notably loans, which require knowledge of your deposits to offer.
Misconfigured CGI handlers in Apache were very common in the late 90s, treating Perl as text/plain. There's no laws being broken, just a bad httpd.conf and no one is getting locked up for malicious intent.
What if you intentionally or unintentionally took down a server that controlled important infrastructure which people depended greatly on? Flood warning system for example ?
Grow up.
Either way, I still think you're in the wrong, kind of like checking out a naked person getting changed because they accidentally left their blind open. It was available, maybe it was clever, but it's a strange way to learn how to code. Why didn't you just buy a coding book, or borrow some from the library ? Was the code really of good quality if the server was configured so badly?
Obviously we have a difference of opinion and that's ok.
Sure, you can argue that they were not supposed to read the code, so they shouldn't have. But without some tangible harm I don't see why we're supposed to disapprove of it. Maybe allow some hacker spirit while posting on Hacker News :-)
From what I understand, it is not proven that the AI uses the knowledge of concepts and logic to write the new code. It is likely that it actually performs instead a very optimized stitching of code it previously saw.
Is my understanding outdated here?
From the ethical point of view, I'd say you're making some assumptions here that result in it being ethical when a human does it, and those assumptions might not hold for an AI.
For example, you're assuming a win/win outcome, where your learnings from copyrighted open source code don't harm the original authors ability to find work, or the value of their code.
With an AI I think there are possibilities we're looking at a win/lose situation, where Microsoft wins big, and maybe some other developers that also profit of their use of copilot, but where the original authors of the code that went to train it see their skills be devalued over time as a direct consequence.
In my opinion, a win/lose is unethical. What I'm not convinced is that we're looking at a win/lose, but I think there's a possibility.
I have seen this line repeated many times, but I never saw it actually explained. A lookup table is dumb and easy to understand/interpret. Deep models are not that. They are also not a linear interpolation of … something. What exactly is the claim being made here? Yes, deep models don’t generalize too well on ood data. How does this make them a “very optimized” lookup table?
Here's the real problem:
We're on a forum in which most of the participants are in at least the top 0.1% of technical ability, and yet here we are waving our hands and speculating on "what AIs think" and how "they probably/likely compute" things.
Last week I met with a director of a new "AI" research group, chuffed to the nines with a massive research grant he just landed. I'm happy for him, with only one little concern - that he knows nothing whatsoever about machine learning or mathematics and outspokenly doesn't believe that "knowledge is that important in the new reality".
Copyright infringement is an inconvenience for people who worry about that sort of stuff. Sure. But don't you all see a much more serious issue? It's bad enough that code is already so precariously bloated and over-complex nobody bothers to debug critical applications any more. Now we want to add "assistance" from tools that nobody understands.
People have special rights and responsibilities under our legal system - we don't send an airplane that has a mechanical fault to prison for a crash nor do we extend right to life to a web server. Humans have an implicit the right to learn by reading copyrighted material, machines have no such right.
Sure — in the same way that hacking into a competitor's GitHub account and copying their private source code is "genuinely useful" to you. As the person benefitting from unlawfully using their source code, of course you wouldn't care that it reproduces it. But you're not really the person we're trying to help here.
That's like comparing grand-theft auto to someone stealing a pack of gum from a convenience store. It's not a useful analogy. The latter is still a problem, but we don't need to be FUDy about it.
And OPs right, this will keep happening until we come up with better ways of solving this problem.
Whether that's educating companies on the legal (and moral) risks their developers IDE tools are exposing them to, better licensing database/indexing, working with future OSS devs building these tools instead of treating them like criminals, suing the for-profit companies like Microsoft who seek to profit from this until they invest in this problem, etc.
> I won't be afraid of accidently violating copyright myself, because I won't be trying to bait it into reproducing heavily copy&pasted cherrypicked examples, and I won't use 20 lines of its output with zero modification.
We could come up with scenarios where there might be some fancy algorithms posted on some public Git repo that's super efficient or unique, and that somehow fits into the size of individual functions that could be auto-inserted into some other person's codebase. But IRL that sort of thing is rarely ever going to be the thing that these IDE tools do. At least in a way that meaningfully contributes to another project.
That is still a concern yes, but it's still a niche usecase, which doesn't justify killing off otherwise extremely useful tools.
Maybe I'm being too techno-libertarian here, but I believe existing courts + public feedback cycles + iterating on how the public code is consumed by these tools + spreading awareness of the issue is enough to address the licensing problems.
The more accurately we explain the problem, the quicker we'll find good solutions.
[1] usually licensing saying commercial projects need to either pay or not use it at all. Or some attribution clause
I think you are, though. You have to automate the justice as well, traditional courts can't keep pace. You'll just end up with more automated DMCA-style takedowns, not less.
And I don't even see how an automated DMCA system could exist because I doubt they'd win monetary damages in court over a 'stolen' function or two (or detect it in most commercial applications in the first place).
Regardless a single class action should be enough to make Microsoft either shut down their project or adapt (via whistleblowers, leaked code, public repos, etc). And regardless if they don't adapt by investing in the possible solutions here, an OSS project could take it's place eventually and the courts wouldn't even be a useful solution.
Ideally a capital-backed company will help solve this, with the obvious legal incentives that already exist. But even if it doesn't this isn't going away.
Do you really believe Microsoft employees aren't going to be using this, illegally or unofficially?
"Yes! We (Microsoft) aren't doing anything illegal, but we are going to turn a blind eye to everyone using it illegally as we directly benefit from it - and here's the kicker! Our employees are legally liable not us evil laughter all the way to the bank"
Of course the legal execs aren't using it, this is classic Microsoft (Embrace, Extend, Extinguish).
These AI systems are highly novel, transformative, and useful. Their development is exactly the sort of thing copyright law was originally created to encourage. If it's hindering them instead, that's a problem.
(And no, I'm not saying people should be allowed to use AI to intentionally launder stolen code; use some common sense here.)
As many commenters have pointed out, no one would have a problem had Microsoft trained Copilot on the Windows source code. The fact that they intentionally left it out of the training set is a huge red flag.
Now let me flip that question around on you: What benefit would society gain from that forcing AI developers to do all that extra work?
Again, if (for argument’s sake) we want to maximize the effectiveness of the AI, why are we okay with Microsoft intentionally omitting one of the most important codebases in human history — which it unambiguously has the right to use — from its training set?
That sounds like a downside to me, not a benefit. You're basically arguing it would be better if Copilot, Stable Diffusion, GPT-3, etc (which all included copyrighted works in their training set) didn't exist. I'm just not seeing that.
Granted, people are not necessarily rational actors, so maybe you could argue it still makes sense to have some protections in place to assuage people's irrational fears. Maybe like some kind of robots.txt for determining whether a page can be used in an AI dataset could serve that purpose. I'd be hesitant to support anything more burdensome than that.
Some other uses are allowed without attribution. Someone can read and learn from open source software without needing to put an attribution anywhere. You could run an analysis of the code on GitHub to find out what percent of code is written in C++. You wouldn't need to attribute every project on GitHub.
Now the debate is whether this applies to training ML models.
They're both more like grand-theft auto, but one involves the valet driver leaving with your car, and the other involves smashing a window.
A 3 line boilerplate is neither novel nor a major part of the original.
The example cited by the OP is not a three line code sample - if you've ever done matrix coding, you know that sparse matrix operations are not simple.
Sure you can reduce it to a function call, but then you have library usage instead of code theft.
I think actually perhaps this is a way copilot could ethically move forward - instead of lifting code verbatim if it merely suggested libraries and approaches "here is an example of sparse matrix filtering and some libraries which do it", that would be both useful and ethical, presuming it does not obscure the license.
From where I sit, the complainant has found an extremely convoluted (and buggy) way to copy-paste their own code and is very upset about it. By similar logic, we should restrict the use of ctrl-c and ctrl-v, because they allow very simple infringement of open source licenses. Find a sparse matrix multiplication library which uses the copied code without attribution and you can take them to court; the law is already sufficient for this.
Even when it comes to stuff that seems reaaaaally close to pure derivative: Googling "How long does it take to boil water?" => "If you're boiling water on the stovetop, in a standard sized saucepan, then it takes around 10 minutes for the correct temp of boiling water to be reached. In a kettle, the boiling point is reached in half this time."
That's a verbatim snippet pulled directly from https://unocasa.com/blogs/tips/how-long-to-boil-water, and yet Google exists and continues to do stuff like this under the fair use doctrine despite massive efforts to attack/monetize their service. [To be fair, Google does link results, which probably insulates them because it's less hurtful to the commercial interests of the source; that said, with open source there generally are no commercial interests to hurt (open source attribution will be a tough sell as an actual commercial interest), and that's specifically called out in the law as a factor]
Copilot is even less explicitly at risk IMO, in that it never even stores the text, nor can it reliably retrieve it. I have no idea what makes anyone think it should be more vulnerable than Google.
From the copyright.gov page on fair use (https://www.copyright.gov/fair-use/, worth reading in detail for anyone who cares about this stuff, also has links to a monumental number of cases with shockingly intelligible summaries): "Additionally, “transformative” uses are more likely to be considered fair. Transformative uses are those that add something new, with a further purpose or different character, and do not substitute for the original use of the work."
Copilot without any shadow of a doubt does add something new, with a further purpose, and does not substitute for the original use of any codebase on Github (it can't create any of the codebases in full, without manual guidance so extreme that you'd have to be using the actual original codebase as a reference, so it clearly cannot substitute for a single one of them, and that's what a lawyer will argue, likely successfully).
In the Google vs. Oracle case (see https://www.copyright.gov/fair-use/summaries/google-llc-orac...), a big piece of the fair use finding was that "its value in significant part derives from the value that those who do not hold copyrights, namely, computer programmers, invest of their own time and effort to learn." and "further[s] the development of computer programs”. It's hard to see where Copilot wouldn't fall into that category, as well, and that's precedent on (multiple) appeal.
By my reading this should be a slam-dunk fair use ruling, unless precedent gets really upended, and Butterick is wasting a ton of time and effort for absolutely zero potential gain other than some bragging rights, but to each his own...I guess we all have to grind our axes from time to time.
That doesn’t matter, IMHO. Once Copilot manages to copy a work, a copy has been created and copyright has been violated. If this occurence is reasonably likely, then Copilot is wittingly assisting in violating copyright.
I do completely understand that a lot of people disagree with me morally, and think that extracting insights from scraping the public web should be illegal. You're free to have that opinion, but I'd recommend you start lobbying your congressman to change the law, because though I'm not a lawyer I hang out with a few who do copyright stuff, and I don't think the law as it stands is on your side. That said, who can say, maybe this will end up bubbling up through many layers of appeals and end up at the Supreme Court someday, this stuff is all certainly wildly outside the bounds of what anyone writing the copyright laws was thinking about back in the 70s (which I think was the latest significant iteration?) so it's fair to say it's a complete gray area.
And yet, someone has done so, and found that to be the case. Therefore, it falls under "normal use".
So make up your mind. Is software engineering only about assembling working programs, or assembling working programs + navigating license minutiae?
One of these, Copilot has a place in. The other it does not.
"I can use a gun to shoot someone" doesn't make guns illegal, even if people do so with some regularity. "I shot someone to make the point that guns can kill people" is worth even less in the eye of the law, and that is literally what you're pointing to here.
With open source code, the harm is much less tangible, since negligibly few open source projects make money from people going to their GitHub pages because they're searching for code snippets (not zero, but almost). My guess is that an honest quantification would put the lost revenue due to Copilot's existence in the tens to maaaaaybe hundreds of dollars. Courts look at that type of thing, which is why I don't think this will end up being an issue, at least in the US. Europe is wild, who knows what they'll do there, and that's where activists on this topic should most wisely apply pressure, you can always convince someone in government there to throw a spear at a BigCo. You won't take them down, but you may get them to negotiate, and I don't even necessarily think that's a bad thing.
That said, even in the US, if enough people make noise then things could change, so I encourage you to speak to your congressperson (I will be as well, but arguing the other side, because I really do think this is fair use and I'd like to see it enshrined as such explicitly, because this fight is going to be extremely common over the next few decades).
You don't accept arguments against the use of Copilot from people unless they... use it?
That's a nifty way to ignore any and all criticism of Copilot, or indeed any discussion about any ethical issue ever.
Smells like: “ I stole this lousy apple that wasn’t any good” Then why did you steal it?
Put your money where your mouth is, Microsoft, train copilot on your own code!!!
Don’t wanna train it with windows 11 code? Prefer to hijack others projects and use their for your needs and then pretend thst insulting others and calling their code worthless will get you off the hook????
Backfire
The lousy code trained copilot in what a switch statement looks like so it can autocomplete mine for me
> You don't accept arguments against the use of Copilot from people unless they... use it?
> That's a nifty way to ignore any and all criticism of Copilot, or indeed any discussion about any ethical issue ever.
"I only listen to people who agree with me, but to make that sound legitimate, I have a somewhat indirect way of saying so."
Because it's useful, then it's not a problem?
Well, it's also useful to send our non recyclable trash to 3rd world countries and every 1st World country should try it. It will definitely make the consequences less serious if everyone does it.
Not apples to apples but I guess you get the picture.
It’s a net boon for Microsoft in their efforts to rule the world.
It’s a net loss for society and ethics.
Open up copilot code, Microsoft, if you are so sure that everyone must wear transparent underwear let’s see you wearing some. Train copilot on windows 11 code. It’s not public domain.
Truth matters. Lies matter
Github make a convenient way to search and contextualise this publicly available code and paste it into your code (adjusting local scope, format, language along the way). Suddenly we have crossed an ethical line!?
Which ethical line? Are you pretending people never copy and pasted open source code before copilot? Are you pretending open source code never copy and pasted other open source code? That we were in an ethically pure world until copilot came along?
This code has different licenses. You can't just copy code randomly without checking license first.
Copilot serves it stripped of the license to unaware users. Even if copilot user wants only to reuse code licensed in a way that allows it copilot will serve him code from restrictive licenses without him being aware.
GitHub doesn’t force you to accept the license in the repository before showing you the code.
What's the harm, specifically?
Say it copies that snippet of workflow scheduling code I made at work yesterday or the greasemonkey script I made in my own time.
How is my life worse?
See your sister comment's child for my reply.
It's sort of like a power tool — sure, you could use a screwdriver, but a drill with a screwdriver attachment will be quicker. Hammers are good, nail guns are quicker. You'd never expect someone to use a drill with sd attachment if they'd never used a screwdriver before.
There are for sure things to be improved, such as the recent post on how you could put in a very specific seed and get out a specific function that it shouldn't. The answer here isn't to shut down the project, at a net loss for everyone, but to find ways to improve it.
As others have said, with Copilot gone and the new demand created, the vacuum will bring in community projects that will happily scrape every public repository they can get their hands on.
I use copilot everyday, I love it, but it still leaves a bad taste in my mouth knowing that people out there worked really hard on their code and harder on building OSS licences just for Microsoft to throw all that out of the window.
Feels like licenses don't matter anymore. My own code doesn't matter much, but it's about principals dude, licenses are there and they should be respected, if not, then it's just anarchy, and we all know anarchy only works in very specific scenarios, Microsoft is not apart of any of those scenarios.
I disagree, and this does not hold up generally: We can, and should, argue things we have not tried or experienced, like heroin and murder. What makes it so that this has to be tried?
> It's the bare minimum to make an informed opinion.
Only if the usefulness is what is in question. But it is not.
I believe the argument being made is that in _actual, real-world_ use of copilot, no copyright infringement happens. So it's not just about usefulness.
How would you know though? The burden of proof is on Copilot. Especially now that it has been shown to spit out copyrighted code.
/s
It's not that you absolutely have to have experience with something, but you'd be foolish to discount the input of people who do. In debates about drug policy I try to be polite to people with zero first hand experience, but their contributions are rarely of interest. Murder is a bit more abstract insofar as anyone who has fully experienced it by definition didn't survive to testify, but I give a lot more weight to the views of people that have first-hand knowledge of violence and crime.
It's not that you shouldn't weigh in on a topic without first hand experience, but that it's a good idea to specify the scope of your understanding, or frame uncertainties as open questions rather than assumptions.
I tried telling that it requires a credit card number to try it but he didn't believe me… I guess the thought that non-microsoft employees have to pay for microsoft stuff didn't occur.
Let's just copy each other's code without attribution.
Some algorithms in scientific computing require lots of effort to implement as nice, reusable, performant function. Those functions often more important than whatever the whole is doing because it's what most other people will be interested in using.
Really? I challenge you to write a correct email validation regex.
If it’s all just disposable code, WHY ISNT MICROSOFT TRAINING COPILOT ON WINDOWS AND OFFICE CODE?
Creative commons maybe.
Just because you, as someone who self proclaimed to have done more OSS than 95% of the commenters here, does not know how to use OSS licenses, doesn't mean that the copyright question being discussed here is a non issue.
The issue is that you don't care about what the licenses in your code mean:
> I publish my code under MIT when possible (...) please train on my work or split out my functions verbatim (...) I don’t even care about the attribution part of MIT
Having been a member of a very high profile permissively licensed project and having started a few relatively popular ones of my own, I’d say I don’t need to take licensing advice or be called “does not know how to use OSS licenses” from someone who laughably advises using Creative Commons, when Creative Commons itself advises against using CC licenses for source code, except CC0, which is entirely different from CC.
I was considering not replying, but here goes nothing....
> not because I will personally pursue every clause in it.
Then you are out of the game, because all clauses should be respected, or else you are committing something very close to an illegality when you violate said clauses. If you don't want a specific clause, consult a lawyer and remove it, or use another license.
> Creative Commons itself advises against using CC licenses for source code
So before your didn't care about respecting clauses and now you do?
I could argue: there are some bits of text in CC that make it not a good license for code, but I don't care because I don't respect those clauses. I'm not going to make that argument, because it doesn't make any sense. You either use a license and respect it or you don't.
I mostly contribute by finding issues and reporting them, does that make me less of a OSS contributor than you?
And yet I do care if my private project is used by behemoth like Microsoft without my consent, even if it's only poorly written fizzbuzz. Why? because if I wanted to share it I would publish.
Feel free to care.
If open source communities are worried about having their source code copied... then don't open the source. Keep it closed, keep it off GitHub... I mean the genie is already out of the bottle, so doesn't really matter what they do now.
You can't prompt Copilot with things like: # Function that detects spam accurately
And get anything useful/sensitive/competitive
Is there really super sensitive algorithms out there that Copilot is exposing that are otherwise unknown?
I have a very severe ADHD and as a result terrible memory, I've been working as a dev for almost a decade now and it always was almost entirely in the browser with a search bar, not IDE, and me working with infrastructure doesn't help as I encounter more than one programming language at a time, daily.
CoPilot helps me not making dozens of jr. dev level search queries daily as I can formulate query right in the IDE and it will fill in the basic algo's, data types, language abstractions that I know exist (as I use them all the time), but I don't remember how to actually invoke them, despite doing so just an hour ago. This is the most useful part of copilot, not the complex and very specific code.
What? You left out the second line of the quote. It never reproduces copyrighted content for me because I'm not trying to bait it into doing that.
If you aren’t paying for the product, you are the product.
Alarm bells of this magnitude haven't been rung about people torrenting films for decades; it's a given that some people are just going to do it and there's little that can be done to stop it.
Producing new data from original data of questionable lineage makes the questionable acts visible. Copilot and the like actively encourage this creation.
If it were possible to peek into the rooms of everyone who downloaded a torrent to admonish them then maybe pirating would have been made a modicum more taboo. But those consumers never intended to leave their rooms. Copilot forces them to leave their rooms if they want their derivative work to be used.
Google does do this though. Just Google for an easy to answer question, like “when was George Washington born”
The same applies to all realistic use-cases for copilot by the way. Whatever is produces is not copyrightable.
Correct.
>The same applies to all realistic use-cases for copilot by the way. Whatever is produces is not copyrightable.
That's a pretty bold statement to make. How do you know how people use copilot? Also IIRC Oracle vs Google essentially determined 3 lines of code can be copyrightable. So I think you statement fails on two points, you can't really predict how people use copilot and you cannot predict what a court would decide is copyrightable (this is much less straight forward than statement of facts).
Google Search is ethically acceptable because for the most part website creators like being in search results and are "compensated" in the form of more visitors, and if they don't like it they can easily exclude themselves. Website creators famously do NOT like it when Google indexes their content and then serves it up independently.
When closed source code leaks: "Copyright infringement by criminals".
What's the difference? There's plenty the world could learn from the source code of Windows or GTA6 and having access to the source of these large projects would move society forward faster. So why are OSS contributors protecting their rights "Anti progressive luddites", while the large copyright owners who guard their proprietary code like a dragon guarding gold are let off the hook?
The value that is being derived here is in the curation of the material, not the material itself.
You don't do this. You get my books, cut out the bibliography, glue all the pages together, and then sell the book as your own.
It is my book and all you did is derive some work from it.
Curation companies have the same problem and there are plenty of high profile lawsuits about it.
I mean I wouldn't expect it, but I think I'd be pretty annoyed if you didn't ask permission and then made a bunch of money off my image. It's easy to find stories from the subjects of famous photographs who feel like they've been exploited. Just off the top of my head there's Afghan Girl, the kid from the Nirvana album, Harvard's collection of photos of enslaved people, and Henrietta Lacks is sort of a similar case.
> if someone compiles a list of the best restaurants in the world and sells that list, do you think the restaurants should be compensated?
No, but here's a better example: you make friends with a bunch of food critics, collect their thoughts and opinions and favorite secret spots, and then publish a book based on that stuff without ever telling them what you were doing or compensating or crediting them.
I'll give a concrete example: I was rock climbing recently and met an old guy who was sort of the local expert, and he told me how some other non-locals had come in and kind of mined him for information about the area, all the routes, etc. and then published a guidebook without crediting him at all. He felt pretty upset and exploited by that, and I felt bad for buying the guidebook because I had assumed it was written by some local climbers and didn't realize they got most of their info from someone else.
It's not illegal, but it is unethical.
The people painting Microsoft as a big, greedy trust conveniently ignore that Copilot would actually be empowering the ecosystem of tech companies to develop services that compete with Microsoft faster and more easily.
Long live Copilot. It’s an amazing product that shows what we are capable of thanks to crowdsourcing and bleeding edge technology. We live in the future, and progress never remembers those who tried to stop it.
A bit ironic that Copilot itself is not open source.
If they open sourced Copilot then it would probably comply with most of the licenses anyways.
Like at the very, least, respect the licenses, that means, give attribution, and provide your own source as open source and under the same terms.
Open source is what allowed this progress in the first place, and the way I see it, commercial interest is actually simply trying to slow it down by keeping it behind proprietary trade secrets.
B..but, Copilot isn't open source though?
Seems very similar to a "tragedy of the commons" type of situation.
I feel like the title of the article is literally written for you: """Maybe you don’t mind if GitHub Copilot used your open-source code without asking. But how will you feel if Copilot erases your open-source community?"""
If you want to keep having useful tools based on open source code in the future, it is in your interest that people still want to write open source code. It is still too early to say how much of a chilling effect projects like Copilot will have on that. But clearly many (just read this comment section, myself included) are having second thoughts.
Maybe you can exercise some discipline when using copilot, but what about your coworkers? Many companies might not want their employees to insert copyrighted codes into their projects accidentally.
They can opt not to pay for this entirely voluntary service.
Plagiarizing is already understood to be a genuinely useful practice.
Copying a passage here and there when making a new work? I don't think the courts have ever ruled that plagiarism.
I get that using function names is an obvious way to get copilot to generate contested code, but has someone tried to get copilot to generate contested code in a way that users might sincerely be using copilot for productivity ? How do you get to the claim it is the "only way" ?
It enables the large scale theft of code. It completely ignores licenses. There are plenty of open source licenses that allow code use with proper attribution yet Copilot doesn't (and probably can't) figure a way to comply with all of them.
Copilot, as the article suggests, is a marketing stunt. To me it's more than that. It's Microsoft pushing the boundaries of law using it's money muscle again. I have been screaming from the rooftops that VSCode was just M$ spyware and I get legitimately made fun of for it. Now we have Copilot as well and they aren't even hiding it.
To address your point more directly if you do any work for monetary gain Copilot is a defacto liability. You can't just take 20 lines of completely stolen code, modify a few things, and call it your own. That's why legal reverse engineering has an entire black-box method of development. QA, researchers, and developers aren't even allowed to talk to each other directly.
I hope Copilot does get shutdown. Along with everything like it. It is one thing to have an AI trained on your workplace's code, or specific code following specific licenses, but the blatant theft of not only licensed open source code, but also private code, is a terrible precedent to have.
Because most software piracy is of saleable software. 20 line snippets are not for sale, mostly because the transaction costs are higher than the snippet value.
Seems like a pretty tangible harm to me.
Now I hope that Copilot sticks around for this exact reason: cause endless inane lawsuits claiming that actual original code was stolen and laundered through Copilot or reverse engineers going cowboy. Make it enough of a problem clogging the courts that they start dropping copyright cases.
Dropping copyright cases is great if you hate proprietary code. That's fine. Open source is also powered by copyright. If we start dumping copyright cases we don't get the "well bob we may as well open source it!" We get large companies like Google just completely ignoring copyright and using open source code without attribution. What you've described (causing enough a problem in the courts) is the exact purpose of GPL. If we start dropping copyright cases altogether the open source movement may as well be dead in the water.
That's exactly what I said in my comment. I wouldn't take 20 lines of code, since it wouldn't actually work. Even if it was able to spit out 20 lines of correct code, they would be tailored to my codebase, and not violating copyright.
The only time you see CoPilot violating copyright is when someone coaxes it into that, in a completely empty codebase with no context. The violation of copyright is not possible when it is used as intended.
You don't have to bait it.
>I won't use 20 lines of its output with zero modification.
Depending on changes that may still be derivative work. The entire concept reminds me that can I copy your homework meme.
How do you know that this is the only way for copilot to reproduce copyrighted code?
Do you also just rip any open source code, violating the licenses? Nice.
If you don't know Microsoft's history, a lot of what more informed people are worried about seems overblown. Copilot was Microsoft's first test of people's trust after the GitHub acquisition. It's going very, very, very poorly. There were ways to do this with consent and collaboration with the people and projects it takes code from, but they're acting like classic Microsoft here.
Too many people are focused on what's legal. It's fine to think of, but law is the last stop before the breakdown of society. Microsoft skipped society and went straight to sparking an inevitable test of and possible reshaping of copyright law.
Maybe it's illuminating of a trait of human nature. On the stable diffusion webui repo many people have stated that they would continue to use the code even if it were stolen or unlicensed. These people aren't a part of a corporation; they are average netizens handed a technology essentially indistinguishable from magic with nothing in place to prevent its use.
If the tech is simply so impeccable as to be irresistible then a higher order framework needs to be in place to teach people not to bite because they will be bitten back.
Or maybe they do know about it, and don't agree with you. Do you allow for such an option?
https://github.com/features/copilot
"What can I do to reduce GitHub Copilot’s suggestion of code that matches public code?
We built a filter to help detect and suppress the rare instances where a GitHub Copilot suggestion contains code that matches public code on GitHub. You have the choice to turn that filter on or off during setup. With the filter on, GitHub Copilot checks code suggestions with its surrounding code for matches or near matches (ignoring whitespace) against public code on GitHub of about 150 characters. If there is a match, the suggestion will not be shown to you. We plan on continuing to evolve this approach and welcome feedback and comment."
Notice that Copilot often gives code that verbatim matches opens source software, even when that filter is on. For example: https://twitter.com/DocSparse/status/1581461734665367554?s=2...
Their approach of "matches or near matches (ignoring whitespace)" is clearly inadequate, and it's honestly insulting that they think this is enough. Even if Copilot just changed the case of a single letter, their filter wouldn't catch it.
I saw a few examples, but I don't see how that extrapolates to often. It's quite possible I've missed something in the article since I kinda skimmed it. :)
>and it's honestly insulting that they think this is enough.
They don't. - "We plan on continuing to evolve this approach and welcome feedback and comment."
That is corporate speak for "we plan to do nothing about this".
Bottom text provided by copilot.
Copyright laws should change if needed, but this is not the process.
I would love to see a lawsuit which requires GitHub to provide their full Copilot dataset.
But of course if the data is already sitting in object storage inside your cloud environment and all you have to do is run some MapReduce jobs to get at it...
Hence: unfair, anticompetitive, intellectual-property-right-abusing behavior. Microsoft GitHub (tm) can prevent anyone else from running the kinds of analysis they do by simple "operational security", while running literally any kind of analysis, model training, etc. they want. Don't like it? But their commercial services and products so you can run Microsoft GitHub (tm) on your very own Microsoft Azure (tm) infrastructure, using Microsoft Visual Studio Code (tm) and Microsoft GitHub Codespaces (tm) so work on _your_ code privately.
Best of all, you can still still take advantage of the huge library of "free" code offered by Microsoft GitHub Copilot (tm) to ensure your private, proprietary codebase still has all of the advantages of Open Source Software, brought to you exclusively by the Microsoft GitHub Platform (tm).
ArchiveTeam has a distributed Github archive project[0]. It's unclear what the status is right now. It seems like a worthwhile idea.
[0] https://wiki.archiveteam.org/index.php/GitHub#Archive_Team_p...
I’ve also built many parallel repo downloaders for CI reasons. You can clone repos all day pretty much with little rate limiting. I haven’t pushed parallelism past 64 per host though
That's not unfair, and not the basis for a lawsuit, it's just business.
> Copilot is not only making money off of open-source, they are making money off of open-source in a way others can't.
Of course! That's why MS paid squillions to buy Github.
most monopolies are completely legal. small town with only one gas station or one grocery store? 100% monopoly, 100% legal.
there are plenty of alternatives to github, paid, free, as a service, and self-hosted. GitHub has a large market share, but no monopoly.
This isn't a company with plenty of goodwill in its sails launching a hip new product. They're not Do No Evil era Google or Apple riding high on the iPod. We can't just pretend there isn't a history.
I posit that they have indeed learned from their mistakes.
1) They train Copilot with repos hosted on github.com. 2) Users who upload code to github.com grant GitHub an explicit license[0] granting GitHub the right to show that code to others. 3) GitHub do not specify which technologies or techniques they may use to show this code to others, meaning they may use any technique they like.
Microsoft have learned. And they have covered their collective asses. Users who host code on github.com agreed to these terms.
Users who don't like their code showing up in Copilot should not be hosting their code on GitHub, because they agreed to have their code delivered to others when they signed up.
[0]: https://docs.github.com/en/site-policy/github-terms/github-t...
> This license does not grant GitHub the right to sell Your Content. It also does not grant GitHub the right to otherwise distribute or use Your Content outside of our provision of the Service[...]
I imagine the crux of this case will be what constitutes "the Service", and whether that includes Copilot. Also whether licensing Copilot counts as selling Your Content.
Which is to say: Microsoft was taking specific action which wasn't a natural consequence of their software, but was being actively enforced to keep out competition. Vendors didn't organically discover consumers weren't interested in getting a PC without Windows and IE, they were prevented from even offering the option lest they be completely denied the ability to offer that at all.
they are not threatening to block access to github.com for all Comcast users if Comcast chooses not to block access to gitlab.com, for example. That would be an illegal activity for a monopoly. simply existing as a monopoly is not itself illegal.
Citation needed. I'm interested to know of examples of how they have stifled competitors. AFAIK Gitlab came of age well after Github had already grown super-large. If Github could have killed Gitlab, wouldn't they have done so?
"GitHub has around 56 million users, whereas GitLab has over 31 million users." - https://radixweb.com/blog/github-vs-gitlab
That is also excluding Azure, Bitbucket, AWS and the plethora of other git repository hosting companies and services. AFAIK, gitlab has more payed (AKA, private or premium) customers than does Github (though I lack a citation there, but that was my understanding when researching this a couple years ago on where to host a private companies code. My take-away then was that Gitlab is actually more popular amongst private companies and Github is more popular for open source projects).
What would be considered Monopoly marketshare? I would suggest that Facebook, Twitter and Amazon have monopolies. If you get shut down as an Amazon seller, that can be 90% of your revenue. Given so many people can leave, and have left Github voluntarily, is implicit evidence that there are very decent alternatives (if not, there would be no alternative and you would be forced to stay with Github for lack of alternatives. That is not the case though, it's easy to just go over to Gitlab. From my perspective, I have no idea how Github can kill Gitlab, I'm curious if there is a vector where Github could use it's network effects to diminish Gitlab, as a specific example. So, how could Github do that?)
Does the law really say you can't include a free web browser if someone else created a paid one?
I don't want to live in a world where potential improvements for consumers get companies sued for antitrust.
https://en.wikipedia.org/wiki/United_States_v._Microsoft_Cor....
Yes. It's called the Sherman Act, and it's the basis of anti-trust enforcement in the US.
https://www.ftc.gov/advice-guidance/competition-guidance/gui...
<<<The Sherman Act outlaws "every contract, combination, or conspiracy in restraint of trade," and any "monopolization, attempted monopolization, or conspiracy or combination to monopolize." >>>
I know lots of people here don't like it, but it is the law and that was the question; "this" in parent clearly meant "use dominance in one market to gain dominance in another" in grandparent, regardless of whether that's actually the central issue here or not.
For instance, Apple already knew about the portable hardware market and extended their reach into the portable music market via iPod. It used the iPod to reach the music marketplace via iTunes. Used its market dominance to create iPhone and the rest is history.
Maybe that's a personal maxim on your part but there's no law that says if you have dominance (what does that even mean?) in one market you cannot enter another. Think about what you're saying. We'd have no multi-product company if that were the case. What does "dominance" or "market" mean anyway?
As you point out, establishing what dominance is in a market is a tricky thing and it is why governments will investigate any potential signs of it. Now, personally I don't think this is really an issue for Github, or microsoft (the parent company) but let's not pretend market dominance and abuse of dominant position are not a real things.
TLDR: If you have a monopoly in a market, you can't use your position to get a heads up in another distinct market by bundling products together. Ex: Microsoft can't use their monopoly position in the OS market to get an advantage in browsers by bundling IE with Windows.
It's my understanding that this doesn't apply if you don't have a monopoly, and it also doesn't apply unless you're actually bundling the sale of multiple things together.
Doesn't seem to be relevant here IMO.
>or instance, Apple already knew about the portable hardware market and extended their reach into the portable music market via iPod. It used the iPod to reach the music marketplace via iTunes. Used its market dominance to create iPhone and the rest is history.
Look at what you're saying here. Apple had domain experience in hardware and launched a new product (good). Then it used the dominance of that product to muscle into an entirely different market (bad). And the combination of hardware and market has led to Apple being able to extract their tax on half the music market, or whatever they have.
This is exactly what we don't want and why anti-trust laws exist.
Just because there is some existing competition which has a few percent marketshare and technically it's not a monopoly doesn't materially change anything besides a pro forma excuse. Which is why Google have been propping up Mozilla, they want the excuse "but technically there's another browser". However for consumers and the market it doesn't matter that technically there's an option that practically nobody uses.
If we define "open source" as "you can't necessarily use this to train an AI", then Copilot itself is illegal because it's using code without permission.
If we define "open source" as "you can use this to train an AI", then Copilot it legal, but GitHub may be illegally misrepresenting itself as a host for "open source", as the policy it hosts code under isn't truly open-source.
If we define "open source" as "you can't necessarily use this to train an AI" but then GitHub's policy explicitly states "by using us as a hosting provider, you give us permission to use your code to train AI" then they are in the clear. But I doubt they have that clause or at least had it when Copilot was first revealed.
>The Google BigQuery Public Datasets program now offers a full snapshot of the content of more than 2.8 million open source GitHub repositories in BigQuery. Thanks to our new collaboration with GitHub, you'll have access to analyze the source code of almost 2 billion files with a simple (or complex) SQL query.
https://cloud.google.com/blog/topics/public-datasets/github-...
I'm pointing out that this limitation is not meaningful because everyone can access all GitHub hosted source code through BigQuery, where they won't be rate limited.
I'm not comparing BigQuery to Copilot.
Their freemium product is useful to many open-source projects and communities, but you do not have any more rights to use Microsoft's GitHub than you have to Microsoft's Windows.
Is it? Data storage would be prohibitive, but I can see ways to download the entirety of Github in a few weeks/months (assuming my size estimate is accurate).
https://www.gharchive.org/ http://ghtorrent.org/
also available on GCP as a dataset, provided by Github itself: https://console.cloud.google.com/marketplace/details/github/...
I pay GitHub / Microsoft to host my code, and that's all I expect them to do with it, host it, as securely as possible. It sounds like Microsoft are doing more than this so what's your actual deal...if you have one?
GitHub considers the contents of private repositories to be confidential to you. GitHub will protect the contents of private repositories from unauthorized use, access, or disclosure in the same manner that we would use to protect our own confidential information of a similar nature and in no event with less than a reasonable degree of care.
"You have bested me in this epic battle of the minds."
Yeah, Microsoft has bamboozled users for years... that's the whole point.
Like someone else said, there was a version of this where they asked people to opt-in and got community involvement. In true MS fashion, they just did it without asking and people are rightfully pissed.
ML training needs to be fair use of copyrighted works, or most machine learning and AI projects will be impossible.
No, learning (by ML) is not learning (by humans). It's just a same word used in different context, and doesn't by itself imply that the meaning is the same. The underlying process is completely different. Neural networks, despite their name, don't share anything in common with human brains.
Do you have any other argument besides "the word is the same"?
> Neural networks, despite their name, don't share anything in common with human brains.
They share enough. Current ANNs are different than human neurons, however, the principle of learning they are operating on is very similar.
Machine Learning is learning. You can't look at ImageNet or just about any other model and say otherwise.
I'm informed on how current artificial neural networks work, how they are made, what they are and aren't capable of, etc. Thanks.
Human learning, at the very least, consists of maintaining internal model of the world around us and integrates all data into a coherent structure. Current neural networks are extremely limited in their capabilities - each model operates on a strict class of data, and is not able to comprehend the world outside of that class of data, and as such cannot be compared to human understanding. CoPilot does not understand how the computer works - it only understands that connections between pieces of text exist.
If you're as informed on how ANNs work as you claim to be, then perhaps you should inform yourself more on how human mind works.
I believe that the vast majority of code that copilot produces is fine. But we have also seen clear examples of copyright violation.
The biggest problem is that it is basically impossible for the user to tell which is which.
This whole website though and "investigation" is some sort of sensationalism that seems ultimately aimed at fair use training, which is very unfortunate. It seems like it involves an ego trip and popularity-chasing with targeting a megacorp as a means to that end.
Frankly, I don't care that ML/AI _needs_ this to work. That's not my problem. You don't get to circumvent existing agreements (and law) because you believe that ML learning is the same as a human reading a piece of code and then typing it up on the side. Tesla manages just fine by generating their own training data. Other businesses have found partners to acquire data from. The only reason this isn't being immediately addressed is because there is near-zero accountability for license violations in software companies, and ML further obfuscates that.
And yes, it is the same thing as learning.
If you have a robot that learns like a human does … you think it should be illegal for that machine to look at GitHub? To watch a Hollywood movie?
Yes.
Now, wanting to minimize the impact of robotic competition on human wellbeing, understandable. But the means to that end is declining to recognize property rights of those who who try to privatize the commons.
It should be illegal to reproduce copyrighted material, but not to “read”, “view”, or “consume” it.
Luckily this is what the law already says.
Good riddance then.
But the moment you start reproducing more than a few lines of prose without attribution I guess you are in for a nasty letter.
I'm probably not the first person to say it through this whole debacle, but that might be me: https://news.ycombinator.com/item?id=33242619
>> "Copilot was Microsoft's first test of people's trust after the GitHub acquisition. It's going very, very, very poorly. There were ways to do this with consent and collaboration with the people and projects it takes code from, but they're acting like classic Microsoft here."
Copilot doesn't help people intentionally launder copyrighted code. It may cause people to accidentally use copyrighted code without realizing it. They're still liable.
This is absolutely incorrect. Independent creation is a complete defense to copyright infringement. Funny enough, Learned Hand gives a near identical example to highlight the opposite conclusion ("if by some magic a man who had never known it were to compose anew Keats's Ode on a Grecian Urn, he would be an 'author,' and, if he copyrighted it, others might not copy that poem, though they might of course copy Keats's").
You wouldn't be able to claim independent creation though by reproducing a work with Copilot.
The latter is obviously a violation of copyright, full stop.
The former, to me, is obviously not a violation. If it were, that would massively tilt the playing field in favor of large corporations. It would become very hard to independently train your own models. Philosophically, I go by the principle that if it's (il)legal to do yourself, then it should be (il)legal to do the same thing with an AI's assistance.
The massive complicating factor is that nobody knows how to do (1) without also doing (2) as a side effect, because we don't understand how deep learning works well enough to control it.
I'm not sure I completely agree w/ the comment (nor do I think it vindicates CoPilot), but I think it does provide insight into why CoPilot is violating copyright.
The only analogy I can see is that copying the code internally to use in CoPilot training could be a violation of copyright (like how backing up your own MP3s is a violation of copyright?), but the licenses on these public repositories probably already allow that...
https://docs.github.com/en/site-policy/github-terms/github-t...
this is separate from the license you specify in the repository and you can't revoke it without removing your code from github.com.
GitHub shows code to those who wish to see it. it is up to the person using that code to use it according to the license. when I buy a car, it is up to me to use that car according to the law. when I buy a gun, it is up to me to use that gun according to the law. etc.
And yet we (modified) the law to mandate speedometers and seatbelts to make you more aware of the speed and more secure against failure. We require car companies to perform thousands of crash tests to validate that the tool the give you is safe for when you inevitably push “according to the law” a little too far.
We mandate mirrors and backup cameras because we know that those who intend to follow the law closely still have blind spots and it’s in everyone’s best interest to mitigate and increase awareness.
> when I buy a gun, it is up to me to use that gun according to the law. etc.
And yet few laws have caused the US (and other nations) to question this principle quite like gun laws.
Gun laws are really both a perfect example and the worst example of why we’re having a debate around CodePilot. We both expect people to be responsible for their decisions (you need to verify legality of that code snippet before using) while also giving them the notion that they can toe the line as much as possible (why regulate the availability of dangerous tools, crime is illegal, users won’t make a mistake).
Guns are used to kill people despite it being illegal. That’s why people want gun control laws. And in a comparison I never expected I would make, perhaps people want AI to be regulated because it will be (is?) used to circumvent copyright.
Edit: I don’t know if I really have a side I stand on in this debate overall, but I think the argument for why it’s copyright violation today is pretty compelling. We wouldn’t make the progress we’ve made without this violation and perhaps the loss of copyright is a worthy sacrifice?
I still don't see how there is any footing for a copyright infringement claim here, given that users who put public code on github.com explicitly grant GitHub a license to use that code to provide services to other GitHub users.
that license grant is above and beyond what any specified license terms the repo itself grants to users of the code.
you literally grant GitHub the right to do this when you put your code on github.com.
Relevant snippet:
> you grant each User of GitHub a nonexclusive, worldwide license to use, display, and perform Your Content through the GitHub Service and to reproduce Your Content solely on GitHub as permitted through GitHub's functionality (for example, through forking). You may grant further rights if you adopt a license.
They key parts are the “through GitHub” portion. GitHub is being careful to not give people rights to your content beyond the right to view it through GitHub. Performance refers to multimedia like music and video assets (according to others parts I didn’t reproduce).
No one is gaining a license to use your code through the inclusion on GitHub.
Section D is the relevant section.
https://docs.github.com/en/site-policy/github-terms/github-t...
>We need the legal right to do things like host Your Content, publish it, and share it. You grant us and our legal successors the right to store, archive, parse, and display Your Content, and make incidental copies, as necessary to provide the Service, including improving the Service over time. This license includes the right to do things like copy it to our database and make backups; show it to you and other users; parse it into a search index or otherwise analyze it on our servers; share it with other users; and perform it, in case Your Content is something like music or video.
>This license does not grant GitHub the right to sell Your Content. It also does not grant GitHub the right to otherwise distribute or use
I would say they would have a pretty hard time to justify using the content for AI training (and selling) based on that license. Copilot didn't exist at the time when many agreed to that license, so an argument saying Copilot is part of the service would be difficult to pull off. Moreover they don't even provide copilot to people hosting on GitHub.
Note that MS themselves are not claiming that they are allowed to use the code due to their terms of service. They claim they can do it due to fair use.
if this is indeed how they charge for Copilot, and I don't know if it is or is not, then they will need to show that they have done their due diligence in making sure that code is not reproduced verbatim when a user requests that it not reproduce code verbatim.
I'm quite sure that GitHub can defend Copilot in court. That's part of the process of offering a new feature to customers; making sure that it is legal and defensible to do so.
All of the armchair attorneys here who think they know better than GitHub's attorneys when operation of the service puts GitHub's ass on the line is ... I wish I had 1 percent of that confidence. I would be a thousand times more confident than I am now.
I don't think you can control it. Machine Learning models do not create anything, they make a prediction of the expected outcome based on the training/validation data. Similar to how human beings are an outcome of their experiences, so are ML models. Ofc human beings are much more complex than a ML model.
> The latter is obviously a violation of copyright, full stop.
It's not obvious to me that (2) is a violation of copyright. Unlike patents, copyright violation is not as simple to prove. My understanding is that, at least in the US, independent creation is a valid defense against copyright infringement. For example if 2 people independently write the same story and can prove that they did, they can both hold copyright over that story.
The analogue to this does exist without AI, when creating something that looks like copyright infringement, clean room design (don't look at similar things) is often done to ensure that "independent creation" can be used as a valid defense in court. Given that, I think (1) is probably not safe to do at all if you can't prevent (2).
Outputting copyrighted material is a violation of copyright, period. Whether that violation is enforceable depends on your means though.
This is the theoretical case but I don't think I've ever seen that actually happen in practice.
> > The latter is obviously a violation of copyright, full stop.
> It's not obvious to me that (2) is a violation of copyright. Unlike patents, copyright violation is not as simple to prove. My understanding is that, at least in the US, independent creation is a valid defense against copyright infringement. For example if 2 people independently write the same story and can prove that they did, they can both hold copyright over that story.
> The analogue to this does exist without AI, when creating something that looks like copyright infringement, clean room design (don't look at similar things) is often done to ensure that "independent creation" can be used as a valid defense in court. Given that, I think (1) is probably not safe to do at all if you can't prevent (2).
I don't think the analogue holds, the AI does have direct view of the actual code. In the most paranoid clean room design you have two teams, one analyses the behaviour of some software and writes a specification (without view of the source code), the other then uses that spec to write the reimplementation.
Copilot turns that on its head, you ask to do something it then looks up the source code how to do it and gives that to you.
The only way to prevent all uses of your code is to keep it secret.
If anyone wants to say me using copilot violates their copyright, then sue me. But if you have no loss of reputation or revenue, and I have an innocent infringer defense - noone can stop me.
Someone recently said most statements on HN should automatically get "in the US" appended to them due to how US centric many of the views are. This is an excellent example. There are plenty of juristictions where "fair use" doesn't exist.
Statutory damages still apply.
> I have an innocent infringer defense
There's no such thing. It's a meaningless phrase, in a legal sense. Claiming that infringement was accidental or unintentional is not a defense. It has no effect on a determination of guilt or innocence. All it affects is the penalty.
Fair use is a defense, but a more limited one than you seem to believe. The usual formulation is "criticism, comment, news reporting, teaching, scholarship, or research" but none of those apply to Copilot. Fair-use claims are also not accepted by default, but only by demonstration that the four factors defining it are all applicable.
Also, since you've brought it up recently, you as the defendant in a copyright case don't get to choose jurisdiction. Usually the plaintiff does, either because it's explicitly defined in the same license that grants anyone rights at all or because it's a place where they do business. If you live in a different jurisdiction that might affect whether the plaintiff or court can collect any penalties, but not whether any are assessed. Having yourself declared persona non grata in multiple jurisdictions doesn't seem like a good long-term choice.
https://copyright.columbia.edu/basics/fair-use.html
Instead of "flooding the zone" with dozens of comments offering nothing but the same few (false) claims - strong echoes of a recently banned user BTW - I strongly suggest you actually read up on copyright and fair use. They're not whatever you want them to be. Courts are unimpressed by your towering intellect.
from the comments here, if I push copilot into giving me code that I would have written for a given problem and that code violates licenses, then who is responsible for the copyright violation? co-pilot for giving me code that looks like copyrighted code or me for tweaking co-pilot commands to give me the code I envisioned which looks like copyrighted code?
also consider that the very tools used for solving problems in code lead coders to a small number of solutions for a given problem. is it plagiarism or parallel original thought?
also consider that when I wrote code, if I was solving a similar problem to what I solved before, I recreated that previously used code fragment (or larger) and use it to solve the problem at hand. I had zero issues leaving a trail of duplicate code behind me especially if the code was a major part of a software patent.
I didn't care, my code was lauded for it's readability and reliability. reuse the same concepts in multiple variations, you get real good and writing code correctly.
maybe co-pilot like programs could scan existing code bases and find examples of code fragment plagiarism with the goal of showing that software copyrights are useless.
I wonder when we'll hear about the first big hack that gets traced back to production code pushed live after CoPilot "suggested" eval(base64decode({webshell}))
(Hell, I've _been_ that incompetent management and engineering leadership at various times over the last few decades...)
It's been a massive productivity improvement to our senior devs, and I got so used to it that it's an annoyance when Copilot doesn't respond.
It would be because it's illegal and violates the licenses, desires, and intentions of the thousands of workers who wrote the code in its corpus.
You act like Microsoft is trying to do a public service and people are angry about it. The reality is that they're taking billions of hours of work and using it to build a product that only they control.
If they re-released Copilot as FOSS, a lot of the valid criticisms would evaporate.
I wonder how many people on HN would be on the side of the creators if we were talking about content created by Walt Disney and whether pirating was ethical?
Disney used its power to distort copyright laws in a self-serving manner. As an individual you don't have equal power to oppose them (you were supposed to have in a democracy, but lobbying is legal and corporations are people).
Disney is a huge corporation that won't even notice if you pirate a movie, which you may not even have been able to pay for anyway, because of their region-locked twisted maze of distribution and DRM.
OTOH you may be screwed if you're a creator making a living from your work, and a big corp can just take it without paying, launder it through "… in the style of $YOURNAME" query, and say they own it now, because unlike your copyright, their Terms Of Service apply.
Even people who think copyright shouldn't exist may rely on using copyright — against itself. You can't unilaterally say "I don't believe in copyright", because the law doesn't care, but if you license something as copyleft, then the law does care about your anti-copyright license.
People probably would have less of a problem if microsoft breached the license on 100 year old code.
But life of the creator + 70 years?
Reminds me of this:
https://arstechnica.com/uncategorized/2007/07/research-optim...
Without copyright, it would be perfectly legitimate and legal for someone to not follow a license (because it would bear no legal weight because of the lack of copyright).
I would argue that open source is best served by strong copyright protections that allow the people created the software to make sure that further changes to it are released back to the community. Weakening copyright law means that it is that much easier for big companies to co-opt some software and not need to release the changes back.
As long as Microsoft can and will wield those laws against me? Darn tootin'
This is also why I fully expect Steamboat Willie to fall out of copyright protection in January 2024 - right on schedule. There's a few countries that have supra-EU copyright terms, but none of them are dealmakers. Nobody is demanding we match Mexico's life+100 terms, for example.
I am 100% on the side of content creators. Regardless of who they are .
The courts tend to take a dim view of theft. Which is what this is.
The article clearly lays out that multiple requests for sound legal basis have gone unanswered . It simply doesn’t exist and Microsoft is operating on a forgiveness vs permission model.
Licensing is 100% about permissions. Clear and explicit enumeration of the permissions (or lack thereof ) for a work.
This class action lawsuit should surprise nobody. It’s a class that is sick and tired of being exploited.
Do not take my work that I contributed with explicit permissions and use it in a way I didn’t grant permission for. Full stop. It isn’t complicated.
You wouldn’t download a car and all that jazz….
https://docs.github.com/en/site-policy/github-terms/github-t...
so much animosity over rights you gave GitHub when you put your code there. "Theft" gimme a break. you license your code to GitHub so they can show it to others. This is separate to the stated license in your code. Nowhere in that terms of service document are the means that the code is shown to users specified.
Also MS themselves don't even claim that training is covered by their terms of service, they claim it is fair use.
Or you are mistaken and "the service" of github, includes all features available on the website including copilot.
Even if you're right and a court rules against them, what's to stop them changing the terms to become compliant?
Moreover terms have been largely unchanged for years AFAIK. If someone agreed to the license years ago, they can't have agreed to copilot use. Also copilot is not a service on their website, it is a separate service and they charge for it, also contradicting the terms.
What does separate service mean? What would Copilot look like if it was not separate?
Elaborate on "not a service on their website", as it is available and listed as a feature "Github Copilot" on their website.
Is the contradiction related to payment for the service, or just because you think it is separate?
Since I thought you were arguing that Githubs own terms prevent them from using the public repositories in Copilot, this is what I argued against.
If you think fair use is involved, then that's the end of the line. If MS claims fair use, then until a court says otherwise, it is. Anyone who thinks their copyright is being violated can get an injunction tomorrow.
Maybe some type of hint shown inside your code when it’s shown at github.com. There is already a text editor.
they claim the training of their model is covered by fair use, but they did not say that was the justification they were using. They don't need to claim fair use.
It's pretty clear from the Terms of Use that they can use code hosted on github.com to provide any service they like, so long as it is a GitHub service. They don't need fair use, they already have the rights to do what they are doing.
Fair use is a major part of copyright law. I do not have to ask permission to use your work.
For you to win in court you have to overcome fair use, you have to overcome innocent infringer, you have to overcome no damages.
Anyone leaving comments saying that there's an obvious way a court would rule on a copyright case involving those 3 things is wrong.
If you use my licensed work then yes, you do need to follow the terms of the license .
The issue of license / contract / copyright is messy. It doesn’t ever seem (in the USA anyway ) to be definitively answered / “solved.”
I chose AGPL v3 only on purpose.
Co pilot and users thereof (so now two levels removed ) utilizing code in whatever work are stealing my work (unless it’s AGPL v3 licensed ). The adding of intermediaries (and the most likely unknown and with no way to know) infringing is going to be very difficult to mitigate. It’s like truly unknowingly buying stolen property,
If I use it under fair use, there is nothing you can do about it.
The fourth factor of fair use is the effect on the value of your work. If I'm not affecting the works value because there is no market for it, because it has no commercial value, you are going to have a very hard time defeating this argument in court.
That changes nothing at all. A FOSS license is, in a legal sense, no different then a proprietary one. If it’s illegal it’s illegal regardless of whether it’s FOSS.
If Copilot was FOSS there'd probably be a few absolutists complaining, but they'd be mostly ignored.
Why haven't they uploaded Windows source to Copilot?
Just how much code reproduced violates copyright?
If, instead of Copilot, Bob was giving me code to copy, and it was a AGPL codebase, am I still subject to the AGPL?
The problem is that their product sometimes produces verbatim copies of licensed works, without attaching licensing information. This not only goes against the licenses under which the original authors made these works available. It can also put the product's users in danger of anything from bad publicity to a copyright lawsuit.
CoPilot is a very interesting research project. It's not yet an acceptably mature product though.
Copilot makes source code much more open, if you think about it. It implements code reuse in a different way than classes and libraries. It offers its skills equally to everyone, skills learned from everyone.
As for the cost of the API - it's expensive to run large language models, I think the price is justified. But there are free models if you like to run your own.
For a commercial version, run it on Microsoft's internal code, the code they actually own!
Yeah, the vocal few.
Do you think I give a rats ass that Copilot is duplicating my OS code?
I have to imagine most people are completely ambivalent. Of course I have no proof, I just can’t imagine anything else.
The lines probably fall somewhere along the MIT vs GPL camps.
"Ambivalent" means "of two minds," but I'm going to assume you meant that you're indifferent.
If people are/were indifferent, their licenses should reflect that. They overwhelmingly don't.
Regardless, Microsoft is legally bound to obey the licenses.
I would accept a claim of license violation if someone used copilot to autocomplete so many methods from one specific project that you have recreated that original project.
I still think it is a matter of scope. It can still be the case that a relatively small module is not cool to lift, but I think in this case we are still talking about such small subsets of functionality that it is completely divorced from the original software. Like, I could see it if a specific method were really key in some way to a unique application, a very novel solution to a difficult problem - but if that were the case, how can an AI possibly use that for a training model? In other words, the auto-suggestions of an AI are going to be common coding solutions to common coding problems that the AI has seen hundreds of thousands of times. That individual proprietary GPL, unique and novel solution is not really the stuff of an AI suggestion. In other words, the code that co-pilot is going to suggest is going to be non-unique, generic, and not really specific to the overall application at all.
> If people are/were indifferent, their licenses should reflect that. They overwhelmingly don't.
Apparently, overwhelmingly they do. At least if the licenses used are any indication.
https://github.blog/2015-03-09-open-source-license-usage-on-...
I'm almost entirely certain you're wrong about the desires bit. 99% of the developers who wrote that code won't mind.
This is completely unsubstantiated. I for one would mind microsoft profiting from the closed source code I wrote.
If I’m writing code for a query optimizer, the SQL Server solution isn’t going to magically show up.
It doesn’t mean you have the right to use any of the code it generates but Copilot itself isn’t illegal in any meaningful sense.
I don’t think it’s accidental that this product is specifically Github Copilot.
But even then I think this is legal overkill. If you use the search box on Github they will display snippets of code from public repositories without the license. Same as what Sourcegraph does same as Copilot does. Nobody here is arguing ripgrep is violating the license by displaying matches without the corresponding license.
Yes if you use a tool to violate copyright it’s copyright infringement. If you prod Midjourney into outputting near exact Starry Night that’s on you too.
So far no one has made a compelling case that Copilot itself is violating copyright.
Saves a lot of "hey how do I do this simple thing again?" memory loss issues.
The people who comment on something are disproportionately those who care a great deal.
What proportion of its capability is derived from the labor of people who don't like it? I get your point about feeling like an unwilling contributor while github/MS harvests revenue from people who like it. But there's an implication here of being in a critical majority, which I am not convinced is the case.
Reread your post. Doesn't it sound scary? You are blocked from even thinking and crafting because a specific web service is down.
Even if Google is down you can go direct to Stackoverflow and MDN, and have a choice of information sources.
Also what is "productivity" ... as in features built / month or lines of code / month?
Life's a lot easier when you can just copy whoever did the hard work without crediting/paying/etc for it.
I would hate to work at a place where advanced-but-untrustworthy autocomplete would, at all, impact the productivity of a senior engineer.
Not only does this indicate that your senior engineers' productivity is measured poorly (lines of code), but also that your senior engineers are paid to type, rather than to think.
People's expectations have already been set by this technology, and they are only going to want more. Also, AI researchers are still publishing their work out in the open for anyone to reproduce.
If there was a Copilot model out in the wild like with Stable Diffusion then this ceases to be a valid question, regardless of the model's legality. All it takes is a single leak or decision by another entity to release their own code generation model.
Incorrect suggestions all the time
I think there is a good, fundamental legal/societal question of how copyright should apply to AI output. I just don't think our existing copyright structures handle this question well.
Note there is currently a very important case before the SCOTUS that is related to this issue, [1] where the original photographer of a Prince photo is suing Andy Warhol's estate for copyright infringement. The fundamental question is whether the Warhol series of painting are "transformative" enough of the original photo. While there are always gray lines on what "transformative" means, if there is any chance that Warhol's painting are legal and not infringing, I don't see how Copilot could be in the wrong. Copilot's output, even if it contains a substantial amount of the original source, appears to me much more "transformative" than the Warhol paintings are compared to the original photo.
1. https://www.npr.org/2022/10/12/1127508725/prince-andy-warhol...
Is it really better to only draft laws that are clear without courts?
Is that provable?
Badly written law and poor legislators are a problem in any system.
Any higher court ruling might well draw some lines between different domains, but be clear that a ruling against GitHub would almost certainly be a ruling for copyright maximization and against fair use in other respects. So be careful what you wish for.
Copilot is a set of trained weight values in a matrix. There is no source code stored in that matrix. The fact that someone can prompt Copilot with specifically chosen text to generate a short sequence of code that matches a corresponding segment of code used to train the model does not mean that it is somehow "just retrieving" that snippet. It is _generating_ that code, guided by the weight matrix, via pattern-matching based on the chosen textual prompt and surrounding context.
That distinction is significant because one of the primary defenses against copyright infringement in US law is if the derived work is transformative. Copilot is a work derived in part from Github code, but it has unique capabilities far beyond returning short snippets of input code, and the work itself is clearly an extensive transformation of the input data.
This is without even considering whether concrete _outputs_ of the model that happen to match code in a repository used to train it are themselves protected via copyright or not, which is another issue entirely (and not as cut and dried as many folks on here seem to think).
This is far from the truth. The main usage of most open-source projects isn't as code, but as a product. The median user of an open-source project wants to think about the project as little as possible. They want to be as unaware as possible of the code that makes up the project. They're happy to add the project to their `requirements.txt`, add a few lines to import and use it and then never think about it again.
Also, if we agree that GitHub copilot enables you to be more productive as a developer. Can we argue that it could help open-source communities by helping them finish projects faster?
Copilot doesn't recognize what you're trying to do and then paste a code sample from a repo it has in its index. Just like DALL·E 2 doesn't produce images that say "I picked these pixels from this image and this part from this other one and these colors from this third one". When a model is trained, it's effectively a set of hundreds of millions of numbers that when combined in just the right way can produce a specific output. In my experience the vast majority of the time Copilot doesn't write code that already exists. It actually uses the variables you declared, the functions that already exist in your code base, etc.
It's not an index of best matches from GitHub for what you're trying to do.
There could be a problem for open source projects (and closed source ones as well) if Copilot could autocomplete with code from private repositories. I can't remember if it looks at them too.
It's not like they don't get it. Even Microsoft will provide the source of its products under certain circumstances for this exact reason.
It's most likely the case that in 1, 3, 5 years, Copilot won't be spitting out code blocks verbatim. It will generate rightsize code, trained on lots of publicly available code, and start reducing the surface area required to code/develop.
Stable Diffusion doesn't get in trouble right now because the artwork looks like permutations of different works; text is easy to copyright, style is more challenging, but artists are facing up against the same reality. There's no rolling this back; ML models are going to remove a ton of cruft from creative/labor based endeavors, and people are going to need to evolve to stay relevant.
They don't even trust the thing to train it on their own code, yet their boss is over here telling us they are "learning". It's a damned insult.
For example, I wouldn't care if a small YouTuber used a copyright song in the background. I would care if Disney stole a small YouTuber's original song and used it in a movie.
This is entirely consistent within my ethical framework: scale and power matters.
> Meanwhile, we open-source authors have to watch as our work is stashed in a big code library in the sky called Copilot. The user feedback & contributions we were getting? Soon, all gone
When a hacker finds a new way to get into systems or root their phone it's cool. When someone uses that technology to steal money or personal information or encrypt your files it's criminal.
Nothing new in this case.
At the very least add a comment in the generated snipped where the code originates from. That won't suffice in all cases but it's disingenuous to profit from others' work without any credit/permission.
OpenAI, Microsoft, Google, et al will STILL train such models that can do all the same things, but it will be much harder for non-industry-backed individuals to navigate the legal minefield where you must ensure you properly attribute your model outputs, only train on opt-in data, etc, etc. Surely no one really thinks that a court case against Microsoft/OpenAI (even if they lose) would stop CoPilot?
Most of these complaints seem to be extremely emotional and cherry-picked. "People's legal rights are being violated!" (you definitely don't know that, no one knows that, the article is 100% right about that), "look I prompted CoPilot for this piece of code that I already knew about and it spit it right out" (that's not how it's going to be used in practice).
It seems to me that the longer-term implications of the outcome of a lawsuit like this are far more interesting, yet almost all the comments I see are nitpicking and whining about how the world isn't the way they want it to be. I wish the conversations around generative AI could be...just better.
Earlier HN thread today on a large chunk of OS code pasted almost verbatim by the CoPilot engine into a project, but stripped of any licensing references.
Within the last few days, another thread on artists who have spent decades developing a unique and valuable style are making parallel complaints about Dall-E/SD/etc., where inputting "Xyz in the style of [Artist]" produces exactly a copy of [Artist]'s unique style, barely distinguishable from the original.
These engines are fairly literally giant collage engines, able to parse language inputs and output a collage of the input works. Maybe some are small snippets so it could be fair use, but they are also evidently capable of outputs of a far larger scope, amounting to wholesale ripoff.
Opting out or not posting on Github or whatever prevents nothing, as stuff is posted everywhere by many, and with code, it's totally legit posting a fork under OS licensing.
Is there a solution analogous to a <NoRobots> flag? How do we verify it? Will there soon be HaveIBeenUsedAsTraining adversarial systems to probe these output engines?
Not sure of the solution, but this seems to rather rapidly overstepping boundaries of creators.
I'm just not convinced by the idea that any penalty less than death isn't a penalty.
The Securities and Exchange Commission (USA) has a history of giving million-dollar fines for crimes that produced billions in profit and/or took billions away from victims. And the lack of deterrence has been reflected in the actions of the US financial industry.
https://twitter.com/docsparse/status/1581461734665367554
An english description plus three characters of a function name is enough to coax CoPilot into distributing LGPL-licensed code out of context, without a proper license. That's neither "emotional" nor "cherry-picked", it's a clear-cut license violation.
Also, this code is really just executing a mathematical operation in what I would assume is fairly standard, so it may not even fall under copyright. IANAL.
Even if it is a copyright violation, that is one out of, IDK, millions, maybe billions already of Copilot completions?
That sounds cherry-picked. Just because some high profile, highly-retweeted person says something doesn't make it not cherry picked.
So let's say I obtain an illegal copy of microsoft windows' source code. Under this precedent, what stops me from just (overfitting) training a neural network to produce the source code verbatim, sans any license notice?
But it doesn't end there. What stops me from making a neural network that exactly reproduces the bytes of Illegally_Ripped_Disney_Movie.mp4 that I obtain from the pirate bay? Copyright need not apply.
At what point can the neural network I've described (which is intentionally designed to violate copyright) distinguishable from a neural network like Copilot and others which violate copyright extrinsically?
I don't know much about AI/SL but I think the logical outcome is that using a transformer to generate source code is going to result in an over-trained system that produces sections of code verbatim because of the relative size of the space of all valid programs vs the space of all possible programs. It's not like art where you get soft failure if a single pixel or word is wrong: the inclusion, exclusion, or replacement of a single instruction or symbol is enough to introduce fatal bugs into a computer program. If the system doesn't have the capacity to understand programs generally (which it probably won't if it's just a transformer), then you're going to end up with a system that spits out samples from training data that do work.
In 999 out of a 1000 cases it’s just spitting out boilerplate though.
If someone is motivated to search through our entire (proprietary, private) codebase. They match it with repositories that are freely available. They’re properly motivated to make a problem out of it (some twitter randos?), and most importantly they gain some benefit out of spending hundreds of thousands of dollars engaging with our legal team.
By the time you satisfy all the conditions required for it to be an issue you are talking nation-state actors.
ahem napster, pirate bay ...
I'm no copyright law expert and I'm certainly not a lawyer, but it seems to me that in your examples you're setting out to violate copyright as a goal, which seems like it would be a factor in a court case.
To answer your last question, your examples are pretty clearly distinguishable from Copilot in their final states that you describe. IDK exactly *when* during overfitting that line is crossed, maybe it's crossed the moment you personally decide to knowingly publish copyrighted content and has nothing to do with the neural network itself?
It doesn't void copyright.
Anyone that uses code that Copilot spits out which infringes on someone else's copyright is still liable. There's no requirement for intent. That may be a mitigating factor in terms of remediation, but it cannot void the copyright itself.
It can, however, produce a plague of completely ignorant copyright infringement, and since the user of Copilot has no idea where the code is coming from, there's no way to check if it was trained on infringing code.
If I used Copilot I would be real worried about the legal implications for me and that I could easily be accused of copyright infringement or plagiarism[*]. Of course people seem to not really care these days if they can cheat to get ahead so this is probably a feature, and 99.9% of user won't have their reputations ruined by using it.
[*] Although I'm personally more worried about the fact that most of the code will be wikipedia/blogs-quality and filled with bugs, edge cases and performance issues.
Between that and the typical quality of Copilot snippets, it's pretty obvious who the real beneficiary of this technology is: large sweatshops like Infosys.
But the only thing that does is make step 3 of fair use harder to clear. Not impossible.
There are fair uses of copyright that use the entire identical work as is.
If you, only once, steal lines of code that you don't have license to do so and use them to make money, that's the same exact thing. "Trusting the algo" and saying "whoops I'm sorry" doesn't make a strong legal defense.
In a company of 1000 programmers, what are the odds that copilot increases the risk of using improperly licensed code because "well because microsoft gave it to us it has to be legit!"
And sure, stackoverflow copying is a thing, but they clearly tell you the license by which you can use said code: https://creativecommons.org/licenses/by-sa/4.0/
If copilot gives you CC-by-sa code, will it tell you so you can properly credit?
So, no. You can in fact "steal" several lines of code and use them to make money and be legally clean as a whistle. It isn't that clear cut.
"Google’s petition for certiorari poses two questions. The first asks whether Java’s API is copyrightable. It asks us to examine two of the statutory provisions just mentioned, one that permits copyrighting computer programs and the other that forbids copyrighting, e.g., “process[es],” “system[s],” and “method[s] of operation.” Google believes that the API’s declaring code and organization fall into these latter categories and are expressly excluded from copyright protection. The second question asks us to determine whether Google’s use of the API was a “fair use.” Google believes that it was.
A holding for Google on either question presented would dispense with Oracle’s copyright claims. Given the rapidly changing technological, economic, and business-related circumstances, we believe we should not answer more than is necessary to resolve the parties’ dispute. We shall assume, but purely for argument’s sake, that the entire Sun Java API falls within the definition of that which can be copyrighted. We shall ask instead whether Google’s use of part of that API was a “fair use.” Unlike the Federal Circuit, we conclude that it was."
In the US, "[T]he fair use of a copyrighted work ... is not an infringement of copyright." (17 USC 107, Oracle at 14). Whether there is a difference between a finding of no infringement or infringement with no liability is academic.
Practically, the copyright on the Java API is commercially worthless because after Oracle anyone may freely copy any and all of it and use it to compete with its creator.
Any other software vendor who thinks it has platform "lock in" because its customers built to their API should take notice. (e.g., Amazon AWS).
In a company of a thousand programmers there are much easier ways to find improperly licensed code.
There are posts under an earlier license which was CC BY-SA 3.0.
There are people who don't have accounts anymore or haven't logged in to accept an updated license.
Only the changes to the post after the 3.0 to 4.0 in the above case are technically licensed under 4.0 (the original post is still under 3.0).
Furthermore, Stack Overflow didn't follow the proper process for updating the license.
https://meta.stackexchange.com/questions/333089/stack-exchan...
For example - https://stackoverflow.com/posts/11574647/timeline
Look at the license and the Aug 22 change and consider if that removing "Hope that helps" was a sufficient change to relicense it.
You want to use my code, without ever knowing I wrote it? You want to use my hard work, regurgitated anonymously, stripped of all credit, stripped of all attribution, stripped of all identity and ancestry and citation? FUCK YOU
There's no need to defend something so obviously harmful, so why do you do it?
The law should be amended to make this kind of theft illegal.
It's not ambiguous.
You can have a totally legitimate business making hacksaws and bolt cutters.
Now if your customers use these tools to break into homes and steal things, then yes, that’s illegal.
But making the hacksaws and bolt cutters is not.
This has the potential to severely damage open source - I would not host my open source project on GitHub, especially if it was copyleft. I'm sure many others wouldn't either. Some of these authors make amazing software that we might not see because of this.
Every artist, every creative individual, must EXPLICITLY OPT IN to having their hard work regurgitated anonymously by Copilot or Dall-E or whatever.
If you want to donate your code or your painting or your music, so it can be used ("written", "painted"), in whole or in part, by everyone else, without attribution, then go ahead and opt in.
Otherwise, you can't use the artist's or author's creative work for training.
All these code/art washing systems, that absorb and mix and regurgitate the hard work of creative people must be strictly opt in.
While I understand the rub about licenses, the fact is the vast majority of code is not all that original or unique. Some fringe amount is, and those edge cases are worth discussing.
But the rest? Likely not all in all all that special. Yes, we get paid good money to do it. But is that a function of what it takes to do the work, or the demand for the skill (relative to supply of that skill)?
Frankly, I think some ppl just plain ol' fear Copilot. And either don't want ro admit it, or they have buried that fear. I'm not advocating ignoring the law / licenses. But putting a licence and lipstick on what is an everyday pig doesn't make that pig a unicorn. Does it?
But that's not the type to fear Copilot. Yes, they might object to the license violation. I get that. I acknowledge that. But when you're that intelligent and that creative you don't fear being replaced - displaced? - by something like Copilot. Nah. That's a fear for the mundane and the common. That's a fear for the rest of us.
It can be both things. (I'm not endorsing either view, just trying to clarify.)
It's his code. This "high profile, highly-retweeted" crap is an appeal to emotion. He has a specific and legitimate interest in protecting his own intellectual property. It's not "cherry-picking" to report a crime being committed on your front lawn.
And as soon as he sees someone actually take the chair from his front lawn he can report it as a crime. He cannot report the people walking past because they could potentially steal his chair.
One could even argue that if he didn’t want his chair taken, maybe he should have locked it in his shed.
Of course, these chairs duplicate, so it’s not as if he loses his own chair.
No, it is not.
For it to be a violation you have to lose a court case. To lose a court case a court has to find against your fair-use defense.
A fair use defense is fact specific to the parties involved. What's fair for you might not be fair for me.
Only a court can determine fair use.
Are you saying that because there are millions of copyright violations, Copilot is too big to fail? Or are you’d saying that Copilot is too big to be held accountable for flagrant violations of the law?
I guarantee Copilot’s developers knew it was spitting out verbatim code. It’s too obvious, and probably would result in a perfect rating for the prompt.
Stealing a car is a large degree of crime for an individual. If they had stolen a penny, it would be a small degree of crime and we'd be more willing to let it slide.
https://en.wikipedia.org/wiki/Structure,_sequence_and_organi...
https://en.wikipedia.org/wiki/Abstraction-Filtration-Compari...
Suppose you want to express what some other writing does, and then you see that writing? What can you do?
(You write your ideas in your own words, and quote and cite your sources)
If I want to describe life with a nature metaphor, and then see you do it with a waterfall, I can probably get away with using a waterfall to the same metaphorical effect in my story.
Can I do that in code?
https://stackoverflow.com/questions/17913191/using-typedef-i...
Found it after 5 mins and a couple tweaks to the search terms.
Another:
https://vdoc.pub/documents/direct-methods-for-sparse-linear-...
Someone copied this guy's book and put it on scribd: https://www.scribd.com/document/514019650/Direct-Methods-for...
Someone put it on a "personal" edu page:
https://people.sc.fsu.edu/~jburkardt/c_src/csparse/csparse.c
A modified version of it here marked as open-source:
https://github.com/rwl/CSparse.py/blob/master/csparse.py
More:
https://tonus.pages.math.unistra.fr/schnaps/schnaps/csparse_...
Google search used to find them:
https://www.google.com/search?q=Sparse+matrix+addition+%22ch...
Could probably find more if I looked harder.
Side note, looks like in a lot of places people do the ""proper""-ish thing and leave this guy's name on the code.
For example, I'm pretty sure this is why some models turn up a near exact version of The Girl With The Pearl Earring.
That said... is that any different from someone copying and pasting into their code vs copilot doing it?
If someone randomly pastes code that has a copyright, and people use it, how are they supposed to know they shouldn't be using it?
I imagine we're talking functions here though. Not sure if Copilot would reproduce entire libraries unprompted if they're not open-source, anyone have an answer?
Artificially starving copilot of context and then showing that it recites parts of its training dataset is mundane.
A better future to me. I don't want pictures of my face training ML models, nor do I want my art, or my code. I don't want my face to be more recognizable to AI, and I don't want my work to contribute to the consolidation of power to a few big firms. And for what, what can ML models even bring me besides surveillance? Cool art? Text-to-speech?
Certainly facial recognition models etc, can be, but those would seem to be appropriately covered by the Google Books ruling dealing with discriminitive models: https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,.... They've also been a thing for quite a while, so I think that cat escaped the bag a long time ago.
Remember clip art in Microsoft Word back in the day? Now there is an infinite supply of that. Stock images? Infinite supply. Solo filmmakers are going to have a much easier time creating their own films that can rival the best movie studios in the world. Any text anywhere will be read to you in any voice or voices you like, with tone and setting appropriate sound-effects. Smaller things will just be better too, noise cancellation on microphones? Easy and free. Image editing? Trivial to remove, relight, reposition, etc, etc, etc.
So many other things too. It's going to be magnificent. If you're not into then I guess to each their own, but I do think we are looking at something that can be a net good for everyone in the world, so long as it's available and cheap for everyone in the world.
Who here is doing a startup to secure licensing rights to every companies surveillance camera videos to make the AI/Surveillance version of Equifax Worknumber? Maybe you offer to give them the surveillance system for free in return for the rights?
Because of this, if the law looks at GitHub Copilot, I would expect that they would find Copilot to be A-OK despite the occasional regurgitation of copyrighted material that isn't fair-use, as long as it is removed upon request.
You argue that desiring ownership of works that you created is “extremely emotional and cherry-picked”, but you do not provide a compelling argument why artists, photographers, and indeed programmers should be excluded from the conversation when it is their art, photography and code that is being appropriated in the first place.
Ultimately my argument is that these tools will allow human beings to accomplish more things with less and that these tools should be distributed to as many people as possible for as little cost as possible. Part of that belief comes from the fact that I think these tools are coming no matter what and I'm slightly concerned about the potential (although unlikely-looking) future where a small number of large corporations are the only ones controlling these tools and they just rent-seek on them.
For programmers, the job market for loud-mouthed posers and plagiarizers will grow and quality will suffer. But those programmers will be fluent in marketing speak.
There will definitely be more cheap rehashed garbage online and we will be forced to invent tools to wade through it. I actually look at that as a bright side because there's already a lot of cheap rehashed garbage, we just don't have good tools for wading through it yet because it hasn't become completely intolerable yet.
So what? I’m not being snarky: does that actually make any difference, legally?
Github/OpenAI should have to pay a licensing fee to use GPL and similarly-licensed code in their closed-source derivative IP (CoPilot).
A future where technology is developed in accordance with longstanding law? Also, maybe a future where my copyrighted works are compensated for when they're being used to automate my job away? If the music industry can deal with royalties, maybe software can, too?
> OpenAI, Microsoft, Google, et al will STILL train such models that can do all the same things, but it will be much harder for non-industry-backed individuals to navigate the legal minefield where you must ensure you properly attribute your model outputs, only train on opt-in data, etc, etc. Surely no one really thinks that a court case against Microsoft/OpenAI (even if they lose) would stop CoPilot?
I'd expect injunctions against Microsoft/OpenAI from further training CoPilot with inappropriately-licensed code. I'd expect damages for all of the instances of copyrighted material that CoPilot regurgitates.
> Most of these complaints seem to be extremely emotional and cherry-picked. "People's legal rights are being violated!" (you definitely don't know that, no one knows that, the article is 100% right about that), "look I prompted CoPilot for this piece of code that I already knew about and it spit it right out" (that's not how it's going to be used in practice).
How are these emotional? They're opinions, like all legal claims, supported by facts. It doesn't really matter if they're "cherry-picked" or not, only whether CoPilot actually violated copyright. If it gives seemingly novel results 999 times out of 1000, but in the other case verbatim generates copyrighted material without proper permission, then that is a copyright violation. At scale, that's a copyright violation with some modest damages even.
Maybe this is not the industry to emulate: https://www.theatlantic.com/business/archive/2011/11/how-mus...
The WAY it was done with copilot is the problem: no attribution, just shoving all legal liability off on the end “programmer” without providing the attribution required TO COMPLY WITH LICENSES as the diligent programmer tries to clear all the code copilot handed it without meta data.
Go read the article before arguing further, please. Otherwise you are wasting all of our time.
I hope that when a case on generative models hits the courts that it's found that training on data from the Internet counts as fair use. I hope that for the reasons I laid out in my comment, because I think that if it isn't then we are all in trouble since these tools will STILL EXIST, but they will be in the hands of the few instead of the many. My main reaction is to how short-sighted it seems like the authors and many others are being about this technology in general. They seem to think they can just wish it away.
I also think that training on data from the Internet is fair-use, but I'm not a lawyer and I haven't studied the law extensively, so who cares what I think about that.
Assuming they will even get off the ground with their copilot system not suggesting vulnerabilities and license traps already.
Though that doesn't justify such a dismissive attitude towards ordinary HN commenters. As the way it's written implies that most are too stupid and overly emotional, which is more likely to fuel complaints instead of dousing them.
That's just not true. If the "industry standard way to do things" is to violate other peoples' copyright, then everyone doing that absolutely can be sued.
And while it's not clear if using these AI tools constitutes copyright infringement, it looks to me like there's at least a very strong case that could be made.
And at up to $10,000 per copy (register your code with the copyright office if you care about this issue!), that starts to add up very quickly. Even for a company like Microsoft.
But in terms of user behavior, it's rather the same. I used to make more stuff publicly available online than I do now, and the mass-scale surveillance and data modeling that big companies do off of publicly available stuff is a big part of that.
Generally that's how you get walled gardens - by abusing the commons - but here you'd need not just a walled garden but a TINY TINY invitation only one if you don't want people doing mass surveillance and data modeling (CoPilot is really more of the latter than the former, but any of this "scrape the whole internet" stuff is just a tiny little sidestep away from being used for more blatantly evil surveillance purposes - here we're training a generative model, they're we're de-annonymizing everything you've written anywhere...).
Is there a good solution to "BigCos are gonna do whatever they want with the shit you make" other than invite-only, paid-content type models?
I'm not sure what the argument being made here is. If you make something opaque enough that no one can tell if it's violating legal rights, no one is allowed to say anything about it? This seems uncomfortably close to "it's only a crime if you get caught"
1. I think if you don't want your code re-used in CoPilot you should have that right
2. I think if CoPilot gets smart enough that it can read your open source code and then reproduce the algorithms without copying your code that should be fair use. It's the same thing a human would do. AFAIK CoPilot can not do that but I can certainly imagine it's not too many years away from that.
3. I think I would opt into sharing all my open source code mostly unrestricted with services like CoPilot. I think the group of people that choose to share their code with AI will do better over all than those that lock their code behind licenses.
Note that I'm referring to snippets of code. I don't know what a good definition of snippet is. In other words, if AI helps me write chunks of 10-100 lines at a time I don't see a problem. Effectively, S.O. answer level of snippets. Whereas, if I tell AI "create LibreOffice" and it clones the millions of lines of code, I think that is a problem. I don't know where the cut off is.
The examples where people can show it reproducing snippits of code are more the exception than the rule. And they are usually done by people who are trying to manuliption into proving that it can reproduce copyrighted code.
Some people tend to think of it as a search engine. Looking though it's database for relevant snippets for the current situation and regurgitating them unmodified.
But that's really not what it's doing. It's more like the AI autocomplete on your phone, but for code.
It might not be able to understand the algorithms. But it seems to be able to adapt simple algorithms that it's seen multiple times in it's training data to match the surrounding code (in style, naming conventions, and actually using the variable/functions you already have).
I don't have number, but from my experience, I would say it generates uniqu(ish) non-copyrighted code at least 95% of the time.
The only question is what to do about the other times when it does occasionally output potentially copyright infringing code, either by accident, or when it's forced.
People already have that right - all you have to do is not host your code on GitHub.
I don't fucking care, I'm not in the business of competing with OpenAI or whatever. If you want to launch and AI startup but you can't that's your fucking problem, not mine. I just don't want them violating the licenses of the open-source programs I have created.
> "look I prompted CoPilot for this piece of code that I already knew about and it spit it right out" (that's not how it's going to be used in practice).
It proves that copilot has the capacity to copy existing code without fulfilling the requirements of the license. I don't care if it's "cherrypicked", this shouldn't happen under any circumstances.
> I wish the conversations around generative AI could be...just better.
I wish that these people making all these complicated language-comprehension machine-learning systems could read the fucking license statement at the top of the file and copy that license statement along with the code. this ought to be a solvable problem. I'm pretty sure i could write a bash script that does it if M$ is looking to hire.
The product would be useless if it prompted you with license approvals. They didn't care and removed them. They consciously decided to prioritize their paid-for product over the rights of their users. I'm amazed that MS's lawyers allowed it out the door. That's the even scarier part.
If I'm writing some code and want the suggestion to be good then why wouldn't I use the name of a top programmer as a prompt?
Imagine if every new piece of software your wrote had to be tested for legality because you don't know that it's explicitly legal. Oh there aren't laws for this new thing, so I guess you should challenge yourself all the way to the supreme court?
I get the author not liking Copilot, but I don't see that GitHub/Microsoft have any kind of obligation to figure this out just because they're GitHub/Microsoft.
If I as an individual had this obligation placed upon me I'd just never write any more code.
Ultimately I think, like open source, Copilot and the tools that will follow advance human progress in novel ways. Software getting easier to make is a good thing. If you don't like this particular implementation of something helpful, feel free to start an open source alternative without challenging yourself in the supreme court.
What's explicitly illegal about this?
The fact that the work is copyrighted, rights withheld in the absence of a license and limited with one. Also, you should know that willful infringement can carry 5x the statutory damages compared to accidental, and spurious claims of fair use would be distinctly unhelpful to your case. Just by posting that comment, you have probably compromised your position in any future copyright case you might be involved in, or you might even have invited one. I really recommend being more careful when anything legal is involved.
That's not how I interpret what's happening.
People who produce things have rights over their products. Be it artists, craftsmen, inventors, entrepreneurs or coders. There is a legitimate question here as to whether CoPilot has infringed upon those rights. I don't see it being about "making something illegal." I see it about answering a valid question as to whether CoPilot is liable for measurable damages caused to creators under existing laws.
But at that point, it would be just like someone cloning the Github code without following the license and in that case, it should become obvious that there is a clear violation harming the creators. But in most use cases of CoPilot, where-in people are just using it to build their own product, I doubt there is a cause for damages.
The question is whether the courts will find damages. Everything else has nothing to do with my comment. You might have your own ideas and opinions, which is fine. So do I. Both are irrelevant. The point is that there is a legal question here that the courts alone are equipped to answer.
For open-source projects, CoPilot is in the realm of fair-use for snippets but it can be mis-used just like Github can be mis-used if someone blatantly copies a repository.
But if either of those are not true, then the GENERAL point is true and your point is not.
---
Gais, gais, I downloaded the code using an automation tool called a browser, so it's fair use and not infringing!
yay for technicalities!
Plenty of people over the years have attempted to skirt the law by putting a proxy in the middle and as such, plenty of clarifications have happened that it doesn't matter.
MS's stance here is specifically that it's the responsibility of the person/company using copilot to ensure the code isn't infringing, it is NOT their stance that the code itself is not infringing.
Their stance is that using it as TRAINING DATA is fair use, so they themselves hold no liability, only their users.
---
It's similar to chicken factories claiming they hold no liability for employing illegal immigrants because those illegal immigrants are the ones who chose to work there. And that they also hold no liability if they go through a 3rd party that exclusively hires illegal immigrants. The law very clearly refutes both stances.
I'd argue that people who open-source code expect other people to learn from it in a small way of snippets and that constitutes fair-use.
https://stackoverflow.com/help/licensing
Copying code samples from a copyrighted work for use in a commercial product is not fair use.
Important difference is you don't know where your Copilot snippets come from.
You cannot sample music without permission no matter how short the sample may be.
Similarly you cannot steal a snippet of someone else's code without permission or the correct licensing.
I can’t use a snippet from a recording no matter how short but I can use a tiny snippet of a composition. You can’t copyright a single note.
Which is a blatantly mistaken court ruling and one which I will not enforce if I am on a jury in such a trial.
> I get the author not liking Copilot, but I don't see that GitHub/Microsoft have any kind of obligation to figure this out just because they're GitHub/Microsoft.
Because its a trillion dollar company with an infinite amount of lawyers and legal resources?
Authors have explicitly and deliberately made it illegal for a person (or corporation) to do what Copilot is doing. Doing it through the legal non-entity of an AI changes absolutely nothing; it's still illegal. To say otherwise is to say that "AI-washing" can be used to nullify any law, which is of course totally absurd. The assumption you lead with is not what anyone is actually trying to argue.
That's what "all rights reserved" means.
So if you trained yourself to only regurgitate github code with wanton abandon and careless disregard for licensing, then yeah, you're liable to violate that default copyright, and certainly going to be violating license rules if you're regurgitating large blocks from memory but not their accompanying licenses.
This is the system that github and Microsoft participate in and willingly and purposely perpetuate. They benefit immensely from copyright law and protection of their code. You can get that they will damned-well avoid letting copilot anywhere near Windows' source, and they would very much enforce their copyright if copilot was spitting that code out for the masses to use.
While there is no US case law that explicitly says "training AI is fair use", the Second Circuit says that scanning books to make a search engine for them is. And the absolute worst interpretation of AI is that it's just a very well-compressed search engine index for its training set data[1]. So I'm not entirely sure if we can even thread the needle to only ban Copilot or AI training as a whole without also creating harmful precedent for search engines. Actual judges may try, I'm not sure if they'll succeed.
Internationally, the EU already legalized training AI on copyrighted works[2]. So if we do win against Copilot in court, all we've really done is shift AI research over to the EU where laws are already more favorable.
I fully agree that Microsoft is shoving too much liability onto their users, though. And this, again, also applies to all generative AI. My personal opinion with generative AI is that it's a nice curio, but not anywhere close to "production-ready", and Microsoft and OpenAI are trying to sell us on a lie that it's better than it really is.
[0] This also implies that all y'all playing around with image generators are just as much of a freeloader as Microsoft is.
[1] This viewpoint is also called "compressionism".
[2] This was part of the most recent EU Copyright Directive update - the one that added a de facto upload filtering requirement. It also added a copyright exception for museums and historical preservation.
I don't think this parallel makes sense because a search engine links to copyrighted works, each of which is still governed by its original copyright, while these AI create derivative works or reproduce the original works without even attribution.
Indeed, if an AI was just and index for the training set there would be less of a problem because the origin of a work could be found and its license honored.
Not necessarily. Copilot is a special case because it is using licensed code and the model is a derived function. There is an interpretation where it needs to be open sourced.
I genuinely think that all of them are. But that's not why I'm against them.
We've seen the effects that text and image generators have on the bottom segment of content generation (SEO pages). As the technology matures, it'll displace more and more, in both arts and engineering.
The copilot service backed by an army of actual humans wouldn’t be a story at all. Nor would anyone be angry, if an individual offered coding skills as a service, and had gone through the exercise of learning great amount to open source software to do so.
No open source license was written with this in mind. Because previously learning was something only humans could do and no one had issue with sharing that knowledge. Until licenses take machine learning use into account I see no problems with Copilot.
Source cannot be open if you restrict any viewing of it.
No, people's disgust is with Microsoft violating their legal privileges.
> The copilot service backed by an army of actual humans wouldn’t be a story at all.
Correct, it would be an open-and-shut lawsuit.
Surely I don’t need to recite the last 50 years of tech legal precedent and case history for you to see that such a blanket generalization cannot be left unaddressed.
Litigants litigate
Do an internet search for “copyright utilitarian” and read up on it if you don’t believe me!
Copyright is about protecting artistic expression which is held in contrast to the useful nature of a work.
"In no case does copyright protection for an original work of authorship extend to any idea, procedure, process, system, method of operation, concept, principle, or discovery, regardless of the form in which it is described, explained, illustrated, or embodied in such work." (17 USC 102(b) [0]).
See also the "Useful Articles" doctrine. [1]
[0] https://www.law.cornell.edu/uscode/text/17/102
[1] https://en.wikipedia.org/wiki/Copyright_law_of_the_United_St...
I predict your attempt at tactically “managing” this copilot scandal will not play well on HN to experienced coders, your Microsoft colleagues chiming in next claiming it boosts their productivity notwithstanding.
Yes, I do indeed suggest astroturfing afoot.
I just think it’s hilarious how fair use is so widely supported on HN when it comes to music, or videos, or interface names, but all of a sudden is a moral crisis when it appears to threaten the value of HN member’s labor.
Definitely an issue...but not as simple as copy and paste.
I can't use co-pilot because if I am stealing someone else's copyrighted code I'm in trouble from a legal standpoint.
I don't see how that invalidates the copyright/license argument. So, instead of just a straight up license violation it's a license violation via plagiarism.
That argument wouldn't hold up even if it was a human that caused the violation. You can't just paraphrase someones licensed work and then lie about looking at and pretend you made it yourself, which is basically what seems to happen with co-pilot, as it doesn't also automatically reproduce the license of the code it reproduces.
The arguments against my point always assume perfect memory of everything this model is consumed. This is the plagiarism position. In reality, some patterns are more common than others and generate a code that looks nearly identical. I can’t speak for the reasons for this, as I’m not familiar with all of the methods. However, I don’t assume that is the current working state or intent of Codex.
It remains to be seen whether ML is true "learning" in the sense of developing a skill the way a human does over time.
It is however irrelevant to the manner in which this model operates today.
Yes you can. That's exactly why you paraphrased it instead of copying verbatim.
At the fringes, your transformation may not be enough to overcome the requirements, but that's an exception. Nearly all paraphrasing is legal by default.
You explicitly agree to this when you upload code to GitHub.
FOSS folks shouldn’t have sold their soul to the proprietary devil but they did and now they have to deal with it.
Also, the section breaks and headers and boxes lack obvious rhyme or reason. It scans a tiny bit like a classy version of Time Cube. You keep getting hit with different font sizes and font styles and lines and ribbons and colors and you're not quite sure why.
This is on a powerful PC with a state-of-the-art graphics card.
After that, what are you left with? A small enough proportion of developers, and Microsoft evidently thinks so, who don't know, and/or don't care, and/or don't have the time to fight their Extend-Embrace phase of take-over of Github.
One could argue that the purchase of Github was Extend, and their involvement with OpenAI, the Codex, and the potentially illegal use of OSS (subject to the legal investigations) is Embrace.
It's my own personal view that Microsoft held-back the progress of software development by probably a decade or so with their shady commingling with academia, blatant crippling of C# .NET to sell Visual Studio, and endlessly so forth. So I am, along with many, upset to see a business like this EEE their way into OSS, something which is dear and special to so many.
In the end, and I must state in my own opinion (since there is an element of speculation here), I am just pleased that there are still people out there who are not letting Microsoft continue their old ways.
Either that, or you wake up one day to see that Microsoft have stole your open source software and Microsoft says "but muh AI".
>> "I'm not too sure yet, but I wouldn't be surprised if we wake up one day, and just like how it went with Facebook buying oculus, we will all of a sudden require some "microsoft account" to log into Github."
See: Minecraft. It was already a goldmine when they bought it, but they built it into an even bigger one before forcing millions to have a foot into their ecosystem. Copilot might be their way of making everyone dependent on GitHub before "moving on" from git and offering a Community Edition of their own source control system.
It'll be easy. A lot of people hate Git.
Extend: Make people dependent on GitHub Copilot. Require a Microsoft account (soon).
Extinguish: Sunset git and transition to Microsoft's own source control system.
- MS absolutely has the authority to copy, use, and even train their models on your GPL-license code, because you agreed to let them do that when you signed their EULA when you decided to host your code on GitHub.
- This authority does not extend to CoPilot users, who cannot republish your GPL-licensed code without respecting the license. But remember that people have always had the ability (not authority) to copy and use open source code in violation of the license. This simply makes it embarrassingly easy for a person to do so unknowingly (although, legally, this would probably be considered negligence, not ignorance).
IANAL but I wonder if the extreme facilitation of copyright infringement here could be considered gross negligence on the part of MS, as they're almost entrapping their own customers in a minefield of copyright concerns. Can't wait to find out.
The logical next step in this arms race is for the GPL camp to build tools to automatically search for copyright infringement in large codebases. Copyright holders could set up hotlines for insiders to blow the whistle on infringement in exchange for compensation, since AFAICT all litigation precedent in the US has so far resulted in settlement.
What about GPL code which you don't own, but post to Github, Like the gcc mirror repo?
you must have the right to publish the code you put on github.com, and by publishing to github.com, you assert that you have the rights to do so. you also grant GitHub the right to show that code to others, no matter what license your code is licensed under.
why does no one read the terms of service or license agreements? these questions are answered there and this "copilot is stealing" stuff won't even make it to court.
>this "copilot is stealing" stuff won't even make it to court
IANAL but I am heavily skeptical of your confidence here.
it doesn't allow GitHub to violate any license; it explicitly grants GitHub an additional license on top of the license you choose for your code.
the only way to revoke this license grant to GitHub is to remove your code from github.com.
> IANAL but I am heavily skeptical of your confidence here.
ok. it's all spelled out in the terms of use. I'll ask you this, though; who do you think Terms of Usage/Service documents are intended to protect?
Maybe GitHub is saying it's not their fault and that they were misinformed.
Those types of arguments usually don't get all that far with copyright violations.
either way, github is absolved. if a DMCA claim is filed on the movie, then that gets quarantined and removed from the list of things that they can show their users.
the user that uploads stuff to github.com attests that they have the right to upload it. by being uploaded, github can assume that it has the rights it asks of users unless and until they are told that they do not have those rights. so, github are covered until they are formally told that they are not, at which point they must restrict access to that data to themselves and others, which they regularly do.
these types of arguments do indeed work very well if github reacts promptly when they are told that they are using rights they do not have. github is not primarily used for piracy, like thepiratebay, and thus has a valid claim that they were lied to by a user. this is when the user who uploaded gets involved legally if the true copyright holder chooses to involve them.
you being mad at github, microsoft, or me doesn’t make any of us wrong.
> Even if copyright holders are constantly playing the game of reporting public repos to GitHub to remove it's not going to be enough
There is no copyright police outside of criminal infringement.
The DMCA gives you the tools to protect your copyright, if 1000 people infringe on your copyright you need to be ready to sue 1000 people OR attack the platform under safe harbor.
If someone uploads your code to github, you need to find it and ask them to remove it.
If someone uses your code via copilot, infringingly, you need to find it and ask them to remove it.
This is the law as it currently stands.
There seems to be -- on the whole -- little respect for the spirit of the GPL and LGPL and it really is quite a change from, say, 20 years ago, when the 'free software' movement was I think more ascendant.
I think we have a generation of software developers who have only known a world where copious quantities of high quality source code has been made available to them under very liberal licenses -- which they in turn make careers and companies out of using / exploiting.
I, too, do this, and I generally open my modest projects under Apache or MIT or Mozilla style licenses. I do this because I want people to use my things, or to be able to use them as resume / portfolio material. Or because my employer at the time has helped fund construction of them.
But I also occasionally use the GPL/LGPL/AGPL, when I want to explicitly avoid corporate entities from exploiting said material without either consulting with me or in turn making their efforts free.
And in turn, I respect the value and power of the GPL for that purpose.
So many of the comments here are trivializing the value of free software and the licenses which make it possible, and acting like there's just this... natural right... to go out there and build on other people's work without recognition / compensation / contribution.
There are too many examples of CoPilot violating the spirit -- if not the actual legal letter -- of the GPL. This is unacceptable. I'm glad that someone is attempting a legal test.
Free software is not your data to mine. It is the blood sweat and tears of thousands of developers who do their work in community spirit, but under explicitly free software principles.
Putting something out under a free software copyleft-style license is not the same as saying "You can do with this what you want." It's "I made this, you can build on it, but what you made also has to be free. Or you negotiate with me."
And what I'm getting from the whole CoPilot fiasco is: GPL / free software does not belong on GitHub. And it might end up having to be put, generally, behind barriers that explicitly (technically and legally) prevent CoPilot & similar systems from getting access to it.
EDIT: I also fully expect a new version of the GPL to be published that includes clauses against this kind of datamining.
0: https://docs.github.com/en/site-policy/github-terms/github-t...
But the show-stopping problem is that copilot is sometimes producing code that is more than fair use of other code that, and is unable to attribute the code or identify how that code is licensed. It is copilot's (Microsoft's) fault that it auto-generates legal minefields, not the person who made an informed decision about licensing their own code.
In spite of the likely downvoting, I'll say that people should be grateful for reciprocal licenses not just because they were and are the foundation of free software (as you point out), but because they shine a light on what it means to license code, and how we are forced to revisit the difference between copyright and licensing when a reciprocal license is violated.
It's easy to forget that protection of creative works is only a means to end, not the goal or ideal state.
I believe AI systems will be able to help us build a new system that can track attribution of ideas and identify predecessor works from derivative products. This attribution could then form the foundation of a reward system. This is just one possible future.
Jesus Christ, dramatic much? Are people that stumble upon a piece of code while googling how to do something, and end up copying and pasting the code from the repo, really building the open source community? Because that's essentially what it is. Whether I use copilot to generate a tedious function, or I copy it from your open source repo I'm on the same level of being a member of your open source community.
This whole thing feels like artists screaming how AI generated art is horrible, trying to figure out how to sabotage it, or how to start lawsuits - just because their value went down just a bit. Same thing with developers.
The entire fucking concept of intellectual property and copyright is flawed from the get go. The issue people are wrestling with beneath the surface is not copyright but the monetary system itself which incentivizes this "chisel off one another" behavior and "MINE!" behavior because otherwise how will you survive if you can't monetize your actions?, but intelligent socioeconomic alternatives exist: https://www.youtube.com/watch?v=lBIdk-fgCeQ
People are trying to solve this problem in an ass-backwards way. Either move to universal basic income or a resource-based economy and make all ideas 'free', 'copyable' and 'remixable' since it doesn't matter either way you have access to some resources (in UBI) or all resources for free (in resource based economy) and don't need to monetize anything since you have access to everything...instead people are content with making life shittier.
"We stand on the shoulders of giants" said Newton, but oh no.. this piece of paper called 'the law' knows better!
Many people are upset because Microsoft is hiding behind copyright and lawyers to enforce it, while at the same time ignoring the concept of intellectual property when it comes to smaller players. I'd imagine that if Microsoft removed copyright on all their code and released it and Copilot as open source, there would be much less outrage.
The issue here, for me at least, isn't centered around copyright as a concept; it's about the asymmetry of the situation. Microsoft is exploiting those without any recourse in order to sell a product.
Your ideas about a new economy and no copyright are interesting, but they will not happen any time soon. In the meantime, in reality, Microsoft is making millions based on an enormous pile of community code while not offering their code back to that community.
In that case, something like the capped profit model OpenAI has (but with less profits) could work. They decide "Okay after we've reached this amount of money for Copilot, we'll both profit and have enough money to sustain it as a service until the next technological breakthrough makes this obsolete"
Then just make it free for everyone forever.
The author seems to be implying that since Copilot can reproduce the code of open source repository X in certain scenarios there'd be no reason for programmers to learn/use/engage with repository X. But this is silly. Maybe some open source repositories could be tab completed with a little prompting but people will presumably choose to add a dependency instead of tab completing the code of express or something.
(Or I guess more technically, the original authors copyright still applies, and the rights granted to use the work under the license as an exception to the strict limitation under copyright - do not apply...)
Perhaps the author means that there’s a possibility that the programmer to whom the code was suggested won’t necessarily know its provenance and how to engage the community from whence it came. If so, that’s a stronger argument, but I don’t know that it’s the best one they can make.
Writing software just auto-completing from CoPilot would be like trying to write a whole novel with the auto-predictive text on your phone. You could do it, but the results would be non-nonsensical and full of semantic errors. I don't think there's any real 'provenance' at play in either case.
This guy is a literal who, who is really over estimating the value of his open source contributions over a general development tool that can reduce the cognitive load of working in some hairy code bases.
I don't think it's CoPilot that's 'erasing' his, Racket, community.
Reminds me of people who defend AI art with "you can tell it apart from real art" yeah no, at the rate we're moving you won't be able to tell at all in a year or two.
Copilot is an AI stunt, an exploration, trying something new and exciting with very mixed and not-so-useful results.
This lawsuit, however, is just lawyers doing what they do for fun. I guess the retained Microsoft lawyers love it too. Glad to see lawyers having so much fun and profit. But we would all be better off with out so much lawyering, can't they do something more worthwhile?
The license and attribution are stripped from regurgitated copied code snippets from code projects. If the people don’t known which project the code was taken from, how can they one day contribute to that codebase? If the code projects on GitHub are not getting the people who use their code at least aware of the project, that project disappears. Copilot is an interloper who doesn’t even tell you which project the code snippet was ripped off from!!
If I solve a problem and I say: Sure, use my code if you want, but be sure to contribute any improvement back to the community - I wouldn’t be happy seeing a tool spewing it out everywhere. And it’s not even free. In this case, microsoft is literally making money on the back of millions of programmers. And without approval.
Generating snippets of code has nothing to do with a fully functional software package/product/service and an organic community around it.
One might argue that a community could be more easily formed thanks to co-pilot because it increases developer productivity and lower the effort to contribute so OS projects actually benefit from co-pilot. If this sounds far fetched, then probably the first case is also similar.
But they deliberately don’t tell you thst… it’s just a code snippet floating in space like they invented it
If the people don’t known which project the code was taken from, how can they one day contribute to that codebase?
If the code projects on GitHub are not getting the people who use their code at least aware of the project, that project disappears.
Copilot is an interloper who doesn’t even tell you which project the code snippet was ripped off from!!
People are conflating their open source license with the one they give GitHub when making a GitHub account, but they are two entirely separate and parallel licenses. The former is for other people to use your code, the latter is for GitHub to host your code.
If you don't like it, you are free to host your code on your own servers.
And anyway, as noted the other day about AI, it is often funny to see people not care about (or even enjoy) AI in other fields that they don't work in, but when it comes for their own field, they are suddenly very worried. See programmers on HN who argue for Stable Diffusion but against Copilot, and vice versa with artists on Twitter. As I commented then, it's an act of cowardice to think our own profession should be immune from AI while we enjoy the fruits of AI in other fields [0]:
> Yes, many of us will turn into cowards when automation starts to touch our work, but that would not prove this sentiment incorrect - only that we're cowards.
>> Dude. What the hell kind of anti-life philosophy are you subscribing to that calls "being unhappy about people trying to automate an entire field of human behavior" being a "coward". Geez.
>>> Because automation is generally good, but making an exemption for specific cases of automation that personally inconvenience you is rooted is cowardice/selfishness. Similar to NIMBYism.
We should want AI. That we then try to use outdated models like copyright to enforce holding back human progress is a true shame. In my view, so what if GitHub uses people's code for training data, we are all getting a better product because of that.
For instance, what if I self-host an OSS project but someone puts a mirror on GitHub? Or just uses GH as a remote for their fork? Does that random person accepting the ToS now mean GH has carte blanche to do whatever they want with that IP?
https://docs.github.com/en/site-policy/github-terms/github-t...
4. License Grant to Us
We need the legal right to do things like host Your Content, publish it, and share it. You grant us and our legal successors the right to store, archive, parse, and display Your Content, and make incidental copies, as necessary to provide the Service, including improving the Service over time. This license includes the right to do things like copy it to our database and make backups; show it to you and other users; parse it into a search index or otherwise analyze it on our servers; share it with other users; and perform it, in case Your Content is something like music or video.
This license does not grant GitHub the right to sell Your Content. It also does not grant GitHub the right to otherwise distribute or use Your Content outside of our provision of the Service, except that as part of the right to archive Your Content, GitHub may permit our partners to store and archive Your Content in public repositories in connection with the GitHub Arctic Code Vault and GitHub Archive Program.
--
[1] the practicality³ of this is a different, though related, discussion
[2] because the user is fully informed and can take responsibility for the decision to use the suggestion or not
[3] or impossibility – given the code could be added by someone who doesn't include that attribution/licence information for the system to be able to pass on even if it were designed to
I'll stick to self-hosting instead of using services like GH. Keeps things a little more simple in that regard.
Copilot on the other hand basically defaults to infringing behavior. Users would have to go to great lengths to be sure they aren't infringing on others work.
And if you didn't have the rights to grant the licenses to Github? Then you are in violation of the copyright holder's rights, not Github.
The only remotely plausible, yes-I-have-graduated-fifth-grade argument against Github is that they ought to and certainly do know that huge portions of their users are in fact granting them licenses without the necessary authority. That's an interesting argument we should be having, and instead we're having this inane screaming match by people who have no clue what they're talking about while some of us are sitting here going WTF is wrong with you?
I am discussing the license grant GH includes in its terms. And that doesn't appear to give them a blank check to do anything they want with code those users have uploaded. Certainly not sell it piecemeal.
"It also does not grant GitHub the right to otherwise distribute or use Your Content outside of our provision of the Service."
Throughout the license, the Content is treated as an indivisible unit, and it specifically refers to the forking functionality. Notice that forking...forks an entire repository, licenses included, etc. You can't fork a single file, and you can't fork a region of a file. GH provides that kind of forking.
Copilot is fine-grained forking.
No significant software company is going to permit copilot to be used and potentially poison their code base in unknown ways, now that this kind of copying is in the open and is clearly a significant danger.
Somebody like Black Duck is going to make a lot of money for trial attorneys by tracing how code was created and finding the "hits". That will be joined with log data indicating who used copilot, when they used it, and exactly what copilot presented as the "hit". This entire process will be performed recursively on the "hit", together with classic source analysis, to find out where something is really from.
The bigger companies are really, really serious about not copying outside code except under really strict conditions -- these conditions mostly look like "no you may not, unless you have one of these specific situations". It's no-by-default, even when it looks like it could be a yes.
You've ignored the important words. Copilot is part of the Service.
I don't know, sounds pretty similar to training on ML programs, even if they don't explicitly say "machine learning" in the ToS.
This would, at a minimum, preclude charging for Copilot.
This is missing the point though. Microsoft claims their use of source code for Copilot is fair use. If they are correct about that, licenses don't matter, this EULA doesn't matter, etc. Everyone should be focusing on this claim, arguing about any other detail before that is decided is a waste of time.
> parse it into a search index or otherwise analyze it on our servers; share it with other users
That is the most succinct and most accurate definition of Copilot I've ever seen.
From the article:
> “Dude, it’s cool. I took SFC’s advice and moved my code off GitHub.” So did I. Guess what? It doesn’t matter. By claiming that AI training is fair use, Microsoft is constructing a justification for training on public code anywhere on the internet, not just GitHub.
And:
> when it comes for their own field, they are suddenly very worried
From the article:
> First, the objection here is not to AI-assisted coding tools generally, but to Microsoft’s specific choices with Copilot. We can easily imagine a version of Copilot that’s friendlier to open-source developers—for instance, where participation is voluntary, or where coders are paid to contribute to the training corpus. Despite its professed love for open source, Microsoft chose none of these options. Second, if you find Copilot valuable, it’s largely because of the quality of the underlying open-source training data. As Copilot sucks the life from open-source projects, the proximate effect will be to make Copilot ever worse—a spiraling ouroboros of garbage code.
If I was writing this website, I would delete this sentence, because it is actually really idiotic.
1. This is an opinionated statement, except that there's nothing backing this opinion. This is fearmongering, that GitHub Copilot will get worse unless we sue Microsoft.
2. Lawsuits are not about concerns about a product's creator potentially damaging their own product. Lawyers suing Microsoft don't get a "we're protecting Microsoft from Microsoft's own bad decisions!"
The argument, completely unfounded, is that GitHub Copilot will undermine... GitHub Copilot. So, if you like GitHub Copilot, you should also be on board with suing Microsoft, so that we don't damage GitHub Copilot. What??? Good luck proving that line of argument in a court - you'd get laughed out of the room. Courts don't react well to hazy predictions about mayhem, from the suing lawyers, that have no historical facts to base them on.
I do exactly this, and it does all of bupkis to prevent someone from downloading my code from my gitea, and uploading it to GitHub. In fact, several people have.
Well... DUH. Why would they? You want to possibly sue them. Why in the hell would they, or anyone, provide crucial evidence for your lawsuit before you've sued them, regardless of the case and circumstances? Of course they aren't going to provide evidence, because you are obviously going to then try to prove hypocrisy, whereas you might not have enough to go on if they don't talk. No corporate lawyer in their right mind would ever grant such a request. (Edit: You are quite literally asking what their legal strategy is going to be, before the lawsuit has occurred, and then trying to spin the refusal as a proof of guilt.)
That's like claiming that an alleged drug dealer who didn't talk without a lawyer present is obviously a criminal, because if he wasn't he would have talked. What a nothing of a point.
If a company use someone's code for a commercial product (a normal app), they do need to follow the license accordingly. If a company use someone's code for a commercial product (model training), they don't need to follow anything.
If a company use someone's art piece for a commercial product (a normal game), they do need to get consent, and pay for the right to use to the hosting platform or artists themselves if it is not royalty free. If a company use someone's art piece for a commercial product (model training), they don't need to get consent or pay for anything.
All the problems actually happen before the technical details, making the entire pipeline questionable.
As a long-time open source software developer, I have favored the 2-clause BSD and MIT licenses because they are the simplest licenses that provide me some liability protection. I would release code into the public domain if that didn't increase the likelihood of being sued, whether for liability, or for someone else claiming intellectual rights to code I actually wrote.
Yet somehow I think most people upset about Copilot would not like that outcome.
If so I think that introduces a lot of other issues. There's an interesting phenomenon of different independent comedians suing talk shows for stealing their jokes. It almost always turned out that those jokes weren't actually stolen. Instead, there's only so many jokes you can make about a given news story and there's bound to be overlap between a whole room of comedians trying to milk every event of any comedic value and random independent comedians doing the same
I don’t think a banner would be sufficient here, though; perhaps some references to the inputs that were used to generate the output, but that’s often very difficult to pinpoint.
Whatever happens, if this ends up setting some sort of legal precedent it will have a big impact on the industry, and I personally hope it leads to more accountability and transparency of the models, rather than the black boxes they are now.
> in very rare cases, an independently generated code recommendation may resemble a unique code snippet in the training data. By notifying you when this happens, and providing you the repository and licensing information, CodeWhisperer makes it easier for you to decide whether to use the code in your project and make the relevant source code attributions as you see fit.
>"Tim Davis gave numerous examples of large chunks of his code being copied verbatim by Copilot, including when he prompted Copilot with the comment / sparse matrix transpose in the style of Tim Davis /."
Copilot regurgitates code and blatantly violates licenses, not even sure what there is to argue about. Not only does it seem straight up illegal and sideline open source communities, I think the next logical step of this is that people who want to avoid having their work vacuumed up and their rights violated simply to move to proprietary software, which would be a huge disaster for open source.
(I am not a lawyer).
For Copilot to blow up, it'd need to be licensed code from a big company demonstrably turning up in a product of a competitor, or some similar event.
If you released AGPL code but never intended to ever sue anyone. Why did you release it like that?
If you did and if someone is able to use your code without any damage to you, without reputation loss, and via a way they have access to the innocent infringer defense after you overcome fair use, after you sue them.
How is that game over?
If you ever went on a hike but never intended to sue anyone. Why did you go out in the first place?
If you did and someone is able to punch you in the face without any lasting damage, without reputation loss, and via a way they have access to the myriad legal defenses you couldn't come up with if you tried, after you sued them.
How is that game over?
Just because someone corporation is, because of its sheer size, over the law (as far as a John Doe is concerned anyways), does that make it a right? We could probably do away with laws at that point and just accept getting punched in the face by Microsoft whenever they feel like it as the new reality.
The court system is the method of enforcement for copyright.
If you want the "right" in copyright, you have to sue people.
To sue people, you need to find infringement. That infringement must be above fair use.
However - if the infringement you find is so minor that you have no loss of revenue or reputation, a court will not award you damages, and may even dismiss the case.
Nobody has any copyright without suing people, there is no copyright police in the general case.
Microsoft has no special rights from its size. Its size makes it a target, it's not beneficial. It's why they have so much trouble with internal rules about GPL. If I infringe on your copyright, the damages will be zero or low, if Microsoft infringes your copyright, the damages could be millions - with the same burden of proof.
CoPilot is a great research work - it is indeed spectacular to see how pre-training can achieve such impressive code completion results. However, in my honest opinion, it should not be a tool for a serious developer.
This is particularly true of fair use, it is very fact specific. A court is much more likely to answer a very fact specific question about copilot, tied to the very specific facts of the case (IE how is this exact thing used/etc) than more broad, abstract questions.
In fact, standard Article III courts in the US are literally not allowed to issue advisory opinions.
Also, programmers please do not hinder on other programmers work. If you do, someone higher up in the ladder with eat your cake at every opportunity.
Is there a reason why an AI being trained on the same open source code isn't a similar situation? I agree that wholesale pasting of code chunks is an issue, but that hasn't been my experience with Copilot.
I'm not arguing for Copilot here...I'm genuinely curious why this would be considered any different.
You are a human. You know what's right or wrong. You know you can't just copy code 1:1 from public repositories without respecting their license. The AI doesn't know and doesn't care. It's a common problem with creative AIs that they will occasionally regurgitate near 1:1 copies of their training data, and I don't think it's an easy problem to solve.
>I agree that wholesale pasting of code chunks is an issue, but that hasn't been my experience with Copilot.
The article provides several examples of it happening. Just because it hasn't regularly happened to you doesn't mean it doesn't happen.
What I don't understand is why the rest of the service (where it doesn't appear to be pasting existing code) is being maligned when it behaves like a more powerful version of autocomplete.
We shouldn't shed any tears for a megacorporation which shows such blatant disregard for the licensed works of people's labour.
Yes, AI is here to stay but we should be able to build AI that respects copyright. Yes, it's easier to just steal data and call it fair use. Whether or not that's stealing will be interesting to try in court.
ML training is akin to reading or learning, and licenses do not apply to that.
You’re not thinking past “megacorp = bad”.
Also, If we had trained some A.I. on the Windows codebase and started freely using suggestions given by it I bet Microsoft would scream copyright infringement in a heartbeat.
First, it would be nice to have a copilot variant that searched only my own work, so I wouldn't need to grep through other code I've written to get a reminder of how I solved a problem in the past.
And, speaking of the past ...
Second, I am old enough to have seen slide rules being replaced by calculators. This was a great addition to the toolbox, but it also had its downside: I've seen many students who have very clouded notions of significant digits, and many more who get quite confused with where to put a decimal point, when I ask them to compute something simple by hand.
Similarly, coding has been transformed with the advent of stack-like systems. There are two communities of coders now: those who learn a language and then can solve problems based on a solid foundation, and those who shorten the learning phase and code by web-search. The latter, it seems, are in danger of creating code that is brittle, limited, or downright wrong.
To the extent that copilot amplifies this habit of searching instead of thinking, I think it may lead to unreliable code.
So, sure, there are copyright issues. I think they have been well-discussed here and elsewhere. And courts may weigh in with new ideas. But my concern is with the reduction in code quality that may ensue. I'd love to see a discussion of the groups that are using copilot. If they are working on something I don't care about, then this is just a copyright issue. But if they are working on the "smarts" behind drug discovery, the control of dangerous machines, etc., then we have another issue, besides copyright.
>Meanwhile, we open-source authors have to watch as our work is stashed in a big code library in the sky called Copilot. The user feedback & contributions we were getting? Soon, all gone.
I don't see how you square the above complaint with this:
> First, the objection here is not to AI-assisted coding tools generally, but to Microsoft’s specific choices with Copilot. We can easily imagine a version of Copilot that’s friendlier to open-source developers—for instance, where participation is voluntary, or where coders are paid to contribute to the training corpus.
Is an AI that was trained on opt-in or paid-for training data any less damaging? How would these choices have alleviated the problems described above?
Also back then, Microsoft had 90%+ share of the PC operating system market and was bundling IE in its operating system. I’m glad the DOJ forced MS to change its ways.
By the time MS bought Nokia, it was already a has been in mobile and the acquisition was a total failure and the game market is competitive.
IMO, if the lawsuit goes to a point where it's likely to be won by copyright owners, Open AI could do the following:
- Use less sensitive code from big corps they have partnership with for training. I bet MS and other have plenty of such code.
- Buy training rights from copyright owners of OSS projects. Many of them have SLAs which allow the owner do much more than the license allows.
- Buy rights to train code, and collect generated code with Copilot from a large number of smaller software companies, likely with exclusions for some sensitive parts. MS has a lot of leverage here (discounts, partnerships, etc).
Locking up code under non-permissive licenses stymies the pace of code development and increases the costs of progress dramatically.
We all stand on the shoulders of others before us. Including the organisations that stand to benefit the most from aggressive licensing.
I put my time and my effort to open source a program for free and I want to make sure that my code creates an incentive to create more free software, by using a copyleft license.
For example, if an overseas firm can just as easily use Copilot as I can write original code (or use Copilot myself), why would any company hire me locally?
And if you dig further, the whole repo is mixed with BSD and LGPL "licensed" packages together. It's probably best that CoPilot does not suggest from code that does not have an explicit license stated.
I think originally Tim Davis was complaining about the non public sources for suggestions which Github CoPilot ignored.
This is the same case as with copyrighted photos in newspapers, a paper prints a photo somebody allowed them to use, but then it turns out that person did not have the right to use it in the first place. Did not stop newspapers from printing photos at all.
Here are the search terms: https://github.com/search?q=cs_transpose&type=Code
Whether or not his style his widely known, his code is VERY widely used. Just look up SuiteSparse and try to find all of the downstream uses of it. It is one of the most---if not the most---ubiquitously used set of sparse linear algebra libraries. If you do anything with numerical linear algebra, there's a good chance you at least know what SuiteSparse is, and possibly also know who Tim Davis is.
The bigger issue here is the effect this has on research. Tim Davis not only programmed this library, he did the basic research leading to many of the algorithms in SuiteSparse. He went ahead and released SuiteSparse open source, probably thinking that it would be a good deal for him, provided that its use was properly attributed. Provide a public service in exchange for attribution. This is a reasonable way to get support as an academic. Clearly he has had a large number of industrial collaborations which likely have provided him with a significant amount of funding over the years.
Speaking for myself, if Microsoft has no compunction against behaving this way, I can no longer see the point in publicly releasing research code that I develop using an open source model. Microsoft is clearly telegraphing that they don't give a f** about licensing, although whether that holds if they are litigated against remains to be seen. I think there's an excellent chance many other researchers feel the same way. If you think openness and reproducibility in science is important, this is a problem.
Can someone help me to imagine a reality in which these points are viable concerns?
> …how will you feel if Copilot erases your open-source community?
> …Copilot will become not just a substitute for open-source code on GitHub, but open-source code everywhere.
> …Copilot is merely a convenient alternative interface to a large corpus of open-source code.
> With Copilot, open-source users never have to know who made their software. They never have to interact with a community. They never have to contribute.
Is the author suggesting that Copilot will be used in place of `npm install next react react-dom` or `cargo add tokio --features full` or `raco pkg install pollen` — that developers will be content to use augmented autosuggest in place of large, well-tested, well-documented open source libraries?
Does he see Copilot's final form as some kind of AI package manager that drops a library of untested unattributed undocumented files into our projects?
Or is it more that he thinks those libraries won't exist because open source contributors will grow to feel more abused than they already do, perhaps quitting the scene or developing in private, like certain artists have already done in response to the AI art movement?
There is already such a huge disparity between paid package consumers and unpaid package contributors. I haven't seen that change since Copilot launched in beta or under general availability. I see the same ratio of help/feature requests compared to code and documentation contributions that I always have. And package usage has not declined so far for the open source things I work with.
It would be nice to learn more about the “Copilot will lead to the death of open source communities” line of reasoning — what is the author's perceived timeline to open source's decline and fall as a result of Copilot's current path?
"Your work is under copyright protection the moment it is created and fixed in a tangible form that it is perceptible either directly or with the aid of a machine or device" [https://www.copyright.gov/help/faq/faq-general.html]
A copy is made whenever that text is displayed, e.g., in GitHub's UI. Even that copy is subject to copyright.
Is there an excuse/exception? In this case, there is no "fair use" exception, because exceptions have to be litigated case-by-case to be recognized, and there are no remotely similar situations. Don't forget: Lexis is a multi-billion-dollar business built on protecting the copyright to the page numbers in the otherwise public court opinions.
Does the law actually protect people if it's too costly to enforce? Not really; hence the blase attitude. Congress is considering a "small claims" system for copyright, to remedy the big-firm bias. [https://www.copyright.gov/title17/92appm.html]
In the ML era, data is the new gold. Many, many firms nowadays get a good chunk of their revenues from selling their private view of "public" data: Facebook, LinkedIn, credit reporting companies, ADP, etc. Microsoft has gone all-in on stealing that gold from open-source developers.
It's not just that the code replication reduces any need to get the code from the source. But removing any link to the source destroys the value most-commonly sought in open-source software: recognition.
Salaries are the biggest expense of tech companies. They do everything they can to increase labor competition and reduce reputational rents: outsource, cross-train, promote open-source (for competition) and destroy any reputation networks or systems that justify higher rates. And, of course, standardize on containerized copy-paste or AI-generated software if they can.
So, no: copilot is not legal, it's socially and economically destabilizing, and it presents structural challenges to developers.
It's not good, but most will keep using it because although the vast, vast majority of developers are wage laborers, they aspire to be founders. They see it can make code fast, and they'll think it make them better.
It is also very telling that they have not included any of their own proprietary code in the training set. If it's merely suggestions that are generated, why not also train on the NT kernel? Office?
I want my code to be used the way people treated text in the old days. There's texts that have been re-written, added to and edited by thousands of people over the centuries and yet they don't come with thousands of pages of attribution notices because why would they?
Whether Copilot is breaching MIT depends on what constitutes a substantial portion, which I am not qualified to rule on.
If tomorrow someone released a StableDiffusion, CoPilot etc with the same functionality, but respecting the provenance of the data (i.e. licensing etc), what concrete difference would this make? Programmers and other creative professionals would still (reasonably) be nervous about the implications for their livelihoods and communities.
At some point it will be possible to prompt a model for music in the style of <random artist>, and having never heard <random artist>, the model will generate a convincing emulation, based purely on statistical knowledge gleaned from millions of unrelated songs and text pairs. (I give it 5 years).
Now what? <random artist> should still be concerned (or not), but at least we're talking about the correct issue: How do we co-exist with generative models that massively disrupt/alter the process of doing creative or intellectual work?
The concrete difference you ask about is that the work derived from copyleft code retains the license and the source code can't be closed. If you scrap the license, then the code created by someone who had clear goal in mind when writing it for not making improvements over it closed source, ends up with possibility of being closed source.
"I used the copyright to destroy the copyright."
That sort of plot never works in practice.
I think you need to explain that more. The problem (or at least one problem) being explored here is that by using any code from co-pilot, you are responsible for making sure the licensing is correct. You could unknowingly be using and modifying GPL-licensed code in your non-GPL project, which is a violation if you don't publish your modifications.
We're not talking about expanding copyright, just protecting the existing copyright systems from being trampled by microsoft.
This is extremely far fetched.
User bases (let's avoid one of the four dirty C words) are organized around something which builds, executes and is documented, not searches for snippets.
Basically, what Copilot (or anything like that) is supposed to do is to speed up your work, i.e., ideally, to write exactly what you'd write, but orders of magnitude faster. How do you write code? Well, you may have a solution in mind — if it's something really original, rest assured, Copilot won't guess it. It can only hope to guess something that, in a sense "has a correct answer" to it. In fact, it does it worse, than it should be: graph traversals, matrix operations and such should be guessed flawlessly (in a perfect world every PL would have some primitives implementing them in the best possible way, but ours is not perfect). If you don't know how to traverse a graph, you'll go and look for a reference. 15 years ago it was likely a book, then looking up on the Wikipedia or StackOverflow became way more likely. For the last 5 or so years literally searching it on GitHub became viable because of better search engines and the sheer size of it.
Now, if I found a matrix transpose function in an open-source project, which I cannot include as a library for some (usually technical, but maybe not) reason, so I memorize it, close the page and re-type it in my IDE, do I have to be restricted by its license? Then, doing so is obviously stupid, so how about me just copy-pasting it, while renaming some variables so that the teacher wouldn't notice? And, given that this is not my homework, there's no teacher and variables are named perfectly as they are — doing that is also really stupid, so I might have just copy-pasted it. So, how about now, do I have to publish my code under GPL3 now? Is this theft? If any lawyers say yes — fuck these lawyers. It is nonsense.
The vast majority of open source code would be almost entirely worthless (or more likely, would straight up not exist) if it couldn't be used in commercial products.
Open source software licenses were a mistake.
Agree about the rest.
I largely disagree with this article, at least for MIT, BSD, etc. training code examples. The small autocompletions, even if they are several lines long, sort of seems like fair use to me.
I do think that CoPilot should have an option to use a smaller model just trained in code that has very liberal use licenses, because I think the use of GPL, etc. licensed code is problematic - at least for me.
For what it is worth, I have a lot of Apache 2 licensed repos on GitHub (largely examples from my books) and I am pleased if my code contributed a small bit to the CoPilot training data. I also publish my recent books under Creative Commons, allow reuse, even commercially licenses: basically anything I do that might help someone, I am all in for sharing.
This would be a better approach than "shut it down".
Speak for yourself. I pay multiple streaming services, music and video, because I prefer creators be able to eat.
We are people, we have ambitions and families. We're not just a faceless corporation.
I want there to be more good music and movies. I want to support artists who create entertainment I enjoy. I go out of my way to buy physical copies of music from artists, wherever possible from the merch table at their shows or from their own websites. I pay to go see movies on the big screen (partly because I like the big screen cinema experience, but also because I understand "opening week revenue" is a key performance indicator for the success of a movie).
I thing copyright is old, outdated, and probably not really fit for purpose for forms of creative work invented in the last 50 years. But I also thing creative workers need to get paid for their effort (juist the same as software developers), and absent a FAANG-style set for companies employing teams of songwriters, musicians, authors, and the like - on FAANG-style salaries, copyright seems to be the option that is working (however badly).
I'll join your "abolish all copyright" crusade as soon as there's an alternative that at least likely to possibly work as well (or better) than the system copyright allows. Just abolishing copyright and erasing the publishing/music/movie/art industries without a transition plan isn't a thing I can support. (At least a transition the artists/editors/producers/writers/etc. I'll admit there's a large chunk of management and legal in the fairly abusive parts of the music industry I wouldn't shed a tear if they all became homeless and destitute overnight...)
Which is ironic given that this is Microsoft. Whatever happened to "don't use programmers' code without paying them" and the whole "proprietary software is better because it sustains the programmer"?
from their terms of service: "Short version: You own content you create, but you allow us certain rights to it, so that we can display and share the content you post." emphasis mine.
that's what they call the "Short version" of the following paragraphs, which are found here: https://docs.github.com/en/site-policy/github-terms/github-t...
they allow themselves the right to display content you upload to others. GitHub does not seem to really put a cap on that in terms of what intentions it needs to have or for what purposes it needs to share your content.
this seems to me that, by putting your code on github.com, that you are granting GitHub license to show it to others. period. IANAL, but it seems like all code anyone puts on github.com is dual-licensed, at least. GitHub gets their own rights to your code.
I read this before I signed up, and while I can't remember if this exact passage was present at the time, I was ok with everything GitHub wanted at the time, and I continue to be.
githubcopilotinvestigation.com doesn't seem to have much hope of doing anything except getting people mad. but you all were already mad anyway, weren't ya?
This seems to be the line of argumentation agreed upon by several waffling pro-GitHub posters. Many comments have some variation on that diversion from the issue.
> GitHub gets their own rights to your code.
This is preposterous and false. GitHub has the right to display the entire work, properly attributed and licensed, to others.
No new licenses are given, no dual-licensing takes place, no code-laundering is permitted.
I suggest you read the terms of service again.
here, I'll link directly to the license grant: https://docs.github.com/en/site-policy/github-terms/github-t...
This license does not grant GitHub the right to sell Your Content. It also does not grant GitHub the right to otherwise distribute or use Your Content outside of our provision of the Service, except that as part of the right to archive Your Content, GitHub may permit our partners to store and archive Your Content in public repositories in connection with the GitHub Arctic Code Vault and GitHub Archive Program.> We need the legal right to do things like host Your Content, publish it, and share it. You grant us and our legal successors the right to store, archive, parse, and display Your Content, and make incidental copies, as necessary to provide the Service, including improving the Service over time. This license includes the right to do things like copy it to our database and make backups; show it to you and other users; parse it into a search index or otherwise analyze it on our servers; share it with other users; and perform it, in case Your Content is something like music or video.
https://docs.github.com/en/site-policy/github-terms/github-t...
Im unsure if vscode etc submit samples or just interact with GitHub.
Edit: and furthermore, make sure it doesn’t import code from third parties. I don’t want my code being infringed upon, but also don’t want to accidentally infringe on others’ work. Legal or not.
For example, there are plenty of academics these days who are at the tops of their fields and open source all their code. They end up considered as experts not because of a black box code base they implement on problems, but because they can think of potential solutions to the problems at all, and one of the tools used is writing up some code. The code is a shovel or a hammer, its not the one wielding it. They have competitors too of course, just that the secret sauce isn't the code but what goes on in your actual brain.
Its too bad most business leaders fail to understand this, and think its a blackbox code base that makes a decent business. Its the ability to solve problems that matters.
It's face-saving.
https://justoutsourcing.blogspot.com/2022/03/gpts-plagiarism...
Copilot may produce results from the training set, but if you're letting it do that, that says more about you than about copilot.
All of these claims use the example "Write me a function to foo the bar that takes baz as an argument". If you prompt it to write entire functions and classes for you, then it will lean on its training set.
But if you actually just write code, then it will complete small single lines in exactly the style you've previously written. With code that is unique to your program because it can synthesize new code.
In this role copilot is no different than a search engine. By prompting it lazily, copilot isn't the one stealing the code, you are.
I’m not even 30 yet and I’ve seen this happen again and again - it’s frankly boring at this point. We’ve seen this with Spotify and music, newspapers and the internet etc.
The practical truth is that Copilot is a useful tool for humanity to have. It is exceedingly unlikely it will be stopped because a small percentage of programmers - themselves a small percentage of people who benefit from code - feel their interests have been hurt. Change or get left behind (but make sure to enrich some lawyers on a pointless suit in the meantime).
If the AI is "learning" how it works by studying public code then using its knowledge to create, that's okay.
But if it's just memorizing code and reciting it back, not okay. Just like if a human were doing this.
Of course we don't currently have ways to know the difference [that I know of] since AI is a black box.
Interestingly, current AI is not capable of truly understanding how code works and how it will execute, so it has to learn in it's own way. I suspect it can learn what valid syntax is, but I doubt it is aware of how the code will execute.
It's possible this is just a case of Overfitting. https://en.wikipedia.org/wiki/Overfitting
1. Redistributions of source code must retain the above copyright notice, this list of conditions and the following disclaimer.
2. Redistributions in binary form must reproduce the above copyright notice, this list of conditions and the following disclaimer in the documentation and/or other materials provided with the distribution.
3. All advertising materials mentioning features or use of this software must display the following acknowledgement: This product includes software developed by the organization.
4. Neither the name of the copyright holder nor the names of its contributors may be used to endorse or promote products derived from this software without specific prior written permission.
5. Use of this source code for the research or training of machine learning models is permitted.
This is no different whether you are honkler or Microsoft (other than in how vigorously or not someone may enforce it).
If you really believe that anything posted to the internet ‘belongs to all’ then I don’t know what to tell you other than you live in a fantasy land where Oracle Corporation does not exist. We might all prefer it if things were that way, but they simply aren’t, and that’s just tough.
Small enough pieces of code can't be copyrighted. No one would support an argument that I violated copyright by using the code "else if {" from some GPL library.
So the question becomes what is the minimal unit of copyrightable code? What if you wrote a nice big function exactly (or almost exactly) the same way as someone else did? Whose copyright are you violating?
For example, from TFA file:
0005e10 o f C o p i 302 255 l o t
https://en.wikipedia.org/wiki/Abstraction-Filtration-Compari...
https://en.wikipedia.org/wiki/Idea–expression_distinction
https://h2o.law.harvard.edu/cases/5004
Most of your code is probably not subject to copyright in the first place, regardless of license.
Copyright is meant to protect “useless” things like poetry and music.
I hope Copilot and similar technologies weakens the copyright establishment.
Do Business WITHOUT Intellectual Property - Stephen Kinsella http://www.stephankinsella.com/wp-content/uploads/publicatio...
Against Intellectual Property - Stephen Kinsella https://mises.org/library/against-intellectual-property-0
Part of the clause would explain that if you are okay with your code being trained on, then you're also accepting being okay with it being copied verbatim at some point down the line during code completion.
You do get a bit of tragedy of the commons where everybody wants to use the AI model but nobody wants their own code trained on.
I don't like the idea of a world where licensing and copyright law prevents us from enjoying the progress of AI. Caveat: I am not an expert on open source.
I don't mind it using my code because in my opinion, we as a software industry are way behind on where we should be and copilot is helping a lot of developers finish their projects quicker.
That said, software licenses should 100% be respected. I would hate for FOSS projects to start being sued over code. It's not in the spirit of FOSS, but neither is stealing code. Copilot should be doing a better job excluding code and none of this would be a problem.
I do understand how ML works. I know it's probably not possible with how it's currently done. That doesn't make it legal or ethical.
It would actually be great for everyone if it showed both the license and repo. Imagine you pull up a great function with Copilot and want to explore the source for more insights. You can't with how they've done this.
But since the use is not for the page author to comment on the comic itself, but the comic is used to support his discussion of another misuse of IP, does it constitute fair use?
The page author is going deep on the content misappropriation theme and on what constitutes fair use, so it seems oddly ironic he'd be so seemingly cavalier about using someone else's content on that page.
If these guys manage to shut down or cripple Copilot using legal mechanisms, you can bet there will be a Chinese/Russian alternative that will be even more indifferent to your LICENSE.md, and you won't be able to get it shut down using the courts.
It's a common advice to not read software patents[1] because the infringement penalties are lower if you did so unwittingly, that is, by reinventing the patented technique yourself.
I wonder if using Copilot doesn't push the penalties back again to wilful infringement. Or worse, patent trolls poisoning the training data with patented algorithms.
Absence of proof is not proof of absence.
They don't owe anyone anything beyond what they agree to provide to users of Copilot via its license agreement or to GitHub users whose code it has used in accordance with that license agreement. Those agreements define what they owe. That's it.
The only way those license agreements don't hold up in court is if they are somehow deemed invalid. I do not see Microsoft making that kind of mistake.
This website is designed to get people angry, and that's all it is going to accomplish.
The code is out there. Millions of people are being trained and writing code based of the learnings of open data.
Designers have "mood boards". Developers have open source. Right now I don't have sympathy for MS, but in a few years any you developer could just do what MS is doing with Copilot in their bedroom. Why would you care about the kid in their bed room training an AI with free (as in public) information?
Imagine if something like google didn't exist, and then it suddenly did. People would be saying: "This newfangled computer algorithm is giving everyone copies of my code with a misattributed licence, just by typing the function name and site:github.com !"
It would be better, of course, if Copilot was opt-in, but they'd never go for that.
It's hidden on both Chromium/Firefox when viewing the page but when saving the page it reveals them in the text field, eg: `GitHub Copi_lot inves_ti_ga_tion`
Plugging the title into a unicode converter shows they're 'soft hyphen' characters
GitHub Copi [0x00AD] lot inves [0x00AD] ti [0x00AD] ga [0x00AD] tion
Edit: apparently they're for indicating to formatters where character breaks should be, though I can't understand the consistency here.
The walled garden bit I get. But I'm lost making the leap to "remove any incentive to do so." Is Butterick suggesting that someone is going to put aside their code and do a deep dive on GitHub looking for a snippet that might not exist?
I'm not trolling. I'm sincerely trying to grasp the argument being made.
https://news.ycombinator.com/item?id=33239706
https://waxy.org/2022/09/ai-data-laundering-how-academic-and...
But this is solely affecting a product where we are the target audience, where if we oppose, thing should change. Now I wonder if it will show that we are actually caring that much to action or we are just as regular consumers as non-tech people in all other cases.
I agree with the article's long term outlook about community and code quality, it is a very long term outlook though. It makes me wonder if humans will actually be writing code.
It's only a matter of time before intelligence agencies will get their hands on the data. And if use of Copilot becomes an industry wide practice then those who wish to preserve their privacy will become uncompetitive.
I really hope we have some decent offline alternatives eventually.
To be fair, record companies were not in the least bit sympathetic. Open source contributors are easier to identify with, though imo it doesn't actually make their concerns more valid
I cannot parse what you are suggesting.
Copilot exists publicly, which also means some copilot-lite thing trained on a smaller subset of repos probably exists privately in many different places. It may not be as good today, but these private instances will improve over time. Since the demand for a copilot-like service exists, eventually a large VC-funded public instance will show up.
In that lens, it is more sensible on the individual level to prepare for a world where copilot thrives than to put all of your eggs in the "ban copilot" basket.
I think 'training an AI' is actually a distinctly new use of IP, and should probably be considered under a specific kind of 'AI-use' license. Open Source licenses should be updated to indicate whether they allow or do not allow AIs to be trained on covered work as well as the other rights they allow.
How cool it is to discuss these kind of issues? What do you think about the "erasing open source community" argument from the historical perspective? What does it have in common with industrial revolution?
Even though the real life implications are real, I find it fascinating and not so simple to unravel.
They are “investigating” whether they should start a law suit. So this is not an investigation, it’s somewhere between “due diligence” and a PR stunt.
I very much disagree with the idea of a law suit that seeks to establish ML training as not being fair use. It is an utterly foolish thing for them to wish for.
Kinda like how gmail was reading everyone's emails and showing ads based on them.
It will be interesting to see where this goes.
I think the fair use violation he describes doesn't happen during training. I do think training AI on anything that is publicly accessible is fair use just as in an example of a person learning by reading/watching the same materials.
However, this fair use rule is being violated the moment the resulting AI starts suggesting verbatim copied code from licensed works without attribution.
So one could argue the source code is not being used in a transformative way but copilot is just more efficient method of retrieval of licensed code. This misses the fact copilot actually is capable of writing new code. I've used it as "an autocomplete on steroids". Letting it suggest maybe half a line, or 1 line of code at a time (or trivial stuff we automate even without copilot like getters/setters in java). But when actual licensed code is suggested yes, this is IMO a license violation.
Therefore one way of resolving this would be to pair copilot with a tool that scanned the resulting code for presence of licensed code then it woukd make a list of "credits" or references. Also there should be measures taken (perhaps during training) to penalise generation of verbatim (or extremely similar) code. Would this make copilot less of a useful tool? I'm not sure.
One thing that's not going to happen is putting tools like copilot back "in the bottle". We now have similar models anyone can download (faux pilot) and I as well as many others have found those tools to speed up mundane tasks a lot. This translates into monetary advantage for users. Therefore there is no way this will disappear, lawsuit or no lawsuit.
First, the huge majority of open-source projects are at no real risk because Copilot offers something totally different from what they offer. Open-source projects generally take highly-complex domains and expose them as simple interfaces or executable programs. This encapsulation is where the value lies.
In contrast, Copilot just dumps code. Never once doing front-end work have I thought "if only there was a way to dump verbatim React internals directly into my codebase." In general, Copilot only replaces tasks I would have otherwise done myself.
The second problem is the biggest loser if Copilot gets shut down is not Microsoft, who can easily take the loss in stride. The real loser is the community of developers, many of them bootstrapping their own projects or trying to develop open-source in their precious off-hours, for whom every minute counts, and for whom tools like Copilot can be the difference between success and failure.
Can this AI regurgitate the vast majority of the creative aspects of an original/novel piece of software with minimal prompting, to the point where the output code looks mostly and directly cloned to a reasonable person trained in the art?
So "Can this AI regurgitate the vast majority of the creative aspects of an original/novel piece of software" is not the test that, for example, the music industry uses when determining if a sample is infringing. The test there is "is a sample, however small, identifiable as part of a copyright work by a reasonable person trained in the art?"
You can't own copyright in a composition of a single middle c note. But lawsuits have been won for copyright infringement of melodies of 2 bars (fewer than about 16 consecutinve notes). Men At Work lost a copyright case for the flute melody in Land Downunder which is the same as a 90 year old tune Kookaburra Sits In The Old Gunmtree https://www.claytonutz.com/knowledge/2010/february/men-at-wo...
Whether that's done by a flute player or an AI, really doesn't make any difference as far as copyright law sees things.
(Whether copyright law is a "good fit" for source code, and whether it makes sense to apply laws meant for books/literature/music/film to software is a different but very good question. I don't have much in the way of other ideas which take original author's efforts and potential rights to benefit from then though...)
If you can't train on data before asking for permission, the data set becomes sparse. Thje only people who will be able to afford this will be, you guessed it, established giants who can build their own sets.
The same thing will happen to source code produced by AI code generators. Github itself, or some entrepreneur, will come up with a way to identify and flag projects containing AI generated code based on models constructed from open source projects, so that those derivative works will not inadvertently be incorporated into other software that is concerned with such a flag. (They probably will also come up with an NFT-based mechanism of some sort to allow open source project rights holders to authorize incorporation of their code into AI models such that derivative works containing those fragments would not be subject to flagging.)
Hey YCombinator, give me $10M to make a billion dollar company that "lives at the intersection of" blockchain and open source. (Haha, No.)
How will you feel if greed of a lawer erases progress of your tools?
Lawyers are a detriment to anything they touch. Letting them into software was the biggest mistake we ever made. We should kept them away same way they are kept away from math.
With all that intelligence if GitHub Copilot can't produce easy to use and manage full stack framework yet with distributed database inbuilt in either any existing programming language or perhaps a new one created by itself then its not useful for me.
Especially with algorithms.
I was rooting for Google when the JVM topic happened and I'm rooting for GitHub with autopilot.
And yes there is src from me on GitHub too but use it! I used so much other code in the last 15 years.
Copyright on algorithm or basic code should be a no go.
I see Copilot as a net positive. Open source is for sharing and learning. Copilot is sharing and learning on steroids.
The license and attribution are stripped from regurgitated copied code snippets. Verbatim with no context, no attribution, no citation, no reference to the project it’s part of…
If the people don’t known which project the code was taken from, how can they one day contribute to that codebase?
Copilot is an interloper who doesn’t even tell you which project the code snippet was ripped off from!!
The following is supposed to be OK: somebody reads your GPLed code, learns abstract concepts from it, teaches it to me, I write code that uses the same algorithm. But it's not OK to abbreviate the process and reach the same result directly with Copilot. That is some Talmudic level reasoning. In a sane legal system, one would note that it is legal to do when jumping through pointless hoops, so it should be legal per se, and the system should be adjusted.
Copyright is increasingly at odds with technological development. Not just since AI applications, at least since Napster or since floppy disks. Of course Matthew Butterick as a lawer would disagree - "It is difficult to get a man to understand something, when his salary depends on his not understanding it."
The trouble is that this apparently is not what Copilot is always doing. If it had only "learned abstract concepts" from GPL'd (or any other form of copyright) code, then that would not be a problem, and of course that is kind-of what Copilot purports to be doing, supposedly learning the association between concepts described in comments and corresponding forms of implementation.
However, apparently Copilot is sometimes NOT generating it's own code based on the concepts it has learned, but is instead just regurgitating chunks of potentially copyright-protected code verbatim. It'd be interesting to know if it is doing this deliberately (to maintain the coherence of what it is generating) or not - I guess the more of something it has already copied exactly the more it is likely to continue copying since that is the best "predict next word" continuation. Of course while it would be interesting to learn more about the mechanics of Copilot, that doesn't change the legality, or not, of what it is doing, another aspect of which (although IANAL) is how much of the original work is being copied.
At the end of the day it shouldn't matter whether it's you or Copilot either learning from or copying someone else's code - exact same copyright protections apply.
I’m also curious to see if/how Amazon CodeWhisperer takes advantage of this whole debacle.
Just imagine how in a lawsuit like this, OpenAI can use GPT-3 to generate eloquent court speech with statistical confidence that it can defeat human lawyers? It just comes down to TPU power.
Will the lawsuit fall on the overseas developer, US developer or Github?
1 - https://www.wsj.com/articles/whats-in-a-hedcut-depends-how-i...
lol rarely see such aggressive use of soft hyphens in page titles
this is why hapas are superior to wh*tes.
Ta da!
Knowledge data should be free to copy and do whatever we want with it
I'm more of a copyleft fan. Feel free to copy my stuff, but you have to make it open source as well.
For some reason when a company benefits from your work instead of some other entity that's bad? Please explain.
it wouldn't cover cases where people illegally copy pasted some code into their projects with dubious / not explicit licenses, but this is the same as using any open source project in general.
It scrolls past all the repos movie-credits-style. Doing it that way takes several days! It shows how abstract and absurd giving contribution to such a large body of works is.
how is copilot doing something fundamentally different?
Microsoft is distributing his software without a license, isn't it?
Fair use in code is broader than just copying. For example, in Google v Oracle, APIs were found to be not copyrightable. Even if you copied the names of, say, 86,000 different functions in a proprietary library, you did not violate copyright.
Then comes the second problem. Let's say there is a function, say, `AddTwoNumbers(int a, int b)`. Just because John Fitzgerald in 1999 implemented that as `return a + b;" doesn't mean you can't too. There's a degree where you can copy the code that made a function work, even if that code existed earlier. It's fuzzy but it is legally real.
Finally, there is your third problem, which is that you risk a "safe harbor"-esque judgement. Just because YouTube has occasional copyright-violating content doesn't make YouTube illegal. Similarly, the person suing here risks a finding that GitHub Copilot is legal as long as any occasional long proprietary code regurgitations are removed as needed.
If your code falls under the first two conditions, copyright be damned, license be damned, it's all irrelevant. See also Linux copying Unix.
I don't know why one can freely pile on, e.g., AirBNB here but Copilot is a sacred cow.
If you think what copilot is doing is ok, and there is nothing wrong with it, I'd love it if you could go through this small thought exercise, and see if it impacts your view at all:
Say you write a bunch of code, and release it under GPL. For the sake of argument imagine it is something complicated that you care about.
Now say another person is trying to do what your code does, and they find your code, a say "excellent". They then copy and paste it into their project, and release their code under a BSD license instead.
Would you consider this theft of your IP? The law certainly would, and I think most devs would as well.
What would you say if they instead release "their" code as public domain?
Now we'll go a bit further. Another person is trying to solve this problem in some commercial software. They find your code, copy-paste it into their project, then sell their software and don't release the source, or even acknowledge you.
Would you consider _this_ theft? again the law would.
Now, what if instead they found your code through the invalid BSD relicense? or the invalid public domain one?
To me every one of these would be theft, and every one would be required to required to release the source of projects that made use of my GPL'd code, under the GPL. That is literally the whole point of the GPL.
But let's imagine a different route.
A person is writing some code and can't work out how to solve a problem, so they ask on StackOverflow. Now another person comes along and answer the question by copy-pasting from your project into SO. The first person says "yay!" and then copies that code, and we repeat the above scenarios.
In an even more extreme case, imagine both of the above people work at the same large company - so neither knows or is even aware of the other - how does this impact what is going on? It's two people, but fundamentally the company is copying the original GPL code into SO, then copying it from SO into its proprietary code.
I get that MS and GitHub try to position it as if copilot is "creating code", but it is simply doing a statistical code completion that is demonstrably happy to copy and paste from the original source into the recipient code. To my mind all it is doing is providing a mechanism to launder GPL (or whatever) code into your own without the license, by slapping "ML" and "AI" on the process and requiring more than 3 keys to be involved.
Let's be honest, copy-and-pasting happens all the time. In software, in engineering, in marketing, in everything. Whether people acknowledge it or not.
Everyone looks at Stack Overflow all the time. You do it. I do it. Nobody reads the licensing terms. We all produce software with reskinned and taped together functions. A collage is still unique, creative, work despite being glued together with other people's art.
Most songwriters will write a song with a part like someone else's song.
The products you buy at a store are rip-offs of someone else's product.
Everybody stands on the shoulders of giants before them. Such is learning, such is life. Get over it.
If I copied code without the rights to it into code I have written at any company I would absolutely be fired. It would not be up for debate.
> Everyone looks at Stack Overflow all the time. You do it. I do it.
I don't - it is very infrequently that I would look at SO answers
> Nobody reads the licensing terms.
Yes, they absolutely do, because again OSS or the GPL is meaningless if people are ignoring the license. More over I would suggest you talk to your employer's legal and IP departments to let them know you're copying code you don't have rights to into their product.
> We all produce software with reskinned and taped together functions. A collage is still unique, creative, work despite being glued together with other people's art.
Wow.
Absolutely not.
Competent engineers know how to write code themselves, they aren't copy-pasting their way to a solution. That's why you get paid a lot - if I was happy with copy pasta solutions I would hire a bunch of badly performing uni students.
> Most songwriters will write a song with a part like someone else's song.
If two people write similar songs that does not mean one copied the other. If one person copies part of another person's song they will end up in court, and they will lose all the revenue from the entire work.
> The products you buy at a store are rip-offs of someone else's product.
The ripoffs that stay on the market are not made by copying the entire implementation.
> Everybody stands on the shoulders of giants before them. Such is learning, such is life.
I didn't say anything at all to imply that we didn't. What I said was you don't get to just copy other people's work and pass it off as your own.
> Get over it.
Just because you apparently can't actually write new code yourself doesn't mean that that applies to other people.
Also, you should really get your employer to tell you whether your proposal of ignoring copyright is ok.
What you are saying is that if I google some problem, and copy some code from, say, gecko, or linux, or gcc, etc into my proprietary closed source product, that perfectly ok and those silly open source people should keep it to themselves if they don't want me doing so.
But there's also a difference between "I searched for this and copied the code I found" and "I typed some letters
Look how Google News enraged news orgs.
Now they come for the programmers. So now it’s a problem.
If you write an article about good writing, and quote a choice paragraph from someone else's work to show an example, and credit that quote, that is fair use.
Is it fair use if you read an awesome paragraph, something that really is the result of the authors unique intellect and effort and craftsmanship, and makes you think "damn", and then drop that same jewel into your book?
You can probably get away with it, because you probably just won't be able to convince a judge that any single paragraph is that big of a theft.
But I don't mean to ask if you can get away with it, I mean to ask if it should be considered fine honorable behavior.
The difference is, the paragraph isn't being included for examination or comment or transformation, it's being included to directly copy and perform it's original function as part of what makes a work a great work, and, it's not being credited in any bibliography or footnotes or directly.
The reader reads the paragraph and is impressed by your deep insight, which you never had, and the original author did.
How about if your new book has many such uncredited snips from other authors, such that your new work is denser and richer than any of the other individual authors?
This is what copilot is doing, or rather it's facilitating people doing it, as far as I can tell.
The original snippets are functional, not there for examination, copied verbatim, not transformed (sometimes), and not credited.
Most of it comes from open source works anyway and most authors would probably be fine with it if the stuff was simply credited.
I think as a tool, in the context of software vs literature, the tool is probably more good than bad for everyone as a whole. It probably results in the generation of more, and more correct software. Since software is more like a machine than a novel, it benefits all of humanity when machines work well.
But it needs to somehow credit the original authors, or if that's not possible then users do not get to claim credit for any work it was used on. Or, they can only claim a sort of tainted credit.
Maybe it needs a combimation of policies that together make a fair system. One element would be, the training set must be composed of strictly open source software (pick some definition). Then another element would be, any work that uses it, is tagged as such. You only get to say "I wrote this, with copilot." not merely "I wrote this". And any work that uses it is itself gpl. The individual snips maybe don't have to be credited because the theory will be the training set as a whole was credited, and those are all available somewhere. You as a contributor won't get credit for being in someone's mp3 transcoder app, but that app WILL declare that it used the training set, and the training set WILL declare all of your material that is in it.
Maybe there can be a special version that only includes code where the original terms did not require anything at all, not even preserving the authors name or the license that says it's free, and that version's output can be used without credit.
If proprietary software wants to benefit from a tool like that, they can pay for licenses from other proprietary software developers to include their software in their ai's training set, just like with normal software licensing for inclusion and re-sale in a new product.
But right now, as copilot currently exists, as far as I can tell it's blowing past and ignoring ANY considerations like that and Github are simply outlaws.
I'll just leave y'all with my favorite of the things you keep telling me to STFU about art AI with: If you're the kind of programmer who feels threatened by this, then you're not a real programmer.
Copilot and Dall-E (and so on) are all bad in the same way.
Many of us agree with you.
> I just got a Dall-E render with a very intact "gettyimages" watermark on it.
Of course, it doesn't mean that all or even most DALL-E output infringes on someone's copyright. The same is true for Copilot. I think both have many legitimate uses if and when the "copyright laundering" issue is solved.
Overall seems like a good trade. (for the world)
Reinventing the wheel, millions of time a day, is an atrocity.
Millions of (wo)man hours, wasted, every single day, on writing solutions to problems that have already been solved. There is a partial solution to this, and it's making people angry, it's crazy.
If you put your code publicly on the internet, you should expect that people will reuse your code at some point, no one broke into your privates repositories.
Why would anyone waste their time to make other people waste more of their time is really beyond me.
Let go of your egos for once.
It is…
I'm assuming you're referring to fair use. In that case whether it's copyright infringement or not is very situational (the legal standard consists of a test with various subjective factors) and isn't as simple as "it's less than a paragraph so I can copy whatever I want".
Note that this may not actually be true, and you may need to pay to license even shorter excerpts of creative work. Copyright is a complex topic. It's not always safe to assume that you have the rights you think you have, in terms of reproducing others' work.
For example: "The proportion of a total work is not the only factor, though. If you are including the most crucial aspect of a work, even if it is only a small part, then the question of “substantiality” comes into play." [1]
[1] https://www.dukeupress.edu/getmedia/3363cb6e-04b6-43ec-b004-...
Whether or not it is infringement depends on if the use can be considered fair use. This is a more nuanced question and is not always clear.
In this case (Copilot) the real question is how transformative the AI training is. Given how verbatim some of the outputs are makes the argument less clear.
I expect we'll see new licenses appear making it clear whether or not the content can be used for training.
1. the purpose and character of the use; 2. the nature of the copyrighted work; 3. the amount and substantiality of the portion used; 4. the effect of the use upon the potential market for the original work.
All of these are quite debatable, and I'll leave it to someone more familiar with the law.
Though if it's not, I believe there are licenses that allow derivative uses of code and licenses that don't. For many of these, the intention is that they create more code, but not be used to fuel AI behemoths.
That's the pitch. You're renting a pair programmer.
If your pair programmer is stealing code, you're going to have a bad time. This has nothing to do with...whatever you're on about.
Seriously though. Did you click the wrong reply link?
Also to your other comment about copyright not being an issue if you just use a paragraph from a book - I am not a lawyer but I would think that copyright applies just the same way it applies to musicians who use portions of the melody of other musicians’ songs.
My guess is that there wouldn't be much to scan ...
https://docs.github.com/en/site-policy/github-terms/github-t...
when you put code on github.com you grant GitHub the right to show that code to others, independent of the license you choose for your code. full stop. doesn't matter if it's on a webpage, a git client, or a github-developed plugin to an IDE.
by uploading code you attest that you have the rights necessary to grant that license to GitHub: https://docs.github.com/en/site-policy/github-terms/github-t...
without the right to grant those licenses to GitHub, by uploading that code to GitHub, you are in violation of the terms of service, and the responsibility of acting in compliance with the license is on the shoulders of the user which uploaded that code to github.com.
Said another way, GitHub has no way to know if the person mirroring SQLite (for example) is acting in accordance with their rights, so the terms of service require that you attest that you are acting within your rights, acknowledge that it is solely your responsibility if you are not, and that by uploading you grant license to GitHub and its users.
the right to allow forking is granted by a user who uploads their code to github.com to other users of github.com. those rights are listed here: https://docs.github.com/en/site-policy/github-terms/github-t...
Read the GPL.
When you upload code to github you give other people the right to fork it... I knew that already. But you license it. You don't give anyone the right to fork it and not abide the license. So if I fork it, I'm still giving Microsoft rights I don't have, I'm giving them the right to violate the license. That makes it illegal for me to fork it.
Let's say I am on a git mailing list, following a project, and I upload that project to github one day. It's licensed GPL. Microsoft says I give them the right to violate the license, and in uploading it I implicitly attest that I have the right to do so. I've violated the license? It's illegal for me to upload the code, with the license, to github, because Microsoft demands rights I don't have to give? Then let's say someone else forks it. They've now also violated the law?
It's nonsensical. The license is the binding ToS here, period, it doesn't matter what Microsoft's lawyers argue. Everything else is secondary.
you are talking multiple separate things here.
when I upload code to github.com I attest that I have the rights required to do so, and the rights required to grant GitHub the licenses I've agreed to grant it by uploading.
> You don't give anyone the right to fork it and not abide the license.
correct, you can't grant a right to violate the rights granted. users of the code hold the responsibility of acting in accordance with the license.
> So if I fork it, I'm still giving Microsoft rights I don't have, I'm giving them the right to violate the license. That makes it illegal for me to fork it.
no. you did not upload code that you forked from a GitHub.com repository. if you are talking about uploading code that you copied somewhere else, and you're calling that a fork, you have violated the terms by uploading code that you do not have rights to upload. remember, by uploading code to github.com you attest that you have the rights required to do so, according to the terms of service. if you lie, you are responsible for that lie and its consequences.
> Microsoft says I give them the right to violate the license
your premise in this part is flawed. see above.
> Microsoft demands rights I don't have [the right] to give?
by uploading to GitHub.com you attest that you have the ability to grant those rights. If you lied, and you don't have those rights, that's your responsibility and your ass if a law suit comes around because of it.
perfectly sensible to me. GitHub gets to say that they require users to grant the rights in order to upload, and that the users necessarily had the rights to give to GitHub. if a user lied, that is not GitHub's fault; the user entered into a legal agreement saying they had the rights needed.
Here's a fun way to see it, suppose someone writes code licensed GPL. I take it, fork it, modify a line in it or not, and also license it GPL because I have to by law. I put it on my github account and what, I now just gave Microsoft rights to the code I don't even have? So by putting it on github I'm violating a license? It doesn't add up. The license to the code is the license to the code, no matter what site it's on and noatter what any ToS says. Otherwise what's to stop me from putting a ToS on my personal website partaining to your use of my eyeballs that says "if your creation becomes viewable by my eyeballs in any way I can use it however I want, publishing your work in such a way that it can be viewed by my eyeballs is consent to this ToS"?
yes. if you don't have the rights to upload code to github.com, including all of the rights required of one that uploads that code to github.com, and you do so anyway, then you are in violation of the GitHub terms of service.
fortunately for you, the GPL allows what you are describing: "1. You may copy and distribute verbatim copies of the Program's source code as you receive it, in any medium, provided that you conspicuously and appropriately publish on each copy an appropriate copyright notice and disclaimer of warranty;..."
> Millions of (wo)man hours, wasted, every single day, on writing solutions to problems that have already been solved. There is a partial solution to this, and it's making people angry, it's crazy.
Following this line of thought, do you think that all code from all software should be open source and publicly available (and free to copy and use), in the interest of saving more person hours from reinventing the wheel?
Why are you asking this as if the answer might be no?
Let's help shape this thought: Copyright should be abolished entirely. It is one of many monetization schemes and its negative effects greatly outweigh its positives.
We know people won't stop writing software in the absence of copyright. We know they won't stop writing books, singing songs, etc. Copyright is not the primary motivator for either science or art.
Will we need new monetization structures? Of course. But generally speaking we already have them where it matters.
End copyright entirely.
Having a way to own works is probably pretty important to either of those endeavors, right?
There is no such as "owning" a work. We use that as a euphemism for owning copyrights, and the only function of copyrights are to prevent others from making copies. To prevent others from sharing.
The question is whether the monetization model presented by copyright is a net positive for the author, after accounting for its chilling effect on communications for all other people in the world.
The answer is almost certainly "no," as empirically demonstrated by entire segments of IP work opting out of copyright. The open source model clearly demonstrates that you do not need to own a work to fund it or monetize it. There are similar models in other areas of art and science which allow for the funding of works without preventing others from copying them.
Even if you're right in principle (and I would love new monetization structures), this will never happen in reality.
Meanwhile, this idealism will get applied asymmetrically in the real world. If you (or the comment I was replying to) say "Copilot is fine, all code should be publicly available anyway", it downplays the fact that this wish will never happen with big players like Microsoft and will only happen with little players like anyone who used Github to host their code. The big player will typically hide their code behind copyright and lawyers to enforce it, whereas the little players have no similar recourse.
So, I see the issue as an exploitation, as Microsoft is selling a product built on the little players and not the big players. The debate around whether copyright should exist at all, while interesting, is not that relevant to most of the concerns being aired in the context of Copilot.
> this will never happen in reality.
Don't be so sure. These kinds of changes start with education.
Training must be opt in, not opt out.
Every artist, every creative individual, must EXPLICITLY OPT IN to having their hard work regurgitated anonymously by Copilot or Dall-E or whatever.
If you want to donate your code or your painting or your music so it can easily be "written" or "painted", in whole or in part, by everyone else, without attribution, then go ahead and opt in.
But if they don't EXPLICITLY OPT IN, you can't use the artist's or author's creative work for training.
All these code/art washing systems, that absorb and mix and regurgitate the hard work of creative people must be strictly opt in.
I understand the argument from an artist's perspective much more, since they don't really have the option to publish their work in a way that any AI or any other artist can't copy off of.
One example of restrictive but public licenses include requiring others to share their source code if it's derived from yours, allowing individuals to use a product but not allowing business to use it (businesses can use it under a different - likely paid for license), or requiring attribution or acknowledgement that they used your code.
There is an argument for fair use if it counts as a substantial derivative, which is a different discussion from why people make it publicly viewable without making it flat out public domain.
The vast majority of open source licenses and copyright terms specifically stipulate the legal requirements for reproducing even just parts of the code. Which at a minimum require reproducing the license and copyright with all software including the licensed and copyrighted code.
Should students need to attribute the copyrighted textbooks and lessons that they learned from for all their future work?
Should artists attribute every reference they've used? Even if they draw stick figures based on the reference? Even if they only use small parts from multiple references?
What's different from a machine learning something and a human learning it?
I think in terms of practical open source/permissive licenses it makes the most sense for new licenses to be made that include no-training clauses for the rights holders that dislike machine learning.
Dall-E's use of training on non-permissive copyrighted web-scraped data seems more complicated and I imagine there will eventually be lawsuits to figure that out.
Stack Overflow facilitates the same thing too, so it's an interesting comparison, but SO makes attribution easy and clear, and it actually made it effortless to contribute back.
I understand that there is a balance, but as an open-source advocate who would love better tools to make their open-source projects better I'm lost as to why this point doesn't counter the "giving nothing back" we hear so often.
As long as some company can improve its bottom line it’s all good though
Eventually someone comes in and takes everything that isn't nailed down and then sells it, and that becomes the problem.
But they have to abide by my fucking license.
I don't use Github, but fuckers upload my code there anyway.
Copyright is evil, but only large corporations having copyright, even more than they already do, is even worse.
The asymmetry that exists in copyright law where large corporations can enforce their copyright to the point of breaking the law themselves (YouTube's content ID is another non-legal, but still very impactful example) is absolute bullshit.
Unfortunately I think that if training ML models on Internet-data is found not to be fair use then things will get harder for individuals training models and corporations will be barely inconvenienced as they can afford to pay for sources, make deals with other large institutions for data, etc.
OR
This place is currently crawling with Micro$oft employees who have been instructed to swamp the place with disingenuous comments basically amounting to:
1) "fair use" is anything I want it to me
2) gimme your code NOW, because I want it, and it's MINE
3) get used to habitual violation of licenses as the new normal
4) you are ruining progress! harming kittens!
I can't see the actual HN crowd all suddenly being copilot users and fans, so that leaves me to conclude the latter.
I find Microsofts continual business model of evil to be rather threatening and annoying and they need to be checked, as they have only gotten worse with the decades. They abuse their market position to stifle any and all tech innovation. Break them up already.
A good start would be to take a leaked code of Windows, and then mechanically adjust all the names, constant values, and code formatting, and then publish it and observe.
> "You are responsible for ensuring the security and quality of your code. We recommend you take the same precautions when using code generated by GitHub Copilot that you would when using any code you didn’t write yourself. These precautions include rigorous testing, intellectual property scanning, and tracking for security vulnerabilities."
I can't help but recall:
"Linux is a cancer that attaches itself in an intellectual property sense to everything it touches."
- Steve Ballmer, while CEO of Microsoft
They have some really good blow in Redmond.
If anybody could win an award for being coked up and sweaty on stage...
See also this Domo video that turned it into a song. :) https://www.youtube.com/watch?v=f7ZDH45OAt8
On the one hand stating plainly that mixing in copy-left code and similar can be disastrously dangerous because it is a rampant virus. On the other hand not understanding why people think it might be a problem that their tool could encourage mixing in copy-left code.
According to the opinions about what inclusion of open source code into your projects does, as per the ex-CEO of the company. That seems a bit of a far fetched conclusion, but then, Ballmer did say it.
With "normal" code I can generally see (or figure out) who posted/published it and reach out for explicit permission. It's not uncommon for me to do this.
How is one supposed to do that for the generated stuff? Seems like an awefully hands-off attitude. As challenging as it is, they really ought to be qualifying the input samples of training code before ingesting.
There are some vendors in this space too (BlackDuck comes to mind) but they're $$$ so only within the scope of large corporations.
If anybody has any ideas relating to this type of analysis, I'd be excited to chat. I am working on a project[1] in this space for "Software Composition Analysis" which could potentially overlap with snippet detection for code like Co-Pilot. (We basically just have a big pipeline of analysis jobs that run on code and store the results. I need to update the docs!)
0: https://yangdanny97.github.io/blog/2019/05/03/MOSS
1: https://github.com/lunasec-io/lunasec/tree/master/lunatrace
Seems like best practice recommendation that everyone should apply when downloading a torrent.
Let's pretend for a moment that your value judgement is reasonable and the advancement of technology should reign supreme over minor things like rule of law. Do you really think that letting people ignore copyright is always good for technical progress? Say, letting people use GPL code in proprietary code that they then refuse to share with others? Because that sounds questionable even if we agree with your casual disregard for the law.
> If you don't believe that, I am not sure what you are doing in a community like this.
Being interested in tech without being a fan of breaking the law and running roughshod over other people's work.
Only in very, very, very, very specific circumstances would I say it is not good. And they involve thinking about the counterfactual: "would this thing be created if there wasn't intellectual rights in place"? Code doesn't pass this test because people enjoy writing and sharing code. Pharmaceuticals, maybe.
I don't want anybody being able to generate "intelligent agents" and will support all legal changes likely to slow this down or halt it.
> Last update of whois database: 2022-10-17T23:07:12Z <<<
Just sayin'...
Using it is exactly like using Google. Google scrapes the internet and trains a model that gives you results for search queries on their website. The results may be copyright protected
Copilot scraped the internet to train a model that gives you results for code snippets in your code editor. The results may be copyright protected
The potential legal issues are there, but that's not why Copilot should die.
Copilot should die for any (or a combination of all) these reasons (and more which I don't mention):
- the operator has to already understand the emitted code to be able to determine if it is what is needed, or to modify it if it is close but not quite right
- the operator may have a false sense of capability, leading to bugs and other problems that would appear later (in production?)
- wrong suggestions are a distraction from the careful mental structures which one maintains while writing software
- any problem that Copilot can solve with guaranteed correctness is probably trivial or already met by a (battle tested) library
Forgive the analogy, but effective automated code generation is like autonomous driving systems. Anything less than 100% accuracy is a risk, and in these examples risk of incorrect behavior is not acceptable.
Copilot seems like a pointy-haired boss fantasy where they can hire only junior programmers and expect successful software products.