Also feels kind of icky to train on open source projects and then charge for the output.
Also feels kind of icky to train on open source projects and then charge for the output.
To be specific, the FAQ states: "It has been trained on natural language text and source code from publicly available sources, including code in public repositories on GitHub."
Some have raised concerns that Copilot violates at least the spirit of many open source licenses, laundering otherwise unusable code by sprinkling magic AI dust... most likely leaving the Copilot user responsible for copyright infringement.
Good, you will almost not be liable for infringement.
It is possible to provide CoPilot with a sequence of inputs that produces some of the input, which was copyrighted. Let's say you want to help people violate copyright, so you as a third party distribute a script that provides that sequence of inputs. Who's violating the copyright there?
Alternatively -- it is apparently legal to produce a clean-room implementation that duplicates a copyright implementation. Supposing you were to use a tool like CoPilot, which has just been trained on that copyright implementation. Is your room still clean? You might even be able to get it to spit out identical functions!
Or, if you have a ML algorithm which has been trained on leaked closed source code, and it is sufficiently over-fitted as to just provide the source code given the filename or the original binary, who is violating copyright when this tool is used? If it is just the end user, then this seems like a really convenient way to launder leaked closed source code.
If I induce you to break a contract with someone else they can come after me for damages.
For example in this case, there are developers who have created GPL code. That code was licensed to some other developer. Github then encouraged people to upload git copies of the GPL code onto github where it was put into the model. That model contains the copyrighted materials and isn't coming with the necessary notices. The output of the model can be code that is a direct stand in for the copyrighted work. Thus Github have become a party to breaking the license even though they themselves never agreed to the GPL.
In addition Github are encouraging (They are advertising it and making it available broadly) other developers to copy that code and use it in their project. Again that's encouraging an action that breaks a contract. Github is well aware that this is likely happening and they continue on. Thus they might be liable. You also might be liable.
All of these things can and likely will be argued before courts but it's not at all one sided.
> That's been established in court with regards to training AIs.
What are you basing the certainty of this statement on? The case law I have seen around this is pretty spotty. Cases around training on copyrighted materials have predominately been about the input, and not the output. With the final output usually being controlled by the model owner. For example Google obtained the books they scanned legally then used them to produce google books' index. There are some major differences.
- The books were purchased, meaning they got a license to use the book. There's for sure code in the model that Github does not legally have the right to use. They are aware of this. Making the input more shaky for github. - Github is making a direct profit off of this service. It's a revenue generating enterprise. That's important since it raises the bar of what they can be expected to do.
There's been nothing that goes to the supreme court yet; it's all per circuit and not settled case law. Also this gets WAAAAY more complex when we start talking about outside of the US and isn't decided at all.
These things are complex and likely you need your lawyer to advise you with any real questions.
This may be a bit nit-picky, but I don't think that is correct.
Most books I've seen don't say anything about granting a license so there would be no explicit license that comes with them.
Maybe you could find an implicit license if normal use of a book required a license but it does not. Copyright law allows all the normal uses of a book without requiring permission of the copyright owner. You only need a license when you want to do something that requires permission.
I was saying that there's some implied license after first purchase. I believe that was part of the court's decision. Paying for a book (or a library paying) gives you implicit rights to fair use. Github's copies of code were not purchased. They were given by sometimes third party.
So there's likely some room to argue that fair use rights are different enough between previous cases and github.
AI is just recomposition of existing snippets of code, art, text, music, etc. Does an AI fall under fair use? What happens when an AI produces something too similar to an existing work or trademark. I know the computer won't get sued, the owner/user will. But still, it's a hard problem.
Even if Copilot was initialized with snippets from Open Source Software (exclusively), it doesn't mean that copyright infringement isn't a concern.
It's not random recomposition, which is worthless. It's useful recomposition, adapted to the request and context. It adds something of its own to the mix.
Been a hell of a decade, hasn't it.
The wisdom of crowds works best when:
1. participants are independent (otherwise you may get failure modes, such as "groupthink" or "information cascades")
2. participants are informed, but in different ways, with different opinions;
3. there is a clear, accepted aggregation mechanism, where individual errors "cancel out" to some degree
I view the topics in James Surowiecki's book (or the Wikipedia summary of it, at least) as required thinkinpg for everyone, preferably synthesized with a study of statistics and political economy.
In particular, the Wikipedia article's section on "Five elements required to form a wise crowd" is a slightly different slicing of the required elements that I offer above.
* If you read that section, trust is listed. I, however, don't see trust as a necessary condition for a "wise crowd". Trust is often useful (or even necessary) when a collective decision is used for governance, decision-making, and policy.
When I'm in the flow, trying to solve some algorithmic problem, I always turn it off because the BS suggestions coming from its little "mind" actually slow me down and mess with my focus. Which all makes sense when you realize what it ultimately is - a philosopher, as opposed to a mathematician.
Most projects are 90% BS glue code and 10% actually interesting code. I don't mind only having help with the 90%.
It helps solve the boring simple shit so I can focus on the interesting bit.
Yea, that makes sense, I agree with that. If your use case is skewed more towards "BS glue code" as you say, you'll find more use out Copilot. Then $10/month can be fair, cheap even.
Yeah, this feels like the same nonsense that scientific journal publishers pull. If your product only has value because of what we made, it's completely unfair to not pay us for our work and then to turn around and charge us to use the output.
https://www.infoworld.com/article/3627319/github-copilot-is-...
E.g., if I take that Disney movie, incorporate it into my own movie, and distribute it, then I'm also violating copyright.
And you might argue that Copilot is also a distributor.
"Intent is not relevant to copyright infringement liability."
"But your honor, I heard on Hacker News that it was."
"I find you guilty."
"But your honor, copyright violation is usually a civil issue, and 'guilty' is a criminal trial concept."
"Well, I also get my legal training from Hacker News."
Sure it has: Time.
In terms of economics it's really simple: Does Copilot free up more than 10$ worth of your time per month? If the product works at all as I understand it (I haven't tried), the answer should be a resounding "yes, and then some" for pretty much any SE, given the current market rates. If the answer is no (for example because it produces too many bad suggestions which break your flow), the product simply doesn't work.
There might be other reasons for you not to use it. Ego could be one. Your call.
> Also feels kind of icky to train on open source projects and then charge for the output.
I don't know why it would feel any more icky than making money off of open source in other ways.
And my problem is : Time.
Cycling through false positives and trying to figure out if it's right costs me way more than $10 a month in productivity.
I cant wait for better versions to come out, but right now, no.
$100/year is a steal for the amount of tedious code copilot helps me with on a daily basis.
In addition, the idea of "derived work" in code snippets is, quite frankly, nuts. There is only so many ways to write (let's be generous on the scope of copilot) 25 lines of code to do a very specific thing in a specific language. If you have 1000000 different coders do the job (which we do) you'll have a significant amount of overlap in the resulting code. Nobody is losing sleep because of potential license with this. Because that would be insane.
I have noticed that upholding oss licensing (at least morally) is kind of a table manner on hs. That's fine, but this is some new level of silly.
It's also not gonna persist, because no matter how much we love our oss white-knightedness, we love having well paying jobs more.
It's quite nice not to have to type generic boilerplate in sometimes I guess but it's very frustrating when it generates junk.
I think it really depends on what languages you use though. If you use something like Kotlin where there's really almost no boilerplate and the type system is usefully strong, the symbolic logic auto-completion is just far more reliable and helpful. If you're stuck in a language where there's no types, and there's lots of boilerplate to write, then I can see it may be more helpful.
For me, this entirely comes down to the philosophy of how a deep learning model should be described. On the one hand, the training and usage could be thought of as separate steps. Copyrighted material goes into training the model, and when used it creates text from a prompt. This is akin to a human hearing many examples of jazz, then composing their own song, where the new composition is independent of the previous works. On the other hand, the training and usage could be thought of as a single step that happens to have caching for performance. Copyrighted material and a prompt both exist as inputs, and the output derives from both. This is akin to a photocopier, with some distortion applied.
The key question is whether the output of Copilot are derivative works of the training data, which as far as I know is entirely up in the air and has no court precedent in either direction. I'd lean toward them being derivative works, because the model can output verbatim copies of the training data. (E.g. Outputting the exact code with identical comments to Quake's inverse sqrt function, prior to having that output be patched out.)
Getting back to the use of open source, if the output of Copilot derives from its training data in a legal sense, then any use of Copilot to produce non-open-source code is a violation of every open-source licensed work in its training data.
What I pity however is that there's no free tier for hobbyists as paying a 10 usd monthly subscription wont make sense when you only code occasionally. For professionals using it everyday, 10 usd / month is inconsequential.
I don't think that would have costed them much more to offer a free allowance to cover say an average coding session of 8 hours per month.
100% of what I do is open source. It's used by millions.
It's free for maintainers of "major" open source projects. I'm not sure what a "major" open source project is, but it's clearly not what I do. The only way to know if your open source project qualifies is to try to sign up. If it does, you're given a free option.
I am the primary author (but not current maintainer) of an open-source project which is reported to be used by over 100 million people, according to (flaky) statistics kept by the current maintainers. That's around 1% of the people in the world.
I don't trust the current maintainers to be honest with numbers (there are lots of ways to estimate numbers of users), but it's definitely in the millions, and it's a project you (and most random people you'll meet in tech, and many outside of tech) will have heard of.
I am currently working on earlier-stage projects, which have smaller communities, but 100% of them are open-source.
"open source is great, except when it's used in a way I don't like"
Also - there is logic in copilot that checks to make sure it is not suggesting exact duplicates of code from its training set, and if it does, it never sends them to the user.
I wouldn't put any hard rules on it, but it does seem very fair for programmers who have learned a lot from GPL code to contribute back to GPL projects. I have learned from and used a lot of open source software so whenever possible I try to make projects available to learn from or use.
Why is it different if we slap a "ml" lable on it
This is particularly the case when we see the emergence of new technologies that use it in different ways. Different people may have a wide variety of equally valid views about how it is incorporated into that system.
There's nothing inconsistent, confusing, or complex about those views.
I think there might be some in this thread who don't consider these derivatives, for whatever reason, but it seems to be that if rangeCheck() passes de minimis, then the output from Copilot almost certainly does, too. That a tool is doing the copying and mutating, as opposed to a human, seems immaterial to it all. (Now, I don't know that I agree with rangeCheck() not being de minimis … and yet.) Or they think that Copilot is "thinking", which, ha, no.
You also have to (slightly) change your flow to get the most out of it, which I know is a deal breaker for many.
I absolutely love it. It's not going to write good code for you, but for an autocompleter it is amazing.
There's a limit to what individuals are willing to pay for a subscription service irrespective of how many hours it saves you. Now if we're talking enterprise and bulk licensing then that's a separate issue.
You wouldn't have an issue with someone making money by using open source software (like a website that is hosted on a server running linux).
It's been really nice for autofilling console logs and boilerplate code...but $10? It's a novelty that is nice when it works, but that's a steep price point for what it is, and I don't see that changing any time soon.
The business model for most of the Internet is to bait people into using things for free and then monetize them without compensation in some roundabout way.
How would you feel if they just provided the software without the model, assuming you could train it yourself on open-source code in an instant?
I just don't like the idea of taking people's work (without asking or checking licenses) and then selling it back to them. It'd be like if Stack Overflow decided to start charging to see answers and not asking or giving a split to the person who gave the answer. I realize they aren't just copy/pasting so not a perfect parallel, but still.