1:1 copies are a different question (a "it depends" on how much gets copied and whether the code is protected, as in not obvious, in the first place), but those are not the main thing that gets focused on in that strange twittersphere (which the article is referencing).
I don't believe that's what they're arguing. They're arguing that they're putting themselves (or their employer's) products at risk by having unattributed copyright protected code in their products. It doesn't matter how much any of us would hope that copyright didn't exist, the fact is, it does.
There are other issues with Copilot, but by far the biggest is the licensing risks for those using it in a professional software engineering environment.
I'm a FOSS advocate and all my FOSS projects are MIT licensed. Anyone can use my code to do anything they like, but my day-to-day professional software engineering life needs to know where the code comes from and what its license is.
> it's official, obeying copyright is only for the plebs and proles, rich people and big companies can do whatever they want
> Wow. It's amazing that nobody has brought up the point that the result of ML is a derivative work from the input dataset. If the dataset has a restrictive license, like the GPL, it is a clear cut case of violating the terms of the GPL because it is not being released in the open
There are other reactions of course, I'm handpicking here, as those are the reactions I was referencing above.
> my day-to-day professional software engineering life needs to know where the code comes from and what its license is.
I understand, I agree with it. But that doesn't mean that code produced by an AI that did learn from GPL code has to be licensed under GPL. Verbatim copies of sufficient high originality, so that copyright applies: Sure. And that's where there is some risk of Copilot going wrong. But that's not the primary output of Copilot nor its goal.
It took no more than 24 hours for the verbatim code to appear in the wild. It's irrelevant if it's the primary goal or not, if it can happen once then it can happen again; therefore (I believe) it's an existential risk to any organisation that uses it for their own products.
This particular algorithm isn't learning anything, it's a big pattern matching/information retrieval system. It wasn't fed tokens, AST, constructs, type information, it was fed pure source code hoping the oblique magic in between would make it develop understanding. Nothing of the sort happened and it is instead regurgitating input data verbatim, which is the obvious copyright issue at hand.
And btw, clowns take? That's no way to enter into a productive discussion.
> Nobody cares particularly if you feed data into an algorithm of your choice. I can feed data into a hash algorithm all day and that won't come back to bite me.
I showed examples to the contrary in my other comment.
It doesn't only do that, but it does do that [1]. And that is the issue. It doesn't matter if it only does this 1% of the time, it's enough to be high risk.
[1] https://twitter.com/mitsuhiko/status/1410886329924194309
It once wrote an API client for a rather idiosyncratic interface I just implemented.