A user once replied to one of my comment[0] about this with the following:
> It's not really an issue when you're a large software corporation; you already have mechanisms in place to check for license compliance in everything that ships, including F/OSS plagiarism checks [1].
IOW, from my understanding, they don't care. Big players do their own checks anyway, and small fish won't be creating problems because it's too convenient for them. Classic Microsoft (as I know from 90s).
The bigger thread can be seen in [2].
[0]: https://news.ycombinator.com/item?id=32534697
Copyright (unlike patents in general) allows for independent creation. If I sit down to write a quicksort routine, it is going to look extremely similar to a zillion other quicksort routines out there.
The other question (IANAL) is whether writing a quicksort routine is even a creative act at this point.
Or they could integrate a way to find the produced output back in the corpus if it's sufficiently close and provide a reference/attribution. Basically whatever tool a copyright lawyer would use to track down original work.
And that's just the engineering solution. The AI researcher solution would be to extend AI learning algorithms to attach attribution metadata to the learned data so that the output could already come annotated with information about the source.
But the latter is much harder to do, so maybe the engineering solution would suffice.
Edit: here's the related tweet: https://twitter.com/DocSparse/status/1581632706693079042
That assumes that the licenses of your code and the original code are compatible which often isn't the case.
In those cases it seems that humans are already copying code without also propagating licenses appropriately. LLMs are more likely to memorize things which occur a lot (and I'd bet rare things that are representative of some conceptual axis).
The main examples presented so far, Davis and Carmack, have the property of having been copied a lot. The generative model is only surfacing an existing pattern of ignoring attribution. Sort of like the code-gen version of generating bigotry if appropriately prompted.
I'll also note that this pattern of retrievable memorizing of copyrighted and sensitive material is present in GPT-3 too and not just for code. As the situation is equivalent, a lawsuit should address the concerns of non-programmers too.
An open equivalent might be wary of being accused of contributing to copyright violations since in that scenario, there is no way to force people to respect it.
Kinda the same problem as YouTube. Lots of people copy movies on the high seas, but if you are as big as yt you cannot easily get away with it.
This is not like youtube because Github is already hosting those violations and people are already inappropriately copying or including such code. It matters not whether the local inclusion was fetched by copilot or a human fetched it using more manual steps through search.
But in defense of copilot, code regurgitation is uncommon in routine use. An editor extension allowing search of github would be at least as easy to use to violate licenses but I do not think it'd be taken down since that would not be its core offering.
Copilot goes far beyond mere search and provides a useful service. GPT-3 can also be prompted into generating copyrighted works of writers but I do not see people talking as if that is its primary utility nor as much clamoring in these forums to end that service.
Well, try doing that for music or movies or proprietary leaked codebase.
If you think copilot is uniquely producing things that are not that different from humans then surely no one would have any problem with feeding it massive amounts of corporate programs?
I am not aware what writers are doing but there have been plenty of uproar regarding stable diffusion. I have a feeling that if any tools like this get built for musicians/film-makers, it will look vastly different from the current situation.
> surely no one would have any problem with feeding it massive amounts of corporate programs?
There is a similar gymnastics done by human engineers today due to the issue of patents. I don't think this is a good trend to uphold.
> I am not aware what writers are doing but there have been plenty of uproar regarding stable diffusion
Yes but mostly in the art community. On HN there were plenty of arguments just the other day how it is not the same for art and programmers have a stronger case. I disagree but regardless, the case is exactly equivalent for GPT-3 and writers but it wasn't an issue generating about a thousand comments on respecting IP and ceasing deployment of LLMs until copilot.
Are you arguing that copyright laws should be abolished? I have no problem with that as long as it's clearly defined, you can't not respect copyright of open source code but enforce it for proprietary code, just because the value is arguably non-monetary.
This is quite important, actually, and I don't think enough people realize this. I am a photographer sometimes and it would be really cool if I could share my photos online under a copyright license that forbids their use in training AI.