His opinion, as well as the current state of law: His code could and shall be used pretty freely, but a comment referring to the original license/author is needed.
The context: Having an option to "leave out copyrighted code", that does not do what the option suggests, is a flaw, both in a functional and secondly with law-as-is, in a legal way.
Addressing this exact premise is the purpose of copyleft licenses like the GPL. I would expect GitHub of all companies to understand this, and it's a spit in the face to disregard it, a spit in the face of the entire FOSS community to which they owe their success.
Disclosure: founder of a GitHub competitor
I’m totally against software patents and totally in favor of software copyright, fwiw.
Open Source software can exist in a world without copyright; that's basically what permissive licenses are. Proprietary software can exist in a world without copyright; just don't release the source code, only distribute build artifacts. But free software is dead without copyright.
Most permissive licenses — and most popular permissive licenses — include a copyright statement and conditions which would not be enforceable without copyright. They usually require redistribution of the original copyright notice and license text to be distributed along with copies or derivatives of the work.
That doesn't mean there can't be other legal mechanisms that would also accomplish software freedom.
Closed-source software should just be illegal. Then you don't need copyright to ensure freedom. With copyright gone, you are now guaranteed free to reuse any code. The law should not prohibit copying, it should only require attribution (i.e., prohibit plagiarism).
For a prominent example, WINE goes to extreme ends to ensure their code isn't derived from any prior direct knowledge of Microsoft code. Replicating functions of released binaries with original code is fair game, but reverse engineering or copying Microsoft code is a hard no-go.
As far as AI is concerned in all this, it's probably even easier to distinguish the unacceptable. Whereas humans can argue for plausible deniability, we clearly know what data an AI is fed to generate subsequent output. So unless an AI is certified to never have eaten licensed materials, literally everything it produces will be license infringements.
The training data doesn’t exist at runtime, and if the developers are competent then it only learned a bit or two from each input snippet. So, though even though it’s certainly a highly capable compression algorithm, verbatim copies shouldn’t be… possible.
My best guess for this bug is that a lot of already plagiarised the code. AI has a habit of holding up a mirror to humanity. Sometimes we don’t like what we see.
I don't think current legal frameworks are equipped to handle these types of generative works. There will need to be either new laws, or at least some legal precedence established. For example, it would make total sense for stuff posted publicly to optionally have a "not to be used to train models" license. But models are being trained on public data that pre-dates anyone thinking that is something that they should worry about.
The problem is that code is text-based and obviously some kind of writing while mostly being a utilitarian invention. The courts of course recognize this distinction (because unlike the average software developer, lawyers and judges have studied the law) so they have come up with a number of ways to filter out what is and isn't covered by copyright with regards to software.
The deal with patents is that you make the way your invention works public knowledge. You document the principle and in exchange you get a limited-time monopoly on its implementation. An alternative for this is a trade secret. As long as you guard the secret, you can remain the only one profiting off it. Of course, some things are impractical to keep secret. It's easier to keep a method of making a fizzy drink under wraps than how the gears are laid out in the gadget you sell.
You also can't really make the contents of a book a secret. Anyone can just look at their copy, arrange words in the same order as you did and sell your book. It also makes no sense to patent the words. The contents are already inherently public knowledge to a certain extent and now people could just buy the book from a patent office. And the clerks don't like filing novel-length applications. Music is not quite as easy to make a 1-to-1 copy of, especially before wide adoption of sound recording technology, but a lot of people can carry a tune well enough to plagiarize one. And sometimes you get a child prodigy Mozart illegally transcribing Miserere.
Since it has been deemed desirable [if not unanimously] that writers and artists also have a period of monopoly rights over their work, copyright needs to work differently from patents.
Software has interesting properties. Unlike a book, you can distribute a program but keep its "recipe" a secret. We'll disregard people who are capable of working with compiled binaries; consider them modern-day transcribers of Miserere. Proprietary code is typically guarded just as any trade secret would be. Perhaps software should have been covered by patents instead of copyright from the beginning. After the patent expires, the algorithm would be unencumbered and public knowledge.
But that's not the world we live in. Software is considered to be like a book, not like a pocket watch. You could mail me a printout of the entire Windows 11 source and there's hardly anything I could ever legally do with it. I could change it for my own amusement I guess, like I can take a red pen and change all the names in my copy of Postmodernism for beginners.
Most software licenses are not full carte-blanche do-whatever copyright waivers. Sure, it would be hypocritical to get up in arms about copilot emitting CC0/WTFPL/Unlicense code, but comparatively little software is released under such licenses.
The GPL (in all its versions) is almost as notable for the specific conditions placed on the freedoms it offers as it is for those freedoms themselves. One possible reason to release your code under GPL is that it you don't want your work used in proprietary developer tools. It doesn't really matter then if the code is used in a proprietary developer tool's training data set instead of its own program code.
Even permissive licenses typically come with some conditions attached. If I release a program under MIT license and someone then uses portions of that code in their own project, I expect to see my name and the MIT license included in some way. Perhaps a line in the README, an ACKNOWLEDGEMENTS.txt or a comment like
// This function taken from quuxifier (https://example.org/software/quux)
// ⓒ bitofhope 2022, used under the MIT license. See doc/licenses/MIT or
// https://spdx.org/licenses/MIT.html for details.
The same, in my opinion, applies if recognizable (for some definition and threshold of recognizable, which can and does get deep into lawyer territory) portions of that code are emitted by a convolution network. If Copilot can write my function, why can't it write my name and choice of license too?Even if you are a full copyright abolitionist, you can still point out the hypocrisy from the opposite side. Github's parent company Microsoft is notoriously protective of its proprietary code and its copyright. The company has historically been explicitly hostile to the free software community and despite its later unilateral declaration of love for open source continues to profit from their proprietary code, including Copilot. If you spend a decade or a few calling open source a cancer and campaigning against it, you can expect that to come bite you in the ass when you later try to sell a neural network trained on a vast corpus of free and open source software.
Or, well, no, "weird" is not the correct name.
I bet the same people would be fine if this code was included verbatim with appropriate attribution and licenses. But that’s not the case.