There are some utility functions that I have written dozens of times, for different projects. Am I supposed to be varying my style each time so that each has its own unique copyright? I do not believe any court would interpret copyright to work in this way.
But note this prohibition is of "code that was not written by yourself" rather that written by yourself twice.
If you write code for one company/client then the IP belongs to that company as a work for hire.
If you write the same function again for someone else then you would be infringing the first company's copyright - if you truly believe in this theory that little pieces of code have copyright.
If you believe in this theory then you need to have a database of all the code you've ever written and be constantly checking every function you write for potential infringement of a previous employer's rights.
It is impossible for code to not be written by yourself under that view because it is just a tool. The fact that it’s mad-libs autocomplete instead of normal intellij auto-complete is totally irrelevant.
> The fact that it’s mad-libs autocomplete instead of normal intellij auto-complete is totally irrelevant
Completely relevent if the former had and gave no permission and the latter did.
There’s a reason companies will use clean-room strategies to avoid poisoning a code base.
So, I’d be okay with a compromise wherein the makers of Chat GPT aren’t liable for copyright for the act of training the model, but are liable for copyright for statements produced by the model. So, when the AI spits out exact copies of copyrighted works in response to a prompt, then copyright has been violated and the AI creator should be liable.
let new_code = f previous_work
where f could be identity. But it could be other transformations such as: rename_all_the_things, move_items_around, object_oriented_to_functional, and what not. Probably a combination. But here's the kicker: all permutations of these functions are all derivatives.
And we don't have to be naive here, thinking that I'm talking about using refactoring tools on an actual code base because I want to hide the fact that I want to steel some piece of code.
I'm talking about the fact that I've seen so much code and this has in fact taught me basically everything I know. I have trained myself on a stream of works. Just like an AI. The difference is in scale, not in nature.
A judge and jury are going to laugh in your face.
An exact copy is completely different than “produced new code based on everything it’s seen”. You can try your “it’s just an identity function derivative” if you want, but… judges and juries aren’t idiots. They know that a copy is a copy. You’re not going to pull the wool over their eyes
I am definitely not trying to conclude that copyright or licensing is bad.
I have read the scheduler implementations in a number of different operating systems. If I were to implement a scheduler in a operating system, would you say that there is a probability that my scheduler will be a derivative to some extent of those that I have read? How is this materially different from training a GPT on those same pieces of code, and then asking it to construct a scheduler?
I would argue that while there are differences, there are also similarities. That means that the dichotomy is not true.
I am saying sometimes GPT doesn’t do the thing you’re describing. Sometimes it produces exact copies of something that it was trained on.
When it does that (not a “superficially similar” one, but an exact copy) can OpenAI be held liable for copyright infringement?
Because it seems like they want to avoid liability on both ends. They don’t want to be liable for copyright infringement based on ingesting works, and they don’t want to be liable when their tools produces a *copy* of existing works.
I don’t think that’s an acceptable framework.
I’m arguing that copying is not allowed. And you’re arguing derivative works are not always allowed.
I originally thought you were arguing that derivative works are allowed, and that since copying was a derivative work, that copying was therefore also allowed.
Sorry for my misunderstanding of what you were originally saying.
Not if neither is derived from the other.
Agreed. No "of course" about it.
While it is technically possible for you to read some GPL'd code, memorize it, and then reproduce it later by accident. That's not how programmers work. What humans remember is not the copyrightable code, but the patentable algorithm. (And there are few algorithms that are simple enough to memorize on a cursory reading, novel enough to be patentable, but not so novel you'll remember it's source)
AI does not work in algorithms. It's a language model. It deals purely in the copyrightable code. (Both figuratively; LLMs are structurally incapable of the high level abstract reasoning required, and literally by way of the training data)
1) You are a thinking adult - or a thoughtful teenager :) - and therefore you understand copyright law well enough to take responsibility for copyright law, and in particular you are capable of having standing in legal matters. AI is not, but it is quite capable of creating infringing code (eg GNU stuff) that human reviewers wouldn't even know was infringing. So it is much better for any honest and competent organization to ban commercial LLMs entirely. (I am fine with in-house solutions with 100% validated training data...but those aren't very good yet, are they?)
2) I am a broken record on this, but the biggest problem with the "stochastic parrots" analogy is that transformer ANNs are dramatically dumber than parrots, or any other jawed vertebrate (I am not sure about lampreys). As applied to code generation: when I first tested ChatGPT-3.5, I was shocked to discover it was plagiarizing hundreds of lines of F#, verbatim, including from my own GitHub. Obviously that's outrageous in terms of OpenAI's ethics. But it is also amazing how dumb the AI is! Imagine a human programmer who is highly proficient in Python, and pretty good at Haskell, yet despite reading every public F# project in GitHub it can't solve intermediate F# problems without shameless copying.
It is a completely misleading comparison to say that humans reading source code is anything like transformers learning patterns in text. The most depressing thing about the current AI bubble is watching tech folks devalue human intelligence - especially since the primary motivation is excusing the failures of a computer which is less intelligent than a single honeybee.