If ChatGPT produces AI-generated code for your app, who does it belong to?
zdnet.com
zdnet.com
> 1. Ownership belongs to the person who arranged for the work to be created.
> 2. Ownership and copyright are only applicable to works produced by humans, and thus, the resultant code would not be eligible for copyright protection.
> 3. A new "authorless" set of rights should be created for AI-generated works.
It seems obvious to me that the answer should be #1. An artist who creates works out of random paint splatters (modern art!) didn't purposefully choose the locations of their paint marks. However, they still own the copyright because they arranged for the creation of the work. You would never argue that a piece like this is uncopyrightable.
I'd like to take a moment to emphasize that whether something is copyrightable and whether it is infringing are actually separate considerations. Both are social constructs the same as which side of the road one drives on, and it makes sense for us to define them in some way that is "fair", whatever that means.
Prompts are really just a UI (with a bad UX IMO, it just happens to be impressive at a first glance to "talk with the computer"), trying to copyright a prompt is akin to trying to copyright a series of clicks and keypresses in an image editor to produce some effect (and just like a "prompt", you'd easily get different results depending on where you apply that effect, what program/version you are using, etc).
But, isn't that copyrightable? Like, if I draw a pixel art sprite, that's just a series of clicks.
You might argue "well you didn't make the AI image", but if not for me doing something it wouldn't exist, so we're back to discussing process.
You doing "something" doesn't mean that "something" is copyrightable - or that it should be (if anything, less things should be).
You said that for pixel art "the copyrightable part is the pixel art sprite," but then couldn't I say the copyrightable part of an AI-generated image is the generated image? Again, both of these creations were made via UI interactions!
If UI interactions aren't copyrightable, then practically no digital art is copyrightable, because all of it can be recreated by a series of UI interactions. I'm not understanding how a digital drawing and an AI drawing are different under your framework.
No, because AI-generated images are machine output which isn't considered copyrightable. This isn't unique to AI-generated images, or even AI-generated output, but anything machine generated. That you did something to cause the machine to generate that output doesn't really mean anything since after all machines do not generate outputs by themselves, they need some human input - even if that input is to press a "Start doing things" button.
(if anything this isn't even unique to machine generated stuff, not all things are copyrightable - for example in many countries bitmap fonts are not copyrightable)
> If UI interactions aren't copyrightable, then practically no digital art is copyrightable, because all of it can be recreated by a series of UI interactions. I'm not understanding how a digital drawing and an AI drawing are different under your framework.
It isn't my framework, it is how things happen right now (if it was up to me, copyrights would be even more lax than they are right now - e.g. i see things like youtubers having to fear DMCA claims and twitch.tv streams getting muted as something that should never happen).
Digital art is copyrightable because it is the direct result of human work, AI-generated art is not because a machine made it - even if it was on a human's "request". You can use the AI-generated art as part of human work (f.e. draw some mustache on a generated face :-P) and that would be copyrightable.
Think of AI generated content (not just art, what i write applies to anything generated) like the result of `for (int i=0; i < pixel_count; i++) pixels[i] = random();` - claiming that the result of that is copyrightable makes no sense - and the laws, at least so far, seem to agree with this.
In addition if you compile a C program that is just the line above (plus a bit of scaffolding to interface with the OS) you could probably claim copyright on the program itself (if you wrote it) but you couldn't claim copyright on the actions of opening a terminal, running the C compiler and then running the program.
If you add some extra parameters to the program to control the randomness spread per RGB channel and X,Y coordinates, like -say- `the_program -rrnd 32..64 -grnd 100..200 -brnd 0..64 -yrrnd 0..100 -ygrnd 100..200` or something like that you wouldn't claim copyright to the parameters `-rrnd 32..64 -grnd 100..200 -brnd 0..64 -yrrnd 0..100 -ygrnd 100..200` pretty much like you wouldn't claim copyright to parameters like `-ffast-math -O3 -march=native` etc to GCC or Clang -- and yet those are the equivalent to a "prompt".
The exception is actually more narrow than it sounds too. Work for hire requires that the work be made as part of employment or that the work was both commissioned for a handful of carveouts and there was a contract in place agreeing that it is considered a work for hire.
So if you commission a piece of art for yourself, the owner of the copyright is the artist. Also my understanding is that while the artist can transfer copyright, they can at any point take it back from you.
I would very much hope that wouldn't hold up in court.
A case before the High Court (roughly analogous to the US Supreme Court) determined that images produced in a video game were the property of the game developer, not the player -- even though the player manipulated the game to produce a unique arrangement of game assets on the screen.
See, I do agree with the court if you're playing through a linear game like Portal or Uncharted. Yes, technically the exact frames are unique to your playthrough, but the creativity came from the game developer.
Minecraft is something very different.
If I take someone's book, painstakingly replace all words with their synonyms and publish it as mine, I am getting sued. Same if I do some other mechanical transformation meant to obscure the true original work.
There's no AI, there's just statistical models of existing work.
Your argument forgets why copyright even exists, it's some a law of nature, it's a rule created to protect people who invest time and effort into creating. AI even if it existed would not be a person, it would be code created by a person (or more likely a corporation) to serve its goals.
What matters is the result, not the process.
It’s hard to swallow with AI, but there is no other solution in my view. If we start using the process as a criteria for copyrighting, then we have chaos. Indeed, you cannot prove which process was used after the fact, especially not with AI.
And the reverse is particularly problematic: anyone could then allege some human work is in fact AI and hence cannot be copyrighted. We’re breaking the copyright system if we focus on the process.
Another example: if you ask an AI to create something “in the style of xxx”, then the result may be something that is infringing copyright, but so would be humans work producing tbd same output.
In the end, what matters is the output, not the fact that it’s math or a human. We’re also just a bunch of atoms in the end, one could argue, very similar to a very complex mathematical model…
Citation needed.
> anyone could then allege some human work is in fact AI and hence cannot be copyrighted
I didn't say that. I said the copyright should belong to the original author. If some new work is provably based on old work to a substantial degree, it's plagiarism. Same as now. Indeed, the tool used does not matter in this case.
And we know today's LLMs produce code based on copyrighted original work because they readily admit they scraped everything available to them, regardless if it was proprietary, copyleft or public domain. Older versions literally produced entire functions copy-pasted from GPL-licensed Quake code. They "patched" that now to be less blatant (mix more) but it does not make it right or legal.
---
Now imagine a rogue employee of Google trains an LLM on all of Google's proprietary code (and only Google's code) and releases it publicly under, say, Apache 2.0. Can you imagine Google saying it's OK and not suing?
What if that person also got his friends from Microsoft, Oracle, Apple and Amazon to pool their companies' source codes? Would that be enough mixing? Clearly the LLM would only be regurgitating code that belongs to one of these companies.
What if instead somebody scraped only AGPL code and released the model under Apache 2.0? Clearly whatever comes out of the model is in its entirety based on work that is licensed under AGPL and therefore also has to be licensed under AGPL.
How do you prove whether or not someone has used AI?
> I didn't say that. I said the copyright should belong to the original author. If some new work is provably based on old work to a substantial degree, it's plagiarism. Same as now. Indeed, the tool used does not matter in this case.
I am a human. I have been shaped by the world around me. Everything I have seen and read has shaped who I am today, and it has shaped the type of work which I produce myself.
Yes, AI is trained on copyrighted work—but so am I! Nothing I produce is ever truly original. But it's different enough that it would not, and should not, be considered copyright infringement.
Now, if I accidentally reproduced an entire function from GPL-licensed Quake code, that would be copyright infringement. And humans do make mistakes like that, when they've seen the same code many times!
(Just to be clear, I am not one of those people who thinks LLMs are recreations of the human brain, I think they work differently. But we do both create output based on copyrighted input.)
---
> Now imagine a rogue employee of Google trains an LLM on all of Google's proprietary code (and only Google's code) and releases it publicly under, say, Apache 2.0. Can you imagine Google saying it's OK and not suing?
This is, in fact, one reason companies are always a bit worried when employees go to work for competitors. But luckily, we don't let Google say "you aren't allowed to work for anyone else after you've seen our source code."
In this case it's rather easy, they admit it officially and publicly. In other cases it might be harder, we might need a whistleblower. I am sure in some cases where it happened, it'll be impossible to prove.
Difficulty of proving it is not a valid reason for making something harmful legal.
> But it's different enough that it would not, and should not, be considered copyright infringement.
Yes, the line has to be drawn somewhere. The (typical) human brain for example has limitations how much code it can memorize and reproduce verbatim. LLMs don't have them (they're orders of magnitude higher on current hardware and models and they're essentially unlimited in principle).
> But luckily, we don't let Google say "you aren't allowed to work for anyone else after you've seen our source code."
But we also don't expect the new employee to build a competitor for one of google's services or internal tools in record time that can be measured in LOC per minute. An appropriately trained LLM certainly could.
---
Ultimately we strayed from the main point. Copyright exists to protect authors and their intentions against parasites.
If somebody's intention is to offer code for free as long as people who build on top of it also release their work for free (a slight simplification and misinterpretation of the GPL), then copyright exists to make sure that happens. That for example somebody who focuses solely on advertising (non productive zero-sum work) can't take that code and profit from it leaving the original author to eat dirt.
There's also the question of income per unit of work vs "passive" income. Many builder professions get paid only as long as they're actively putting in work. We as programmers are privileged in that our product can often run without constant work and produce positive value automatically (but we only get the continuous value out of it if we own the product, not if we produced it for somebody else).
Now, I believe that giving somebody fixed amount of compensation for building something that produces continuous value is fundamentally unfair, exploitative and abusive. Many people will no doubt disagree, must of the world has probably not ever thought about it by the looks of it.
LLMs take this to a whole new level. They can bring in enormous amounts of money (both in subscriptions to their services and in value produced by their output). Yet the people who built the training data that made it all possible don't get _any_ compensation at all.
> OpenAI grants you ownership of output
That’s only meaningful if it’s legally established that the ownership of the output was theirs to grant in the first place. There’s no law directly on point and precious few legal decisions in this area, so it’s still very much an open question.The UK and EU are different. The EU allows copyrights on databases.
[1] https://constitution.congress.gov/browse/essay/artI-S8-C8-3-...
Leave the no to the naysayers.
Ship your app, generate traffic, usage, income. Leave the discussions to other people.
Long ago I went through the company-approved process to link to SQLite and they had such a long list of caveats and concerns that we just gave up. It gave me a new understanding of how much legal risk a company takes when they use a third-party library, even if it's popular and the license is not copyleft.
No, you should be compensated for your labor. This does not entail a market product.
And we can't ignore that ebooks exist.
That doesn't mean you own any software or ebooks obtained by ill-gotten means. Nor does it mean you can replicate and sell an application, as though it was your own, that you paid a licenses or rental fee for. But you still own your hard-drive and are free to erase it and sell it.
1. There may be people who cannot use/afford some software, although there is technically an infinite supply.
2. Collaboration becomes awkward. Either all contributors give up their rights (Open Source), or one contributor holds all the rights and the rest is being treated unfairly. The latter decreases the incentives to make software modular and reusable.
3. The resulting software typically gets worse due to some copyright enforcement mechanisms. For example, no closed source software will ever have a good debugger, because that would allow viewing and changing the source code.
4. It creates a power imbalance between software owners and software users. Nearly all software has to be adapted over time, but the software owner has a monopoly on performing such adaptations. The result is enshittification, surveillance, and basically a return to feudalism where daily life is governed by a small number of overlords.
5. It is not clear how to price software fairly, and there is also little incentive to do so.
6. My impression is that high-quality software converges to formal proof, which is AFAIK not copyrightable.
For all these reasons, I think it is time to consider a world without copyright on software.
To those that worry about salaries in such a world: Negotiate payment in advance (contracts, crowdfunding, bounties, ...), or get a job where software is created as a byproduct (consultant, researcher, tester, ...).
If I were to guess, I'd say the output of an LLM isn't copyrightable (it's not the creation of a human), unless it's a verbatim copy of some copyrighted training data in which case it belongs to the authors of the work(s) used in training. This creates the most annoying combination of legal problems around using it, so by Murphy's Law it must be correct!
The question we need a legal determination on (ideally globally consistent) is more on training data side.
Or put differently this is a "Fruit of the poisonous tree" type legal issue. Wondering about the output as the focus is all back to front
There are several ongoing cases, probably most prominently the Thaler v. Perlmutter case. [2]
[1] https://www.copyright.gov/ai/ai_policy_guidance.pdf
[2] https://www.copyright.gov/ai/docs/us-brief-for-appellees.pdf
The article is about whether the human using ChatGPT can claim copyright.
LLMs are not individuals automatically covered by copyright laws, as they are simply tools based off other (often) copyrighted work. This means that the initial copyright infringement is still a valid concern, hence these discussions.
If it was a easy and clear cut as just shrugging, the conversation wouldn't be so prevalent.
> General-purpose AI models, in particular large generative AI models, capable of generating text, images, and other content, present unique innovation opportunities but also challenges to artists, authors, and other creators and the way their creative content is created, distributed, used and consumed. The development and training of such models require access to vast amounts of text, images, videos, and other data. Text and data mining techniques may be used extensively in this context for the retrieval and analysis of such content, which may be protected by copyright and related rights.
> Any use of copyright protected content requires the authorisation of the rightsholder concerned unless relevant copyright exceptions and limitations apply.
> Directive (EU) 2019/790 introduced exceptions and limitations allowing reproductions and extractions of works or other subject matter, for the purpose of text and data mining, under certain conditions. Under these rules, rightsholders may choose to reserve their rights over their works or other subject matter to prevent text and data mining, unless this is done for the purposes of scientific research. Where the rights to opt out has been expressly reserved in an appropriate manner, providers of general-purpose AI models need to obtain an authorisation from rightsholders if they want to carry out text and data mining over such works.
(note that the "appropriate manner" is meant to be some machine readable way, AFAIK the way this will happen is still in works - the Act wont become law until 2026 anyway)
Under EU copyright law the machine generated output (like code, etc) cannot be copyrighted. Essentially this means that:
1. ChatGPT et al. can can train on copyrighted code, text, etc unless the authors opt out via some (machine readable) way.
2. ChatGPT et al. can then reproduce a bunch of code from whatever it was trained on, that code by itself is not copyrightable (but it can be modified and become part of a copyrighted work - think of it as combining public domain code with some other project).
AFAIK the only muddy aspect is what happens when ChatGPT (or really any AI generative algorithm) reproduces already copyrighted works without the knowledge of the user. Again AFAIK this is something that is currently being worked on.
There was an AMA on Reddit recently[0] by someone who worked on the act and answered a bunch of questions. IMO it is a great AMA on the topic (at least if you ignore the trolls that ask "why do you want to destroy EU", etc).
Also (unrelated to the above AMA) i think both UK and US are likely going towards a similar direction.
[0] https://www.reddit.com/r/ArtificialInteligence/comments/1fqm...