Yes. Weights probably aren't copyrightable in the US. See Feist vs. Rural Telephone, in which the Supreme Court ruled that telephone directories are not copyrightable. The copyright clause in the Constitution ("To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries.") is understood to require human authorship. The US does not have database copyright, or "sweat of the brow" copyright. That it was expensive to produce some collection of data does not make it copyrightable.
Outputs from LLMs, machine generated art, and machine generated music probably are not copyrightable either. US Copyright Office: "Based on the Office's understanding of the generative AI technologies currently available, users do not exercise ultimate creative control over how such systems interpret prompts and generate material. Instead, these prompts function more like instructions to a commissioned artist."[1]
[1] https://www.reuters.com/world/us/us-copyright-office-says-so...
A collection of numbers is copyrightable if it's the encoded result of a creative process. Just because it's represented as a bunch of numbers does not make it non copyrightable. That's why it says " original works of authorship fixed in any tangible medium of expression, now known or later developed, from which they can be perceived, reproduced, or otherwise communicated, either directly or with the aid of a machine or device. "
You can't just classify the weights as facts simply because they are numbers. If they are creatively made by a human they would be copyrightable. Mechanically computed from random numbers, no. Somewhere in the middle? Harder
>Just because it's represented as a bunch of numbers does not make it non copyrightable.
Can you give an example of where the bunch of numbers is copyrightable when it's not just a numeric encoding of something that was already copyrightable? Taking music and encoding it as a wav file is not a creative work, but it's a representation of a copyrighted work.
Maybe you could create a long list of numbers and call it an artistic impression, but that's clearly not what AI weights are. I'm interested to hear an example of your copyrightable numbers.
Here, there's a whole lot of creative decisions in labelling and guiding of training that produces the weights, so it's reasonable to think it might be copyrightable.
That is, the numbers are a whole lot more original than the issuance of phone numbers or part numbers.
(Agree that the skill in knowing how to code and guide the training of a model is probably very different though. It’s not just access to compute time that separates me from OpenAI :) )
Often, labelling is part of large public datasets that are chosen for use for that exact reason, and/or is otherwise not the work of the party claiming copyright in the model.
Labeling training data may qualify for copyright, but if the underlying training data doesn’t taint the output as a derivative work then labeling isn’t going to qualify by itself.
Thus without some new and very generous interpretation AI companies are at best not going to benefit from copyright and at worst may be forced to create all training data in house. My suspicion is this generation of AI companies are in a very difficult situation.
It depends. If each individual training item has a small impact on the output coefficients, then perhaps it's not a derivative work of them. But if there's a large creative process in determining model training procedure, deciding labelling strategies, and applying those-- perhaps those numbers are strongly derived from those things.
Anyway, suppose you’re building an AI to walk, there’s nothing creative about selecting 9.8m/s/s for gravity that’s simply the ideal value to achieve a desired goal. Labeling an elephant as “Elephant” rather than “coat hanger” is similarly a functional choice.
Just because a person is holding a camera and taking a photo doesn’t mean the result is copyrightable.
Suppose you're not building a strawman, but instead building an AI to be an LLM. The exact sequence of what you choose to do for instruction tuning, and the metrics and labels that you choose, the prompt/response pairs you write, and the loss functions you employ are quite creative. They greatly affect the coefficients and are not simple mechanical steps and are the result of a large amount of creative choice.
We are nowhere near a point where they are an uncreative, mechanical recipe to follow.
> Just because a person is holding a camera and taking a photo doesn’t mean the result is copyrightable.
No, but in the overwhelming majority of circumstances it is. What it depends upon is whether the person holding the camera is making a significant, original creative choice.
I am not sure what courts will decide, but I am certain that there is more creativity and originality employed than you are giving OpenAI et al. credit for.
Creative choices requires intentional control over the output across a meaningfully different range of viable possibilities. A brick layer has a huge range of viable options in the specific brick and its alignment in a wall but none of those choices are artistically meaningful.
The coefficients are also not in any meaningful sense chosen based on instruction tuning. It’s no more under direct control than the specific arrangements of atoms in the brick wall and is instead the output of a purely mechanical process.
> We are nowhere near a point where they are an uncreative, mechanical recipe to follow.
Thus: We are nowhere near the point where the output is under creative control rather than being the result of a poorly understood mechanical recipe.
Here's what the supremes said in Feist V. Rural:
> Factual compilations, on the other hand, may possess the requisite originality. The compilation author typically chooses which facts to include, in what order to place them, and how to arrange the collected data so that they may be used effectively by readers. These choices as to selection and arrangement, so long as they are made independently by the compiler and entail a minimal degree of creativity, are sufficiently original that Congress may protect such compilations through the copyright laws. Nimmer ss 2.11[D], 3.03; Denicola 523, n. 38. Thus, even a directory that contains absolutely no protectible written expression, only facts, meets the constitutional minimum for copyright protection if it features an original selection or arrangement.
Alphabetical order wasn't quite enough. But people directing the work that produces the coefficients are doing considerably more creative work than that.
> Thus: We are nowhere near the point where the output is under creative control rather than being the result of a poorly understood mechanical recipe.
No one requires complete creative control of the output. I can spatter paint and have relatively poor control of what's happening, but I am certainly generating a copyrightable work when I engage in creative choices as part of this.
I agree it’s more effort but the metric isn’t effort so I disagree that qualifies the coefficients as copyrightable. The SHA256 hash of a movie isn’t copyrightable even though the movie itself was.
> No one requires complete creative control
That’s a strawman, there are requirements for creative control. You don’t own copyright to your normal dumps, but you can get copyright from looking down and selecting to take a picture. That’s the low bar for a creativity requirement, but it exists.
I know the metric is no longer effort. But there's a lot of creative choices that I've mentioned that greatly affect the coefficients, even if we don't know what those creative choices are going to do to each film grain in the photograph or coefficient in the matrix.
> You don’t own copyright to your normal dumps
Yes, there's an explicit exemption in LOC's guidelines for things that are the direct output of natural processes.
If you have a lot of choices affecting output, then the output is subject to copyright. Indeed, the Supremes above said that factual contemplations can qualify if they involve a "minimal degree of creativity".
In the end, we'll see.
Sure, there are "poems" that consist of just a groups of numbers that are copyrighted. They are not encodings, it's just a string of numbers. It's indistinguishable from a bunch of numbers. This is just one example, there are lots.
They are enforceable to the degree it's creative, and to the degree the infringing use is also creative.
So you would not be able to sue me for using those numbers in a math equation. You would be able to sue me for reproducing your poem in a book of poems :)
As feist says, the creativity required for copyright is quite minimal. But it's still only as protectable as it is creative.
Look - AI is not the first thing to have this "issue". The answer remains the same as it always was - it's mostly about the process not the output.
The output mostly matters is if the output is not intended to be creative (or it's de minimis or ...).
Copyright as it currently exists is weird.
Like if you go to the copyright office and try to register your ssh public key and say "this was generated by ssh-keygen i had nothing to do with it" you may get a different result than if you said "this is my new visually stunning masterpiece, my ssh public key, which was generated with computer help but I used 37 precisely timed keyboard smashes to do it. Prints are available from my gallery for $500"
> Like if you go to the copyright office and try to register your ssh public key and say "this was generated by ssh-keygen i had nothing to do with it" you may get a different result than if you said "this is my new visually stunning masterpiece, my ssh public key, which was generated with computer help but I used 37 precisely timed keyboard smashes to do it. Prints are available from my gallery for $500"
The important thing, of course, isn't whether the copyright office denies to register your copyright, but instead what courts will ultimately do when you attempt to enforce your copyright.
We know the current administrative algorithms used by the copyright offices. We have less clarity on what courts will ultimately do.
There are random number books and I know one of them has a copyright registration [1].
[1]: https://publicrecords.copyright.gov/detailed-record/7060844
Even random numbers are copyrightable.
Below is an implementation of Marsaglia's invention, from p348, courtesy infamous NR[1]. Its a MWC (multiply with carry) random number generator, with two parameters, variable a and base b=2^32. --- For a, "The values below are recommended with no particular ordering." ID a B1 4294957665 B2 4294963023 B3 4162943475 B4 3947008974 B5 3874257210 B6 2936881968 B7 2811536238 B8 2654432763 B9 1640531364 --- as we all now know, the whole thing is copyrighted - you can't redistribute that code and can't use those specific numbers to generate random numbers without purchasing a license, which only allows you to use it once in your personal machine; that's why GSL[2]. The pseudorandom numbers you would get from MWC if you use above numbers are also copyrighted since they are work-product.
[1]http://numerical.recipes/book/book.html [2]https://www.gnu.org/software/gsl/design/gsl-design.html
Ie. if ClosedAI says you can't use output of their API to train competitive models - is that enforceable or not?
Ie. if somebody creates company that sells milkshakes and they say you can't use them to feed employees of competing milkshakes companies - it wouldn't fly, would it?
For LLMs it really depends on the situation. If it's presented in a EULA scenario, where you already bought the rights to use the LLM and ClosedAI gave you the EULA with additional terms afterwards, then the logic above applies. But then everyone knows EULAs aren't very enforceable these days, and nobody buys packaged software any more, so this scenario is quite unlikely these days.
So, if the clause is just one of the many conditions in their main contract of service, of which you had ample opportunity to review before purchasing/agreeing to use their service, then as long as the terms are legal (eg. don't contradict some law), parties are generally free to agree to whatever they want in a contract, and courts will generally uphold those terms.
"can't use them to feed employees of competing milkshakes companies" is probably enforceable. Sounds silly, but I can't think of any reason why it wouldn't be upheld. Unless there's antitrust factors involved.
"Can't use output of their API to train competitive models" is most likely enforceable. Unless there's antitrust factors involved. These kinds of terms are pretty common too. Nobody seriously thinks they're unenforceable per se.
Of course there are practical barriers to enforce a contract -- the aggrieved party has to discover the breach, gather sufficient evidence, and file a lawsuit. As an average Joe individual, you're probably not worrying about getting sued by a company for trivial breaches of service agreements. Most likely the service provider will just cut the service instead of spending thousands of dollars tracking you down (and risk taking a PR hit for going after the little guy). But between businesses, the risks of getting sued by a competitor is real, and no sane lawyer would advise the business to ignore such contract terms.
(Btw, I am not a lawyer. I've studied these things a bit though.)
I don't have a strong sense of whether this is reasonable (I see arguments both ways) but I do think it's pretty strongly at odds with how we treat photographs. There are a bunch of photos on my phone where I unquestionably own the copyright, despite putting in much less creativity than I did for some AI images I've generated.
I don't think it's clear how to resolve this, but I do think that if we are going to protect photos and not prompted AI images, the distinction needs to turn on something other than whether "sufficient creativity" was applied to the input of the mechanical system.
Edited to add: It's probably also worth calling out that the question of whether we protect the work produced by a person's use of mechanical system is a separate one from whether we protect the work of others when it is (in various ways, to various degrees, with various likelihoods) reproduced by use of those mechanical systems.
The interesting part is, those controversial case are pretty recent when the art of photography is century(ies?) old now. I wouldn't expect super clear guidelines regarding AI art before a few decades of weird cases fought tooth and nails in court.
https://petapixel.com/2018/04/24/photographer-wins-monkey-se...
https://www.copyright.gov/comp3/chap300/ch300-copyrightable-...
313.2 Works That Lack Human Authorship
As discussed in Section 306, the Copyright Act protects “original works of authorship.” 17 U.S.C. § 102(a) (emphasis added). To qualify as a work of “authorship” a work must be created by a human being. See Burrow-Giles Lithographic Co., 111 U.S. at 58. Works that do not satisfy this requirement are not copyrightable.
The U.S. Copyright Office will not register works produced by nature, animals, or plants. Likewise, the Office cannot register a work purportedly created by divine or supernatural beings, although the Office may register a work where the application or the deposit copy(ies) state that the work was inspired by a divine spirit.
Examples:
• A photograph taken by a monkey.
• A mural painted by an elephant.
...
Similarly, the Office will not register works produced by a machine or mere mechanical process that operates randomly or automatically without any creative input or intervention from a human author. The crucial question is “whether the ‘work’ is basically one of human authorship, with the computer [or other device] merely being an assisting instrument, or whether the traditional elements of authorship in the work (literary, artistic, or musical expression or elements of selection, arrangement, etc.) were actually conceived and executed not by man but by a machine.” U.S. COPYRIGHT OFFICE, REPORT TO THE LIBRARIAN OF CONGRESS BY THE REGISTER OF COPYRIGHTS 5 (1966).- whether the monkey has the copyright (PETA's argument): this was smacked down by the court twice, the second court explicitely setting a precedent.
- whether the photograph has coypright: as far as I know he doesn't, as the work was deemed non copyrightable (ruled as not created by a human)
https://en.wikipedia.org/wiki/Monkey_selfie_copyright_disput...
The output is not.
The photo you take involved choices of composition and timing and equipment choice.
Just because you don't feel you put in a lot of consideration does not mean at a fundamental level that you still put in creative choices that give the resulting product copyright protection.
But if you took that photo and put it into software which made a derivative image without human creativity, then while the original image would be copyrightable the resulting derivative output would not.
The ideal path forward in copyright would be no infringement in use of materials for training and no protection in AI output with infringement possible against output too close/derivative of protected images.
It's not infringement if you learned to draw tracing Mickey Mouse, your Gerry Gerbil cartoons are fine, but if you draw Mickey Mouse and distribute it, you'll hear from Disney's lawyers.
AI should be the same with the exception that the Gerry Gerbil cartoons would not be copyrightable.
Excellent! I'll put the Inheritance Cycle through a synonymiser, and have a copyright-free (if somewhat degraded) version. Take that, Christopher Paolini!
… wait.
What you say might well be correct: the law is often foolish. But I'd imagine the creativity-free derivative work still counts as a derivative work of the original, copyright-eligible work.
One way it might not be is if Stable Diffusion is seem more like a hash algorithm than a synonymiser. But I don't see why it should, because there's a meaningful correspondence between the input and the output of the system.
I can't legally pirate Windows just because the source code was run through a compiler. Even though the compiler itself adds no additional creativity, the underlying source code is still a creative work[0], so pirating the binaries still infringes a copyright. Just one that's in a slightly different place than what we're normally used to thinking about.
Just to drive the point home, there's a few other situations in which copyright "flows through" to things not subject to copyright. Back in the days of copyright formalities, if you published before properly registering something, your work would be born into the public domain. And this occasionally happened to serial media - e.g. someone might just forget to register the third season of a TV show. In that particular case[1], seasons one and two are still copyrighted, and because season three is a derivative work of the prior season, nobody but the original owner can actually make any use of season three. The only practical difference is that the company that owns that TV show lost one year of copyright ownership over the third season.
[0] By definitions of law. I honestly think most software shouldn't have been made copyrightable, but once Congress said "software is copyrightable" that put that question to bed.
[1] I don't remember the name of the TV show or the court case, but this IS a thing that happened and this theory IS court tested.
Presumably that general rule would also prevent your instructions to DALLE from giving you copyright ownership of the output either. The AI isn’t getting ownership, so it’s either in the public domain or a derivative work from the artists creating the training data.
If a museum can include a small portion of a frame around a public domain painting and claim new copyright as a result - surely any smallest spark or creative influence qualifies, including choosing a single word and choosing the model and time and which output is selected does as well.
The idea of work for hire, and the notion of copyright assignment, applies to people and not machines or processes employed in the creation of a work. Your brush manufacturer would never dare try to claim that their creative selection of fibres and thus their contribution to the unique brush patterns in your painting constitutes a creative contribution to your work. Why is a complex digital model which does the same any different?
Perhaps it is copyright as a whole that is wrong and is nothing to do with AI. This is what we get for creating imaginary property as a means to finance speculative creative endeavours in a capitalist system. So, yeah. Fun times there - once again technology challenges another economic status quo.
There’s no way to map DALLE prompts into any kind of obvious picture from the input. Even DALLE itself can produce a wide range of outputs from a single input.
Further, people have programmed in languages before any compilers where created which worked after the compilers where created.
Early CPU’s didn’t compile anything they directly executed the instruction pipeline.
Also, the number of CPU’s manufactured heavily favored very simple designs.
As to your point that’s not what CPU’s do though, they have both a set of instructions and a set of IO with the outside world. A compiler always results in the same output from a given set of instructions, but with CPU’s you can run the same code and get wildly different output due to that IO.
The only way you can call a CPU a compiler is as a subset of its capabilities. If they they have internal microcodes where a given instruction gets translated into a different internal representation, but that’s not the end it also executes those microcodes.
And the AI image involved choices of prompt and model, and subsequent selection from among several generated images.
I recognize that what you said here:
> Your prompt for the AI image generation is copyrightable.
> The output is not.
... probably represents the state of the law at the moment (with meaningful amounts of uncertainty), but I don't think there's a principled difference based on the amount or nature of creativity involved. IMO the equivalent would be "you own the specification of (position, equipment, relevant world state) but not the photo" which obviously doesn't do anything we want for photography. And I guess that's a part of my point. We should pick the policy we want to make sure we capture the incentives we want. Maybe it is best that AI assisted art (past some point?) not be copyrightable. But I don't think basing the distinction on the amount or nature or... propagation (I guess?) of creativity makes any sense in distinguishing flippant and bullshit photographs (at least a third of my photos, although I would hesitate to apply the labels to any particular photo by someone else) from prompt-driven generative works.
Are you arguing here that because the weights come from an optimization program, they are not "human authored"? If so I find that to be a strange assertion. If I'm working every day on my model and training algorithm to ensure it produces the best weights possible to solve my problem, I would be very surprised for someone to tell me I have no ownership over those weights because they are generated from a program I wrote and data that I own.
It might end up uncopyrightable due to other reasons, but probably not this.
Btw, the phone directories are quite different -- they're just compilation of raw factual public domain information. The LLM weights are anything but. In fact one of the leading theories as to why LLMs are not copyrightable is that they infringe upon the copyrights of the source training materials. (I also don't want to guess whether that argument holds)
Let me put a straw man, and try to find a middle point, when the copyright argument stops being applicable:
1. A painting was done by an artist.
2. On a computer.
3. With a help from an image processor software.
4. Using some advanced filters, like super-resolution, that utilize computer vision techniques. Like neural networks.
Many smartphones already automatically process your* photos with some advanced CV algorithms. That can be called "machine generated art".
I'd personally prefer to stop saying "neural network did X", same way as we don't say "a bulldozer built a road, a crane built a house".
Even non-generative-AI inside Photoshop only mutates images. Generative AI is the source of images.
Or, is the distinction you are making based on there being an image before the model is used?
Maybe a published copy of the weights might be copyrightable, in the exact form of a “creatively” ordered listing, but the weights themselves would almost certainly not be if the US judicial system rules consistently.
This bypasses the entire argument of whether it is human authorship as weights themselves in bulk are just straight up non-copyrightable regardless of origin according to this reading of the law and precedent.
An entirely reasonable, if not fully tested, statement is the following:
Every single one of these AI weight things itself is a result of unencumbered, massive, law-breaking, right-violating copyright infringement -- accordingly, it's extremely difficult to say anything morally justifiable or authoritative about anyone elses "rights" downstream, and to try to inject the word "ethical" makes the whole thing even more ridiculous.
Why? Copyright covers expression not information, AIs can learn information from any source regardless of copyright. They should just not regurgitate copyrighted content, that's all. And much of what organic content is online is common knowledge, thus can't be copyright-controlled.
"You wouldn't look at a car and then remember what that looked like when someone asks you to draw another"
*this is not legal advice, dangit commenter person below
They could be completely biased, could completely ignore everyone and everything else and rule however they want.
I'm almost surprised they still bother to write any kind of "legal reasoning" in their ruling and don't simply focus on what the ruling is rather than why they ruled that way. But I guess such "reasoning" still serves a propaganda purpose and still provides a fig leaf for those who still believe in the quaint absurdity that "we are a nation of laws, not men."
Plenty of companies who have legal teams will keep an eye on the legal landscape of court decisions, and use them to decide if our T&C's or contracts need rewriting, or if any precedent puts us at legal risk.
Sure - the supreme court could overthrow its precedent anytime, but until it does, a lot of people will act as if what they say is the law.
Not really for practical purposes. In the long term, the Supreme Court can and does overrule its own precedent, so the first case on the specific issue to get to the Supreme Court doesn’t end the discussion.
In the short-term, cases get resolved by lower courts and parties either lack funds to do the maximum level of appeals, or the Supreme Court chooses not to hear appeals (they tend to prefer an issue to be well-developed with circuit case law, often waiting till there is a conflict between the Circuit Courts of Appeal, before taking it up), so the state of the law prior to any specific ruling on the narrow topic by the Supreme Court matters quite a bit.
They don't necessarily do. Think about that. You can take some copyrighted material and transform the information contained in it (for instance a fictional book). You can then write a summary. The summary contains information that was present in the original but it has been transformed and hence it's not a copy. The ML model contains information that has been generalized by some degree. So it's just a grey area IMO.
The summary also contains original thought, something is added to it by a human to make it unique. AI models are primarily deriviative.
A better example would be: if I take 1,000 different copyrighted works and put them into a ZIP file, does that resulting file violate copyright?
The copyright aspect makes more sense when you start thinking of AI training models as lossy compression for the original works. Is a downsampled copy of the new Star Wars movie still protected under copyright?
Just tabulating the word counts would not violate copyright as it is considered facts and figures.
Like, if one has access to such a model, and doesn’t count it towards the size cost of a compression/decompression program nor as part of the compressed size of the compressed images, then that should allow for compressing images to have substantially fewer bits than one would otherwise be able to achieve (at least, assuming that one doesn’t care about the amount of time used to compress/decompress. Idk if this is actually practical.)
But unlike say, a zip file, the model doesn’t give you a representation of like, a list of what images (or image/caption pairs) it was trained on.
Or like, in your analogy with the lower resolution of the movie, the lower resolution of it still tells you how long the movie is (though maybe not as precisely due to lower framerate, but that’s just going to be off by less than a second, unless you have an exceedingly low framerate, but that’s hardly a video at that point.)
There is a sense in which any model of some data yields a way to compress data-points from it, where better models generally give a smaller size. But, like, any (precisely stated) description counts as a model?
So, whether it is “like lossy compression” in a way that matters to copyright, I would think depends a lot on things like,
Well, for one thing, isn’t there some kind of “might someone consume the allegedly infringing work as a substitute for the original work, e.g. if cheaper?” test?
For a lower resolution version of Star Wars movie, people clearly would.
But if one wanted to view some particular artwork that is in the training set, I would think that one couldn’t really obtain such a direct substitute? (Well, without using the work as an input to the trained model, asking it to make a variation, but in that case one already has the work separate from the model, so that’s not really relevant.)
If I wanted to know what happened in minute 33 of the Star Wars movie, I could look at minute 33 of the compressed version.
A mono audio version of Star Wars, compressed down to 320x240, filmed from the back of a theater on a VHS camera, converted to Video CD, would under any reasonable interpretation be just a copy of the original.
I assume it starts getting murky when there's some sort of transformation done it it. What if I run motion capture on it, and use that motion capture data to create a cartoon version of Star Paws (my puppies in space epic)? What if I do a scene for scene recreation as the animated cartoon (removing any mentions to copyrighted names -- Luke Skywalker is now Duke Dogwalker, for example)? In this case, there's been no actual data transfer -- all the sprites are hand drawn, backgrounds etc.
What would be an interesting exercise would be to try and create a series of artifacts that each on their own are considered non-derivatives, but can be used together to reconstitute the original. For example, create a compression method that relies heavily on transforms / macroblocks, but strip out any of the actual pixel data from the film. That info might be supplied as palette files which are themselves not really copyrighted data, but together with the compressed transform stream can be used to recreate the original video.
However, saying "the first word of the book is 'The'" would not be a violation, while repeating that for every word in the book, as a whole, would be one.
All that to say, the models have potential to memorize, but they don't, and if they do it's an undesirable failure mode, not some deliberate copying.
Once it starts doing well, will DC come after me? You bet.
But also, an unintentional copy of a copyrighted image is not a violation of copyright. (eg: an executable binary which happens to contain the bits corresponding to a picture of Batman -- but which are actually instruction sequences and were provably not intended to encode the picture -- clearly doesn't infringe.)
LLMs are somewhere in-between #1 and #2, and the intent can happen both in the training and also the prompting.
Stack on top of this the fact that the models can also definitely generate content that counts as fair use, or which isn't copyrighted.
It's the multitude of possible outputs, across the copyright spectrum, combined with the function of intent in training and/or prompting, which make this such a thorny legal issue for which existing copyright statute and jurisprudence is ill-suited.
Taking your Batman example: DC would come after you for trademark as well as copyright, and the copyright claims would be very carefully evaluated with respect to your very specific work. But here we are talking about a large model that can generate tons of different work which isn't subject to copyright or which is possibly fair use.
I don't think that existing jurisprudence (or even statute?!) can handle this situation very well, at all, without tons of arbitrary interpretative work on the parts of juries/judges, because of the multitude and vague intent issues described above.
(...Also presumably the merits of the DC case wouldn't matter because your victory would be pyhrric unless you are a mega-corp. Which from a legal theory perspective is neither here nor there but from a legal practicality perspective may inform how companies go about enforcing copyright claims on model weights/outputs.)
Anyways. I think we have a right mess on our hands and the legislature needs to do their damn jobs. Welcome to America, I guess :)
Curious to hear your thoughts on these issues.
Like, yes, but it's not very likely to happen and it's not a particularly horrible thing if it doesn't; the law is slow and little-c conservative and you're just expecting it to be something it MOST often just ain't.
(Yes the idea of rights is also unnatural and absent from visions such as anarchy)
What does “more principled” mean of a moral code? How does one quantify “degree of principledness”?
> and use that to investigate what might be considered "more natural".
What does the preceding (being “more principled”) have to do with being “more natural”? And what significance does being “more natural” have?
And none of that has any relevance to what is usually described as “natural rights”; its like taking existing words and coming upnwith entirely novel meanings and then a whole architecture around them, which is pretty advanced equivocation.
Edit to add I'm not saying I agree with the justification or am trying to argue for it, only that the point above is commonly raised as the justification, implying that the intrusion on a person's rights is known and accepted.
1. Evaluate every output from the model to ensure that none of the outputs are copyrighted
2. Evaluate every input to a model to ensure that the inputs are either not copyrighted or properly licensed
3. Change the definition of copyright so that ML models can do whatever they want
Nobody is doing #1, because that makes the business models not work. Established brands (like Adobe) are doing #2. I get the feeling that there are a lot of ML startups that are hoping that #3 will happen, but it seems unlikely
A model training being rendered fair use doesn't mean any of its output can be used for whatever regardless.
That's what I listed as #1 - evaluate each individual output of the model to see if it violates copyright.
PS: your "#1" is really hard to do and I'd guess it is infeasible. Even Google (esp. Youtube) with their vast data capabilities, often gets it wrong.
I suspect we're going to see the same kind of rethink about intellectual property in the age of AI.
Not that we really had all that much privacy in the past, as anyone who's browsed old newspapers knows.
The act of choosing to place images in a certain arrangement, such as a collage, can be copyrightable. The same could be said for the "act" of choosing what images to include in a training set and which parameters to use to train the model.
There is in fact a whole art form where people cut out words from different newspapers and books, for example, and re-arrange those words to form new and interesting art.
So there are ways in which such a work would be a creative work, and ways it which it would not, and it would depend on the particular instance and example.
So your phonebook modifications may or may not be considered "creative" depending on the judge and your ability to convince them. The more your modify it, the more likely you are to convince a judge it is a creative work, though.
The US has the "threshold of originality" as its principle. Under that doctrine, it requires some human (and this has been emphasized many times over the years) originality in order for something to be copyrighted. It's a low bar for how original it needs to be, but it must be human (monkeys taking selfies are not human).
https://en.wikipedia.org/wiki/Threshold_of_originality
In England, the doctrine is "sweat of the brow" instead.
https://en.wikipedia.org/wiki/Sweat_of_the_brow
> Under a "sweat of the brow" doctrine, the creator of a work, even if it is completely unoriginal, is entitled to have that effort and expense protected; no one else may use such a work without permission, but must instead recreate the work by independent research or effort.
The definitive case for this in the US that set the two apart is Feist Publications, Inc., v. Rural Telephone Service Co. ( https://en.wikipedia.org/wiki/Feist_Publications,_Inc.,_v._R.... ) where it was deemed that a telephone directory is not copyrightable in the US as there is no originality in it... but under the sweat of the brow doctrine it would have been.
So the "[c]opyright is for things that are the result of human creativity" gets an "it depends" and it would be curious to see if companies that are firmly in the "models are valuable" camp go to the UK for what I believe would be a more favorable copyright protection.
... However there are other IP laws around trade secrets that may be better for it in the US (I'm not as familiar in that domain - I would be curious to find out).
A photo is presumed to be copyrightable. Even horrible photos taken by somebody without any aesthetic sense are presumed to be copyrightable. The argument (AFAIK) is that the photographer chooses the time, location, object, and tweaks various settings of the camera (exposure, aperture, etc.), and these choices are considered sufficient for a photo to be copyrightable.
How about LLMs?
The hyperparameters of LLMs are hugely important in training LLMs, as is the choice of source training data. To me the "degrees of freedom" (and hence room for "creativity") in training LLMs are larger than that of a photographer taking a photo. And as of today, training a good LLM is probably objectively harder than taking a good photo, even if we forget about hardware costs for a moment.
It's easy to convince judges and juries that copying phone numbers into a phone book doesn't require human creativity. But we're talking about the most bleeding edge tech companies producing a bleeding edge new product here. I think it's going to be really hard to convince judges and juries that making this new shiny thing doesn't require human creativity. Maybe in say 20 years when even a 10 year old can train a LLM the situation might change, but as of today, quite unlikely IMHO.
Maybe the popular and free ones. Adobe has a product in beta that uses "ethical training data" as a selling point.
Who am I kidding this is Adobe of course they're fucking over their users
Is there an official ruling? Or is it just a Reddit-style over exaggeration?
They are trained on a lot of text. News sites, comments, books etc. Most books and news sites fall under copyright. Is this fair use? Who knows. Fair use is also an American thing. ChatGPT can be used in the EU, which doesn't have such a broad view of fair use.
If you make a game only out of a lot of copyrighted assets without paying it isn't fair use. Are LLMs different?
What about image generation, which you can prompt the models for specific styles of artists, which works are all copyrighted, but still used for training?
IANAL, but if weights are IP, wouldn't they constitute a "derived work" of the training data?
If I read a few books about a subject as research, and then I write an article about the subject, it's my own copyright. The fact that I did research doesn't make it derivative of those books (correct me if I'm wrong, IANAL).
Perhaps a model created from copyrighted material be treated in the same way?
Yes, because in that case you'd be the "author" doing "creative work".
> Perhaps a model created from copyrighted material be treated in the same way?
Who would be the author doing creative work in this case? The people who decided what training material to use? Perhaps, but it seems a stretch for the people who selected the training material to be authors but not the people who created the training material.
Basically, that world ignores the AI model completely. If your resulting work wouldn't be fair use if you directly were working with something from the training set, it wouldn't be fair use if you fed it through an AI model first.
The inputs could be copyrighted and the weights could be copyrighted if creating the weights from the inputs is (legally) regarded as a transformative use. And I think it could reasonably be considered to be transformative - the weights don't look anything like the input data.
Disclaimer: IANAL. So far as I know, no court has ruled on whether this qualifies as a transformative use. I take no position on how the courts will actually rule. I merely say that they could regard this as transformative use. (But see jerf's "creativity" argument for another hurdle that weights must pass to be copyrightable.)
Google's thumbnails are a purely mathematical transformation on images (no copyright themselves), and yet are considered a transformative use.
I believe that trained models are similarly a purely mathematical transformation of {data}, but is transformative in what that can be used for going forward.
"Can" bearing a lot of weight in that sentence.
It's how the human, with agency, uses the model that may be a derivative or copyright infringing use - not the model itself nor necessarily the output.
The output of a generative AI may be similar enough to an existing work that it is derivative of that work. It is possible to construct a prompt that infringes on an existing work even if that work wasn't part of the training data.
For that case, consider you drew a picture. That picture that you just drew isn't part of any training data. I could presumably look at it and describe it with sufficient detail that something similar enough would be generated... and that may be considered a derivative work. The same test could be applied to me describing it to someone on Fiverr with the same outcome.
If I were to publish that work by the generative AI or Fiverr - who would be infringing on copyright? me? or the black box that may be AI or Fiverr that created a picture based on my prompts?
Their recent decision that implies that anything that AI is used to produce is non-copyrightable is silly, sad, and not sustainable.
On the other end of the spectrum, AI generated content couldn't be copyrighted if there is no human involvement. If someone asks GPT to write 1000 poems, it couldn't be copyrighted.
See “What color are your bits?”: https://ansuz.sooke.bc.ca/entry/23
>> And very much of intellectual property law comes down to rules regarding intangible attributes of bits - Who created the bits? Where did they come from? Where are they going? Are they copies of other bits?
What happens if we train a neural network on a single, copyrighted work? Say it has one input node (or even zero, if you like), and regardless of this input, its output is always exactly the copyrighted work it was trained on. What do its weights represent? Clearly, its weights represent a direct encoding of the original work. Those weights are copyrightable, but not by the person who trained the neural network -- the copyright is held by the owner of the original work.
What if we train the neural network on just two copyrighted works? If its one input node is 0, it outputs the first, and if it's 1, it outputs the 2nd. Almost certainly, its weights are a complicated, tangled mix encoding both, like a compression algorithm that completely rearranged its input. Who owns the copyright to those weights? To whatever extent the weights can be "factored out" into a set representing the first work and a set representing the second, clearly the copyright holder of the first work holds the copyright on the first "factored set", and the 2nd on the 2nd. It seems obvious that we must be able to do this "factoring out" somehow (even if the topology of the factored networks is different), because we know both works are exactly represented by the weights, and the neural network itself can use this information to reconstruct them both, so they're in there ... somewhere. So is there a sort of "joint copyright" on the combined weights, where nobody is really allowed to do anything with it without approval of the other? Regardless, it's still clear that whoever trained the neural network has no claim on any copyright.
Where is the breaking point extending this from 2 works to a billion? People make arguments like "drawing a car from memory isn't infringing on copyright design of that car", which ... are you sure? Reproducing a piece of music from memory (and selling it) is usually copyright infringement. You're allowed to learn a Taylor Swift song as part of your musical training, but you're not usually allowed to then play it back from memory and sell that recording (I'm not sure I morally agree with this treatment of covers, nor if it's globally applicable). So the argument that "surely neural networks are allowed to learn from copyrighted works" misses the point: they can learn all they want, but as soon as they reproduce verbatim (or close enough) a copyrighted work, they're infringing. And if they're representing a complete copy of the work within their weights (which they obviously are if they can reproduce it), then the original copyright holder has a claim on those weights. And never in this process has the trainer of the NN acquired any copyright to anything. The real trainer is a bunch of GPUs, after all.
If the neural network cannot reproduce any of the copyrighted works verbatim, then we're getting closer to "fair use" territory. Yes, it's permissible to write a summary of a copyrighted work. That is so lossy as to not "compete" with the original work in any meaningful way. If it could be demonstrated that neural networks do not encode completed works (no matter how hard the factorization would be), then one could make this argument. Unfortunately, the evidence is that LLMs are more than happy to completely regurgitate copyrighted works verbatim. It seems to me the copyright holder of the original work therefore must hold a share of the claim on the weights. Still, the GPUs that trained the network do not magically acquire copyright over anything.
I wonder if the real answer is that the weights are copyrighted, and that copyright is held jointly by hundreds of millions of people, and nobody can do anything with those weights without the approval of all the others. I'm not saying I like that universe, but I am saying it's the most internally consistent answer I can think of, and seems to follow from the above argument.
On the other hand, we need the original input vector for this to work, and one could argue that the network weights are simply the algorithm for decoding the input vector into the copyrighted work. So the originator holds copyright on the input vector, not the weights. Does it matter if the input vector has smaller information content than the original work? Clearly this argument relies on the input vector being the "actual encoding", and therefore must have at least as much information. If the input vector is an embedding of "please show me the latest Tom Clancy novel in full", this argument breaks down.
Okay, this is hard.
AI weights might be considered a method of production, but that isn't clear yet.
In particular, taking other documents and shoving them through a process that generates a lot of other numbers with no human or creative interaction is definitely something I'd be concerned the courts would judge as not sufficiently creative to be copyrightable. The process itself would certainly consist of copyrightable code, but the output doesn't necessarily. This would be somewhat similar to the observation that there is no copyright to be had in a big table of files and their MD5 hashes (or other hashes), such as a Linux distro might use for integrity checking. Lots of copyright in the original file contents, copyright available on the process for producing these tables, but the tables themselves would likely be ruled not itself copyrightable as there is no creativity in that output.
Note this also has absolutely nothing to do with the question of whether AI output is copyrightable, this is about the huge table of numbers that make up the neural net weights being copyrightable. (Though it would be sort of an interesting question for the legal system to grapple with as to how a non-copyrightable set of numbers could then produce something copyrightable. Call it a philosophical variation on the "copyright washing" argument; can copyright spring from a non-copyrightable source other than a human brain, thus somehow "flowing uphill"? Would a human brain be copyrightable? Stay tuned for those questions, I guess, or if not you, your grandchildren.)
Per your other comments, "work" is not the bar, "creativity" is. "Size" is not the bar either. Merely being a much larger table of numbers than a list of hashes or a phone book is not the question. No human is in that table of numbers creatively saying "no, wait, this neural weight should be -1.5 instead of 2.0 to produce this creative effect". No human is even capable of working in the medium of neural net weights in a creative manner.
If you want to go the "novel legal theory" route, you could play with claiming creativity in the selection of input material and claim the resulting neural weights has a copyright in compilation: https://en.wikipedia.org/wiki/Copyright_in_compilation That's a long way from a slam dunk though. Way out on a legal limb there. It isn't entirely clear to me what exact rights would result from such a claim either. It would be a landmark copyright court case for sure.
> Is a document not copyrightable based on its contents?
Creative process is the bigger issue.
> Weights are just a different kind of a document.
And who sits down and writes this document of weights?
Yes, exactly. It's copyright 101.
For example, if you write a random number generator, and print 10000 randon numbers in a document, it's not copyrightable.
Even if you invented a specific random number generation algorithm, the document is still not copyrightable. Your code is copyrightable.
Again it's just copyright 101. If any of above surprises you, maybe you should read a few copyright case studies.