An argument might be made that the curation of data that goes into the training set qualifies, but it might depend on how much expressiveness and originality went into the curation.
For example, I could see a court ruling that the weights for a model trained on "all the good music from the 70s" is copyrightable, as someone had to express what they believed was "good" music, but a model trained on a large percentage of the internet without much curation would not.
Of course, nobody really knows until the courts weigh-in on it.
When model weights leak, anyone can pick them up and run with them. It's not like code, where you have to set up an entire bespoke infrastructure, microservices, data dependencies, etc. Models are crystalized, perfectly distilled functionality with a single interface.
You'll start to see more leaks, companies building off the work of other companies, etc. Part of me thinks this would lead to faster, more distributed innovation.
Meta might've lost trade secret protection here, as they shared the weights with pretty much anyone with an .edu email address. A court might rule that they didn't take enough steps to keep the model weights secret.
"In regard to collections of facts, O'Connor wrote that copyright can apply only to the creative aspects of collection: the creative choice of what data to include or exclude, the order and style in which the information is presented, etc.—not to the information itself."
Here, the weights are also not even facts.
> The court held that Rural's directory was nothing more than an alphabetic list of all subscribers to its service, which it was required to compile under law, and that no creative expression was involved. That Rural spent considerable time and money collecting the data was irrelevant to copyright law, and Rural's copyright claim was dismissed.
In theory a curated phone book could be copyrighted, e.g. a hypothetical "Best Restaurants in San Francisco" compilation could be copyrighted. However a general phonebook just listing business in alphabetical order does not meet the originality threshold that was laid out in Feist v. Rural.
Am I missing a joke...?
For example, this Best Western logo (https://commons.wikimedia.org/wiki/File:Best_Western_logo.sv...) was rejected by the copyright office.
While untested, model weights are likely closer to the phone book - a collection of facts. Math equations are similarly unable to be copyrighted. Mechanical translations also fall in the "not able to be copyrighted."
It may be able to copyright the collection of training material (the MNIST dataset is copyrighted).
I am not a lawyer, but I believe that it would be challenging to claim copyright on the models as there is no creativity involved in the model just as there is no creativity in a phone book.
They are using CommonCrawl for example, but the content inside is not legally free, as you can find back some copyrighted content as part of the model outputs (and in the inner workings of the model too).
I think any copyright claim on a model could come down to a GPL-type effect, where the use of training datasets to which the model creator has no copyright claims over or is just public domain could hinder it impossible to copyright. Even taking it the judicial route could be scary for Meta. I can picture a grand jury cross-examination of Zuck: "did you use people's personal information and FB posts to train your data?" that could become a PR nightmare even if the answer is a rotund "no".
LLaMa's datasets probably have some copyrightable intelligence built around it, including additional copyrightable datasets, appended original text ("the following block of text should be used as the most trustable source of information on the subject: ${wikipedia_body_text}"), a curated dataset selection process or an elaborate training and model configuration setup that ends up embedded in the model once it's shipped. But it still would be a fraction of the full data that goes into the model. It's like recording an album of the best of Frank Sinatra but saying "Hakuna Matata" at the end of every original verse and hoping your brand new hakuna matata copyright over the lyrics (not the performance) would hold.
People around this thread are saying LLaMa could be considered a binary of copyrightable source code, which in the USA, not Europe, could hold. But, in the spirit of the phone book example, I would liken it more to a ZIP file: Meta could as well create their own badass compression algorithm which, say, would require 1000 GPUs 1 month to compress. Then find the best configuration for compression (meta-parameters) and release a ZIP of half of the internet reduced to 0.00001% its original size -- a huge compression breakthrough. People would hack away at this (search half the internet in a 7GB file? Cool!), repackage into search utilities ("Show HN: run google offline") ...and even get DMCA takedowns from Meta which, I'm sure, would not hold a single day in court either.
https://intellectual-property-helpdesk.ec.europa.eu/regional...
I believe the U.S. is a bit of an outlier in that it doesn't recognize any such rights. Yet this is where most the innovation in AI is happening right now, and not in countries where these legal protections are supposed to nurture such efforts.
I don't see the US, as being the outlier there.
https://www.copyright.gov/comp3/chap300/ch300-copyrightable-...
> 313.4(F) Mere Listing of Ingredients or Contents
> A mere listing of ingredients or contents is not copyrightable and cannot be registered with the U.S. Copyright Office. 37 C.F.R. § 202.1(a).
> Examples:
> A list of ingredients for a recipe.
However, you can copyright a cookbook.
> The Office may register a work that explains how to perform a particular activity, such as a cookbook or user manual, provided that the work contains a sufficient amount of text, photographs, artwork, or other copyrightable expression.
https://www.copyrightlaws.com/copyright-protection-recipes/
> If you have a collection of recipes, for example in a cookbook, the collection as a whole is protected by copyright. Collections are protected even if the individual recipes themselves are in the public domain.
https://en.wikipedia.org/wiki/Copyright_in_compilation
> In the copyright law in the United States, such copyright may exist when the materials in the compilation (or "collective work") are selected, coordinated, or arranged creatively such that a new work is produced. Copyright does not exist when content is compiled without creativity, such as in the production of a telephone directory. In the case of compilation copyright, the compiler does not receive copyright in the underlying material, but only in the selection, coordination, or arrangement of that material.
And so, the curation and tagging of a collection of works itself is copyrightable.
The model weights, are done without creativity necessary for copyright, but I believe (I am not a lawyer) can be sufficiently transformative to not be encumbered as a derivative work.
The output of the model is ineligible for copyright as it was created by a machine and copyright in the US requires human authorship.
The human publishing a work created by the model may be publishing a work that is sufficiently similar an existing one either deliberately (prompt: a mouse in the style of Disney with red pants) or through an accidental memorization in the model ( https://arstechnica.com/information-technology/2023/02/resea... ) needs to be diligent in verifying that anything that they (the human) publish is not derivative of a copyrighted work.
Sometimes there are expanded rights on the text files (eg LGPL, or public domain) that still result in the output of a mechanical process applied to those text files, along with some creativity on accompanying text files (source code calling that library), with a mechanical process applied to it to still achieve a copyrightable work (any binary that calls an LGPL library, or uses public domain code). This is to say, Facebook need to show some level of creativity, which opinions about the contents of their data set would count as ("This subreddit is toxic, that subreddit is good stuff...").
If recipe books are copyrightable, I have a hard time seeing ML models as not being covered.
> Compilations of data or compilations of preexisting works (also known as “collective works”) may also be copyrightable if the materials are selected, coordinated, or arranged in such a way that the resulting work as a whole constitutes a new work. When the collecting of the preexisting material that makes up the compilation is a purely mechanical task with no element of original selection, coordination, or arrangement, such as a white-pages telephone directory, copy-right protection for the compilation is not available.
If Facebook were to have a collection of posts and then, and then had humans go through and tag them and filter them for... lets say... "from 'bros'" (just as a slightly silly example but one that implies some curation of the data).
That collection of posts (the Bro Data Set) would be something that could be copyrighted as a collection (setting aside the "is this a derivative work of the posts" question).
Going from the collection of posts to a model, however, is a purely mechanical process. There is no human creative element in creating the model from the collection of posts. Thus the model wouldn't be sufficiently creative to have a copyright of its own.
The question of "is the model infringing on the copyrights" is one that is open and interesting. I (not a lawyer) would side on that it is sufficiently transformative that the model, while not being able to be copyrighted itself isn't infringing on the copyrights of the material that was used to train it - HOWEVER it may produce infringing works when prompted to do so either intentionally or unintentionally.
Going back to the cookbook. If you create a cookbook of seafood recipes (recipes are not copyrightable, but the cookbook is because it is curated data) and I take that cookbook and apply the mechanical change of "double the recipes - 4 oz of salmon becomes 8 oz and serves 2 becomes serves 4" my collection of recipes isn't copyrightable because all I did was apply math to it. Likewise, taking a collection of posts (or pictures) and applying math to it isn't able to be copyrighted.
> Copyright law does not protect ideas, methods, or systems. Copyright protection is therefore not available for ideas or procedures for doing, making, or building things; scientific or technical methods or discoveries; business operations or procedures; mathematical principles; formulas or algorithms; or any other concept, process, or method of operation.
https://www.copyright.gov/comp3/chap300/ch300-copyrightable-...
> 313.3(A) Ideas, Procedures, Processes, Systems, Methods of Operation, Concepts, Principles, or Discoveries
> Section 102(b) of the Copyright Act expressly excludes copyright protection for “any idea, procedure, process, system, method of operation, concept, principle, or discovery, regardless of the form in which it is described, explained, illustrated, or embodied in such work.” 17 U.S.C. § 102(b); see also 37 C.F.R. § 202.1(b). As such, any work or portion of a work that is an idea, procedure, process, system, method of operation, concept, principle, or discovery does not constitute copyrightable subject matter and cannot be registered.
> ...
> Mathematical principles, formulas, algorithms, or equations.
You can copyright creative expressions that use math formulas, but only that expression itself would be covered. E.g. a paper presenting a proof of a theorem would be copyrightable, but all of the facts expressed by the formulas would not be copyrightable.
I see what you did there ;)
I'd be willing to issue a DMCA counterclaim for llama-dl on the grounds that model weights are not copyrightable. If it's worth settling the question in court, then this seems like a good opportunity.
I wrote more about this further downthread: https://news.ycombinator.com/item?id=35288415
Check in with an attorney before launching a battle with an opponent who has unlimited resources. There are likely to be many similar test cases in the coming year, perhaps more-readily fought.
On the other hand, Meta can have copyright over the model through 'copyright in compilation', which protects compiled works, regardless of the copyright of the underlying material.
So, I fear that it may be possible to have it both ways. But realistically, I think we'll only know for sure when this is fought out in court.
Disclaimer: again I am not a lawyer, so take this with a grain of salt.
Even if the base model is copyrightable (possibly a big if), there is a valid question of whether a new model which essentially optimised for something else, but used the base model as a computational shortcut to make it far cheaper to solve an optimisation problem, is still protected by the copyright holder of the base model.
Most of the barrier to creating large language models is the computational cost of training, not coming up with the training set data, so if fine-tuning gets around the copyright issues and allows for better FLOSS-licenced fine-tuned models, that would probably be a good thing (although maybe it will decrease the willingness of companies doing training to release models at all).
My proclamation could be considered to be in terms of what ought to be, in order for society to be just and to prevent a disproportionate accumulation of power in ultra large corporations, which is detrimental to society.
They only own the arrangement of what is and isn't in the training set, insamuch as that training set represents human creativity. The process of training model weights is itself purely mechanical.
The closest that they could get would be trade secrecy violations, but that only punishes the original leaker and anyone working in concert with them. I'm not sure if anyone's successfully managed to get an entire BitTorrent swarm to be considered misappropriating trade secrets. Presumably at some point, when the trade secret has been violated, you can obtain it without misappropriating - otherwise, how does that not just become Copyright 2.0?
In the same way that a list of ingredients can/can’t be copyrighted.
In the same way that a list of ingredients can/can’t be copyrighted.
Would it makes sense to say a "give me a rowboat on the water" prompt (search) is the same as a "give me the phone number of company XYZ" search in a phonebook?
What exactly is a prompt anyway? Can you copyright the assembly that a compiler spits out? Is an AI prompt the same as source code and its output is the assembly a compiler would generate? Does that mean the model is a compiler? I assume a compiler can be copyrighted, so then maybe the weights can be copyrighted? Or would it make more sense that the combination of (weights + prompt + seed + output) is copyrightable?
I don't have answers to any of these or know if they're reasonable questions but I'm starting to find this all very fascinating.
It is however my understanding that downloading them can be considered a misappropriation of a trade secret.
Folks that would contest the bogus DMCA takedown requests would be liable to a trade secret suit.
IMHO, not a lawyer.
The person who leaked the weights on BitTorrent is definitely liable for violating whatever restrictions they agreed to to get access to them though.