Facebook is going after LLaMA repos with DMCA's
twitter.com
twitter.com
one great move that FB could have done would be to ride the wave of positive PR + get all investors hyped:
---> "Meta is a credible alternative to OpenAI, the company is switching from "Meta"-bullshit to an "IA"-first company",
and get the investors to pump the Meta stock,
and then dilute some of the shares to raise some cash (or issue new shares to newly specialized IA hires).
But no, FB is still going after the VR gimmicks and NFTs.
4 billion USD per quarter wasted on Oculus (!), while they could use this money to fund and support a whole ecosystem around LLaMA.
My feeling is that a lot of Meta's AI/ML work actually ties into the AR/VR long-term dream. How do you make the so-called metaverse alive? By having people design it themselves. They're not going to do that in Maya, that's for sure. But if they could create virtual spaces and virtual people with Holodeck-style natural language instructions...
Yes, it's totally AI, and IA is the French version :)
Programmers in particular tend to overuse anglicisms, and often I end up mentally doing spanglish in my head,
"Necesito este value" "Este command deberia hacer esto" "Este iterator va a hacer $something a este object"
That's interesting. For me, the port metaphor makes me think data port, which seems like a pretty apt parallel to the flow of goods at a shipping port.
Different metaphors for different folks, I suppose.
Computer = tietokone = knowledge machine
Swap file = heittovaihtotiedosto = throw exchange record set (or some such).
So yeah, Finnish programmers rarely used the official words.
In abstract terms, those weights are by far the most dense form of meaning we've ever dealt with.
This is probably why they hate/fight these leaks. Once the model weights are available, they lose much of their competitive advantage.
I feel like soon, ai will be it's own researcher and go far beyond what we ever could do on our own.
However it does open up a point which is: should we allow people to make huge amounts of money by infinitely copying and distributing their own work? Should we be protecting this model?
Not only do they have no street smarts whatsoever, but their book smarts start to disappear when you deviate from the training data.
You just can't compete with Steam. It's been tried. All competitors are mercilessly drowned in the giant fucking moat surrounding GabeN's castle.
PC gamers are fiercely loyal, "generally" more informed then other consumers and absolutely despise any forms of dark patterns. Because Steam has always done right by PC gamers Steam has benefited from that loyalty for a good 2 decades now.
Everyone is welcome in GabeN's castle but y'all gotta pay the toll to cross the drawbridge.
It took a lawsuit to get refunds for Steam.
>absolutely despise any forms of dark patterns
Steam has been doing lootboxes and NFTs for more than 10 years
The new batch of kids don't seem to mind using Game Pass or picking up Fortnite of the epic games store, either
* "How Valve is Profiting from Steam's Back-Door Casinos": <https://www.youtube.com/watch?v=eMmNy11Mn7g>
* "Working at Valve: 'A Fearless Adventure' or 'Lord of the Flies'?": <https://www.youtube.com/watch?v=s9aCwCKgkLo>
VR has a legitimate niche in gaming but outside that it's just not appealing. It's dystopian and depressing. Nobody wants to spend time in a social network with a helmet on their head being served ads.
In what way does this NOT describe the situation in America?
That’s one big reason the stock is up: $4B in the metaverse is nonsense. $4B in AI, though? Transformational
I'm not sure if Horizon falls into "virtual reality" or "social platforms" but it seems to be the latter: "For example, we have launched Horizon Worlds, a social platform where people can interact with friends, ..."
Let 'em burn it all.
This could have been their “contribution to the society” - yielding much better PR than they could have ever hoped for.
If not, then whatever they say is effectively law.
Also, it could be a maneuver to prevent the genericization of the word LLaMa, which they may want to continue using.
(Also: great username!)
EDIT: DMCA here [0]. It does sound like they're asserting copyright on the "content," i.e. presumably the weights themselves.
[0] https://github.com/github/dmca/blob/master/2023/03/2023-03-2...
The other interesting aspect of this is that they're classifying it as unauthorized content distribution. Meta was already distributing the weights, but limited their distribution to "researchers" with approved credentials. It was one of those researchers who leaked the weights originally. So it's not like they were reverse engineered from a binary or exfiltrated out of FB HQ. That might be an important bit of nuance.
The EU is different, it recognizes copyright-like rights in databases and database works, which is why the cavalier attitude of U.S.-oriented organizations to these matters tends to annoy me. For example, the FSF does not actually check that certain non-code data files are legally unencumbered. They merely disclaim any copyright of their own. But for all we know, that could be wishful thinking.
> If a work's traditional elements of authorship were produced by a machine, the work lacks human authorship and the Office will not register it.
However, this guidance seems primarily about the stuff you create _with_ AI, not necessarily the model weights used in the AI itself.
[0] https://www.federalregister.gov/documents/2023/03/16/2023-05...
[1] Discussed on HN: https://news.ycombinator.com/item?id=35191206
(Also, another element to this case is that GitHub is owned by Microsoft, who has a conflict of interest with Meta in terms of ChatGPT vs. LLaMA)
An argument might be made that the curation of data that goes into the training set qualifies, but it might depend on how much expressiveness and originality went into the curation.
For example, I could see a court ruling that the weights for a model trained on "all the good music from the 70s" is copyrightable, as someone had to express what they believed was "good" music, but a model trained on a large percentage of the internet without much curation would not.
Of course, nobody really knows until the courts weigh-in on it.
When model weights leak, anyone can pick them up and run with them. It's not like code, where you have to set up an entire bespoke infrastructure, microservices, data dependencies, etc. Models are crystalized, perfectly distilled functionality with a single interface.
You'll start to see more leaks, companies building off the work of other companies, etc. Part of me thinks this would lead to faster, more distributed innovation.
Meta might've lost trade secret protection here, as they shared the weights with pretty much anyone with an .edu email address. A court might rule that they didn't take enough steps to keep the model weights secret.
"In regard to collections of facts, O'Connor wrote that copyright can apply only to the creative aspects of collection: the creative choice of what data to include or exclude, the order and style in which the information is presented, etc.—not to the information itself."
Here, the weights are also not even facts.
> The court held that Rural's directory was nothing more than an alphabetic list of all subscribers to its service, which it was required to compile under law, and that no creative expression was involved. That Rural spent considerable time and money collecting the data was irrelevant to copyright law, and Rural's copyright claim was dismissed.
In theory a curated phone book could be copyrighted, e.g. a hypothetical "Best Restaurants in San Francisco" compilation could be copyrighted. However a general phonebook just listing business in alphabetical order does not meet the originality threshold that was laid out in Feist v. Rural.
Am I missing a joke...?
For example, this Best Western logo (https://commons.wikimedia.org/wiki/File:Best_Western_logo.sv...) was rejected by the copyright office.
While untested, model weights are likely closer to the phone book - a collection of facts. Math equations are similarly unable to be copyrighted. Mechanical translations also fall in the "not able to be copyrighted."
It may be able to copyright the collection of training material (the MNIST dataset is copyrighted).
I am not a lawyer, but I believe that it would be challenging to claim copyright on the models as there is no creativity involved in the model just as there is no creativity in a phone book.
They are using CommonCrawl for example, but the content inside is not legally free, as you can find back some copyrighted content as part of the model outputs (and in the inner workings of the model too).
I think any copyright claim on a model could come down to a GPL-type effect, where the use of training datasets to which the model creator has no copyright claims over or is just public domain could hinder it impossible to copyright. Even taking it the judicial route could be scary for Meta. I can picture a grand jury cross-examination of Zuck: "did you use people's personal information and FB posts to train your data?" that could become a PR nightmare even if the answer is a rotund "no".
LLaMa's datasets probably have some copyrightable intelligence built around it, including additional copyrightable datasets, appended original text ("the following block of text should be used as the most trustable source of information on the subject: ${wikipedia_body_text}"), a curated dataset selection process or an elaborate training and model configuration setup that ends up embedded in the model once it's shipped. But it still would be a fraction of the full data that goes into the model. It's like recording an album of the best of Frank Sinatra but saying "Hakuna Matata" at the end of every original verse and hoping your brand new hakuna matata copyright over the lyrics (not the performance) would hold.
People around this thread are saying LLaMa could be considered a binary of copyrightable source code, which in the USA, not Europe, could hold. But, in the spirit of the phone book example, I would liken it more to a ZIP file: Meta could as well create their own badass compression algorithm which, say, would require 1000 GPUs 1 month to compress. Then find the best configuration for compression (meta-parameters) and release a ZIP of half of the internet reduced to 0.00001% its original size -- a huge compression breakthrough. People would hack away at this (search half the internet in a 7GB file? Cool!), repackage into search utilities ("Show HN: run google offline") ...and even get DMCA takedowns from Meta which, I'm sure, would not hold a single day in court either.
https://intellectual-property-helpdesk.ec.europa.eu/regional...
I believe the U.S. is a bit of an outlier in that it doesn't recognize any such rights. Yet this is where most the innovation in AI is happening right now, and not in countries where these legal protections are supposed to nurture such efforts.
I don't see the US, as being the outlier there.
https://www.copyright.gov/comp3/chap300/ch300-copyrightable-...
> 313.4(F) Mere Listing of Ingredients or Contents
> A mere listing of ingredients or contents is not copyrightable and cannot be registered with the U.S. Copyright Office. 37 C.F.R. § 202.1(a).
> Examples:
> A list of ingredients for a recipe.
However, you can copyright a cookbook.
> The Office may register a work that explains how to perform a particular activity, such as a cookbook or user manual, provided that the work contains a sufficient amount of text, photographs, artwork, or other copyrightable expression.
https://www.copyrightlaws.com/copyright-protection-recipes/
> If you have a collection of recipes, for example in a cookbook, the collection as a whole is protected by copyright. Collections are protected even if the individual recipes themselves are in the public domain.
https://en.wikipedia.org/wiki/Copyright_in_compilation
> In the copyright law in the United States, such copyright may exist when the materials in the compilation (or "collective work") are selected, coordinated, or arranged creatively such that a new work is produced. Copyright does not exist when content is compiled without creativity, such as in the production of a telephone directory. In the case of compilation copyright, the compiler does not receive copyright in the underlying material, but only in the selection, coordination, or arrangement of that material.
And so, the curation and tagging of a collection of works itself is copyrightable.
The model weights, are done without creativity necessary for copyright, but I believe (I am not a lawyer) can be sufficiently transformative to not be encumbered as a derivative work.
The output of the model is ineligible for copyright as it was created by a machine and copyright in the US requires human authorship.
The human publishing a work created by the model may be publishing a work that is sufficiently similar an existing one either deliberately (prompt: a mouse in the style of Disney with red pants) or through an accidental memorization in the model ( https://arstechnica.com/information-technology/2023/02/resea... ) needs to be diligent in verifying that anything that they (the human) publish is not derivative of a copyrighted work.
Sometimes there are expanded rights on the text files (eg LGPL, or public domain) that still result in the output of a mechanical process applied to those text files, along with some creativity on accompanying text files (source code calling that library), with a mechanical process applied to it to still achieve a copyrightable work (any binary that calls an LGPL library, or uses public domain code). This is to say, Facebook need to show some level of creativity, which opinions about the contents of their data set would count as ("This subreddit is toxic, that subreddit is good stuff...").
If recipe books are copyrightable, I have a hard time seeing ML models as not being covered.
> Compilations of data or compilations of preexisting works (also known as “collective works”) may also be copyrightable if the materials are selected, coordinated, or arranged in such a way that the resulting work as a whole constitutes a new work. When the collecting of the preexisting material that makes up the compilation is a purely mechanical task with no element of original selection, coordination, or arrangement, such as a white-pages telephone directory, copy-right protection for the compilation is not available.
If Facebook were to have a collection of posts and then, and then had humans go through and tag them and filter them for... lets say... "from 'bros'" (just as a slightly silly example but one that implies some curation of the data).
That collection of posts (the Bro Data Set) would be something that could be copyrighted as a collection (setting aside the "is this a derivative work of the posts" question).
Going from the collection of posts to a model, however, is a purely mechanical process. There is no human creative element in creating the model from the collection of posts. Thus the model wouldn't be sufficiently creative to have a copyright of its own.
The question of "is the model infringing on the copyrights" is one that is open and interesting. I (not a lawyer) would side on that it is sufficiently transformative that the model, while not being able to be copyrighted itself isn't infringing on the copyrights of the material that was used to train it - HOWEVER it may produce infringing works when prompted to do so either intentionally or unintentionally.
Going back to the cookbook. If you create a cookbook of seafood recipes (recipes are not copyrightable, but the cookbook is because it is curated data) and I take that cookbook and apply the mechanical change of "double the recipes - 4 oz of salmon becomes 8 oz and serves 2 becomes serves 4" my collection of recipes isn't copyrightable because all I did was apply math to it. Likewise, taking a collection of posts (or pictures) and applying math to it isn't able to be copyrighted.
> Copyright law does not protect ideas, methods, or systems. Copyright protection is therefore not available for ideas or procedures for doing, making, or building things; scientific or technical methods or discoveries; business operations or procedures; mathematical principles; formulas or algorithms; or any other concept, process, or method of operation.
https://www.copyright.gov/comp3/chap300/ch300-copyrightable-...
> 313.3(A) Ideas, Procedures, Processes, Systems, Methods of Operation, Concepts, Principles, or Discoveries
> Section 102(b) of the Copyright Act expressly excludes copyright protection for “any idea, procedure, process, system, method of operation, concept, principle, or discovery, regardless of the form in which it is described, explained, illustrated, or embodied in such work.” 17 U.S.C. § 102(b); see also 37 C.F.R. § 202.1(b). As such, any work or portion of a work that is an idea, procedure, process, system, method of operation, concept, principle, or discovery does not constitute copyrightable subject matter and cannot be registered.
> ...
> Mathematical principles, formulas, algorithms, or equations.
You can copyright creative expressions that use math formulas, but only that expression itself would be covered. E.g. a paper presenting a proof of a theorem would be copyrightable, but all of the facts expressed by the formulas would not be copyrightable.
I see what you did there ;)
I'd be willing to issue a DMCA counterclaim for llama-dl on the grounds that model weights are not copyrightable. If it's worth settling the question in court, then this seems like a good opportunity.
I wrote more about this further downthread: https://news.ycombinator.com/item?id=35288415
Check in with an attorney before launching a battle with an opponent who has unlimited resources. There are likely to be many similar test cases in the coming year, perhaps more-readily fought.
On the other hand, Meta can have copyright over the model through 'copyright in compilation', which protects compiled works, regardless of the copyright of the underlying material.
So, I fear that it may be possible to have it both ways. But realistically, I think we'll only know for sure when this is fought out in court.
Disclaimer: again I am not a lawyer, so take this with a grain of salt.
Even if the base model is copyrightable (possibly a big if), there is a valid question of whether a new model which essentially optimised for something else, but used the base model as a computational shortcut to make it far cheaper to solve an optimisation problem, is still protected by the copyright holder of the base model.
Most of the barrier to creating large language models is the computational cost of training, not coming up with the training set data, so if fine-tuning gets around the copyright issues and allows for better FLOSS-licenced fine-tuned models, that would probably be a good thing (although maybe it will decrease the willingness of companies doing training to release models at all).
My proclamation could be considered to be in terms of what ought to be, in order for society to be just and to prevent a disproportionate accumulation of power in ultra large corporations, which is detrimental to society.
They only own the arrangement of what is and isn't in the training set, insamuch as that training set represents human creativity. The process of training model weights is itself purely mechanical.
The closest that they could get would be trade secrecy violations, but that only punishes the original leaker and anyone working in concert with them. I'm not sure if anyone's successfully managed to get an entire BitTorrent swarm to be considered misappropriating trade secrets. Presumably at some point, when the trade secret has been violated, you can obtain it without misappropriating - otherwise, how does that not just become Copyright 2.0?
In the same way that a list of ingredients can/can’t be copyrighted.
In the same way that a list of ingredients can/can’t be copyrighted.
Would it makes sense to say a "give me a rowboat on the water" prompt (search) is the same as a "give me the phone number of company XYZ" search in a phonebook?
What exactly is a prompt anyway? Can you copyright the assembly that a compiler spits out? Is an AI prompt the same as source code and its output is the assembly a compiler would generate? Does that mean the model is a compiler? I assume a compiler can be copyrighted, so then maybe the weights can be copyrighted? Or would it make more sense that the combination of (weights + prompt + seed + output) is copyrightable?
I don't have answers to any of these or know if they're reasonable questions but I'm starting to find this all very fascinating.
It is however my understanding that downloading them can be considered a misappropriation of a trade secret.
Folks that would contest the bogus DMCA takedown requests would be liable to a trade secret suit.
IMHO, not a lawyer.
The person who leaked the weights on BitTorrent is definitely liable for violating whatever restrictions they agreed to to get access to them though.
That is what we have been told when they stole our open source software.
That said, the inclusion of GPL-licensed code in training sets may yet force the release of those models under the GPL.
I'd really like to see this tested at a court.
If I make a better search engine than Google and take away their ads business, I have damaged them, but I am not liable.
The argument here is that Facebook does not actually own LLaMA, because they don't own the training data, they didn't have humans curate the training data in a creative way, and the actual training process is purely mechanical. If LLaMA is not copyrightable then you cannot be liable for copying it.
But financial damage on programs distributed for free is...nothing. So even if you sue, what are you going to get out of them?
I wonder if you can sue for some form of specific performance and cause them to remove your and all other equally licensed code from their training
Not necessarily.
For instance, if the program is distributed free for a limited set of purposes under a license, but available for a negotiated license (with payment) for other purposes, then the reasonable market value of a license without the restriction would be actual damages.
2) the value of machine learning training on any one particular piece of source code also approaches zero
If a model replaces a business case for software you were giving away for free, how are you financially harmed? Even if you won a lawsuit you can't demonstrate any financial impact of software you give away for free.
This is the whole principle that allows the GPL to work.
That aside, we should really stop misusing "steal". Not only is it legally inaccurate (Dowling versus United States), it's semantically inaccurate as well. The conflation of that with mere copyright infringement is a campaign driven by bad faith actors.
Depends on what they're afraid of.
If they want to commoditize the space, then I think you're right.
If they think they can catch OpenAI, and they want to charge for AI services, then what they're doing makes sense.
Then they wouldn't release anything to the researchers. Reason LLaMA weights are spreading in the first place is because they let a big group of people get access, someone is bound to upload it as a torrent.
Seems they gave people with .edu emails access relatively quickly too, researcher or not.
I'm sure FB has shared enormous amounts of data with researchers in the past. And they've probably had very little trouble with said researchers uploading data to the internet.
I think it's a pretty bad take that FB sharing data with researchers means that they wanted it spread freely. Even if that data is LLM weights.
But I'm sure that Facebook when sharing data that is more precarious, like social connections, have been more selective with who they share the data with, compared to the LLaMA dataset which seems to be shared 100% with everyone with a .edu domain, and almost without review for others.
Current Microsoft strategy seems so push into this direction. Open Source certain technologies and acquire important pieces e.g. Github / Stake in OpenAI etc. to build a bigger picture that they can monetize later.
OpenAI is a real threat to Google & Facebook
However, that could be an unfortunate corrolary of the fact that in the US if you do not enforce your IP you give up your rights over it. Overall, giant lose-lose, and I wish there was a truly open source model to build on top of.
That's the case with trademarks but I have never heard of it being the case with copyright.
There have been a lot of challenges, based on “derivative copies,” but these generally seem to be from where someone uses a photo done by someone else, in their own creative work (like the Obama "Change" poster). Some of these challenges succeed, some do not. I would assume someone extending these weights might be considered a “derivative.”
Copyright law is odd, and there’s a lot of “fuzziness,” especially with creative works.
I’ve always been a bit skeptical of applying copyright to compiled and opaque binaries, but I guess I’m in the minority.
You're not in the minority - copyright requiring human authorship and creativity is standing law. A compiled binary cannot be copyrighted on it's own[0]; the copyright flows from the source to the binaries via derivative works. We just don't normally think about this because until very recently all binaries were derivatives of copyrightable human creative expression. Applying a mechanical process to a creative work doesn't make it non-creative, after all.
[0] Hand-assembled binaries would be considered "source code" in this case.
My goal was merely to warn everyone in the LLaMA community that Facebook appears to be trying to shut down the ecosystem that sprang up around LLaMA since the beginning of March.
For a bit of background, I created llama-dl on March 5.
Show HN: Llama-dl - https://news.ycombinator.com/item?id=35026902
Announcement tweet - https://twitter.com/theshawwn/status/1632238214529400832
Since the repo is now offline, you can find an archived version of the README here: https://archive.is/7t3it
The intent with llama-dl was to kickstart an open source movement related to LLaMA. If you're curious about my personal motivations for this, I did an interview with The Verge about that: https://twitter.com/theshawwn/status/1633456289639542789
Over the next two weeks, llama-dl grew to 3k stars, and (according to my bucket metrics) distributed 4M files. Thanks to the availability of a reliable, high-speed download link to LLaMA, other hackers were able to launch projects such as Dalai:
Dalai: Automatically install, run, and play with LLaMA on your computer - https://news.ycombinator.com/item?id=35127020
Dalai has been making headlines all over the place, and especially on ML tiktok. (ML tiktok is surprisingly interesting.)
When Facebook knocked llama-dl offline via DMCA on the 20th, my primary concern was to ensure that Dalai stayed up. After all, the whole point of llama-dl was to encourage the creation of a "killer app" such as Dalai.
After a quick huddle with Dalai's author @cocktailpeanut via Twitter DM, they launched a decentralized distribution mechanism for LLaMA, powered by bittorrent: https://twitter.com/cocktailpeanut/status/163903613304778342...
This ensures the availability of LLaMA in the short term. However, there's a broader issue at stake.
The question is whether model weights themselves can be copyrighted. It might seem obvious that since compiled binaries can be copyrighted, ML models should also be able to be. But the U.S. Copyright Office recently denied copyright to AI generated outputs: https://www.smithsonianmag.com/smart-news/us-copyright-offic...
> Both in its 2019 decision and its decision this February, the USCO found the “human authorship” element was lacking and was wholly necessary to obtain a copyright, Engadget’s K. Holt wrote. Current copyright law only provides protections to “the fruits of intellectual labor” that “are founded in the creative powers of the [human] mind,” the USCO states.
If the model output isn't copyrightable, is the model itself copyrightable?
It's an interesting and important question, and answering it in court is a necessary step. The outcome will determine how models are treated over the next decade.
Now, all that said, Facebook is proceeding under the (untested) assumption that LLaMA is copyright Meta. If that assumption is correct, then they're well within their legal rights to issue these DMCAs. Llama-dl was little more than a bash script pointing to a download link, yet that's sufficient grounds for DMCA, since the whole point of llama-dl was to circumvent a copyright protection mechanism.
My overall goal here is to simply bring awareness to all of these issues. We're entering an era of closed-source ML. I think the history of computing shows that open source is generally a better bet.
Facebook, if you're reading this, I urge you to reconsider your approach. You had the opportunity to gain an incredible amount of momentum. By killing it off, you're sacrificing your foothold into the hearts and minds of ML hackers. Wouldn't it be a better idea to harness the ecosystem rather than stomp it out of existence? There are so many ways this can facilitate your business in a positive way. Are you sure that being an adversary to your own community is the best way forward?
(edit:) Any other interpretation would require Facebook to have provided some other input with expressive intent; and, AFAIK, they did not. (Mere random perturbations of your computational equipment or a random seed that you chose at random without even any attempt at curation are not adding expressive intent.)
https://news.ycombinator.com/item?id=35292445
But like, to repeat some of it in a different way for this slightly different context: for that argument to work, Facebook would have to be actually doing something as the author that wasn't just automated in this codebase.
Like, if you use GIMP to work on an image, and then I download GIMP and merely run it... it doesn't do the same thing, right? You--the artist--were actually important in that story, because you provided the expressive intent that led to the resulting image.
But, in this case, the model is merely the result of running that code, not someone using that code as part of their own work: the model is a reasonably-deterministic output of code licensed under the GPL being run on data notably owned by people other than Facebook.
Imagine if, instead, you wrote a program that used GIMP to automatically create a really fancy image. You put a lot of work into the script to generate that image... and then you released that code under the GPL. I think you would be hard-pressed to argue that the GPL wouldn't have something to say about this output.
(I could maybe see an attempt at an argument that the GPL is an awkward license to apply to things that aren't programs; but, a model is in fact a program designed to execute in an interpreter on still yet other data as input, which makes this whole thing feel like a program which algorithmically generates the code for another program, which is actually quite a common use case for the GPL.)
To be clear: I think Facebook doesn't own the copyright, not that the GPL infected it; but, if Facebook DID own the copyright, AFAIK their only expressive input comes in the form of this GPL codebase. (See above linked comment for more exposition of the possibilities. Also see that comment for a more extensive "IANAL" disclaimer, but: I am not a lawyer, no matter how much I focus on copyright issues.)
In stable diffusion land, for instance, the community is pretty much stuck on the "old" architecture. Newer innovations, like Huggingface diffusers and various optimizations derived from that like PEFT, torch.compile support, AITemplate/TensorRT compilation and various other bits are largely unused.
They are also pretty much stuck on SD 1.5, even though 2.1 is a good base for finetuning.
This has happened in the past too, with ESRGAN.
That doesn't hold up for video/audio, and I think it wouldn't hold up for weights.
That being said, I don't feel like I know for certain how a court would rule. And I wouldn't like to fight Facebook about it in court regardless, that sounds like a pretty bad time even if the court sides with you. But on the other hand, Facebook might not be keen to test this either.
That page itself is for storing the magnet URLs in a bitcoin transaction to make them live forever.
1) You can't copyright weights. A lot of people believe this. I am not sure this is true. I think it might be that there is a fair use argument that the weights are transformative, but having fair use on your infringement doesn't imply a lack of the original copyright being owned by someone. But like, this might be true, and it is not an unreasonable stance.
2) The model weights are a derived work of the training data. This feels right to me, frankly, as much as it irks a lot of people on Hacker News who are excited to use Copilot (or owns shares of Microsoft ;P). In this case, Facebook does not own the copyright as they purposefully used an open training set from third parties (including Wikipedia and OpenCrawl) as a counter to OpenAI's proprietary one.
3) The model weights are a derived work of the training code, in the same way a binary is a derived work of its source code (vs. #2 where the code is a compiler and the source code is the training set). In this case, Facebook would own the copyright... but the resulting binary program--as distributed in the machine interpretable (and executable) format of the model weights--must be GPL as Facebook amazingly used that as the license for their training code.
The interpretation of events that would allow Facebook to have some hope of arguing that the model weights are their own work--that maybe the code is more of a tool like Photoshop and they have a fair use claim on the training data--would imply something that simply is not true: that they are adding some (hopefully extensive) form of expressive input above and beyond those two inputs, in the way someone using Photoshop does when they remix someone else's art into a transformative work.
However, Facebook definitely isn't doing that: they have provided absolutely no expressive input or intent on top of those two inputs, one of which they do not own and the other of which they chose to license for us under GPL! To make this a bit clearer, separate Facebook into two parties for a moment--one which developed the tooling and the other of which ran it--to determine which of the various parties you think owns this result: if you download a program someone else wrote and click a button to run it on some data someone else owns, you simply do not own the copyright on the result. (edit: I wrote some more on this argument in the following linked comment.)
https://news.ycombinator.com/item?id=35293068
The only thing I can come up with, if I try really really hard to steelman Facebook here, is: maybe, if I were to release a binary to the world that is a compiled copy of code that I also simultaneously released under GPL, the binary might technically have been compiled from an internal pre-licensed (but identical) copy of said code; and so, while if you compile the same binary from the licensed code you get an identical output that is a derived work with GPL rights, when I do that I don't... but this feels like a perilous argument to make as you are going to have such a hard time showing that this was a reasonable way to infer your intent with the simultaneous release.
(Of course, the person who originally agreed to the terms of service attached to the download they got is a totally different matter, but that doesn't mean there are none. There are a lot of limitations to what people can extract out of you if you violate a contract. So like, by having explicitly agreed to those terms with Facebook that person should not have given the world the weights--not because they are copyrighted but merely because they were secret--and there might be some kind of ramification... but, I do not believe that would possibly apply downstream to sillysaurusx.)
(Note: I am not a lawyer. I spend a ridiculous amount of my time vs. a normal engineer working with copyright and both hearing and making arguments about copyright, including directly to the Copyright Office at the Library of Congress as part of my work on Cydia and my efforts to push back on parts of the DMCA alongside lawyers from the EFF... but like, it would be foolish to read my comment here and then embark on a project to do something that might massively infringe on someone's copyrights without running it past a real lawyer. I, certainly, have actual lawyers.)
the fork disappears when the parent repository is taken down, clones stay on your machine
upload to IPFS. you can pin on IPFS for free with filecoin. Use filecoin nodes to pin on ipfs at https://web3.storage