My goal was merely to warn everyone in the LLaMA community that Facebook appears to be trying to shut down the ecosystem that sprang up around LLaMA since the beginning of March.
For a bit of background, I created llama-dl on March 5.
Show HN: Llama-dl - https://news.ycombinator.com/item?id=35026902
Announcement tweet - https://twitter.com/theshawwn/status/1632238214529400832
Since the repo is now offline, you can find an archived version of the README here: https://archive.is/7t3it
The intent with llama-dl was to kickstart an open source movement related to LLaMA. If you're curious about my personal motivations for this, I did an interview with The Verge about that: https://twitter.com/theshawwn/status/1633456289639542789
Over the next two weeks, llama-dl grew to 3k stars, and (according to my bucket metrics) distributed 4M files. Thanks to the availability of a reliable, high-speed download link to LLaMA, other hackers were able to launch projects such as Dalai:
Dalai: Automatically install, run, and play with LLaMA on your computer - https://news.ycombinator.com/item?id=35127020
Dalai has been making headlines all over the place, and especially on ML tiktok. (ML tiktok is surprisingly interesting.)
When Facebook knocked llama-dl offline via DMCA on the 20th, my primary concern was to ensure that Dalai stayed up. After all, the whole point of llama-dl was to encourage the creation of a "killer app" such as Dalai.
After a quick huddle with Dalai's author @cocktailpeanut via Twitter DM, they launched a decentralized distribution mechanism for LLaMA, powered by bittorrent: https://twitter.com/cocktailpeanut/status/163903613304778342...
This ensures the availability of LLaMA in the short term. However, there's a broader issue at stake.
The question is whether model weights themselves can be copyrighted. It might seem obvious that since compiled binaries can be copyrighted, ML models should also be able to be. But the U.S. Copyright Office recently denied copyright to AI generated outputs: https://www.smithsonianmag.com/smart-news/us-copyright-offic...
> Both in its 2019 decision and its decision this February, the USCO found the “human authorship” element was lacking and was wholly necessary to obtain a copyright, Engadget’s K. Holt wrote. Current copyright law only provides protections to “the fruits of intellectual labor” that “are founded in the creative powers of the [human] mind,” the USCO states.
If the model output isn't copyrightable, is the model itself copyrightable?
It's an interesting and important question, and answering it in court is a necessary step. The outcome will determine how models are treated over the next decade.
Now, all that said, Facebook is proceeding under the (untested) assumption that LLaMA is copyright Meta. If that assumption is correct, then they're well within their legal rights to issue these DMCAs. Llama-dl was little more than a bash script pointing to a download link, yet that's sufficient grounds for DMCA, since the whole point of llama-dl was to circumvent a copyright protection mechanism.
My overall goal here is to simply bring awareness to all of these issues. We're entering an era of closed-source ML. I think the history of computing shows that open source is generally a better bet.
Facebook, if you're reading this, I urge you to reconsider your approach. You had the opportunity to gain an incredible amount of momentum. By killing it off, you're sacrificing your foothold into the hearts and minds of ML hackers. Wouldn't it be a better idea to harness the ecosystem rather than stomp it out of existence? There are so many ways this can facilitate your business in a positive way. Are you sure that being an adversary to your own community is the best way forward?