That is what we have been told when they stole our open source software.
That is what we have been told when they stole our open source software.
That said, the inclusion of GPL-licensed code in training sets may yet force the release of those models under the GPL.
I'd really like to see this tested at a court.
If I make a better search engine than Google and take away their ads business, I have damaged them, but I am not liable.
The argument here is that Facebook does not actually own LLaMA, because they don't own the training data, they didn't have humans curate the training data in a creative way, and the actual training process is purely mechanical. If LLaMA is not copyrightable then you cannot be liable for copying it.
But financial damage on programs distributed for free is...nothing. So even if you sue, what are you going to get out of them?
I wonder if you can sue for some form of specific performance and cause them to remove your and all other equally licensed code from their training
Not necessarily.
For instance, if the program is distributed free for a limited set of purposes under a license, but available for a negotiated license (with payment) for other purposes, then the reasonable market value of a license without the restriction would be actual damages.
2) the value of machine learning training on any one particular piece of source code also approaches zero
If a model replaces a business case for software you were giving away for free, how are you financially harmed? Even if you won a lawsuit you can't demonstrate any financial impact of software you give away for free.
This is the whole principle that allows the GPL to work.
That aside, we should really stop misusing "steal". Not only is it legally inaccurate (Dowling versus United States), it's semantically inaccurate as well. The conflation of that with mere copyright infringement is a campaign driven by bad faith actors.