I don't think distillation as 'stealing IP' has any legal legs. They can probably claim violation of ToS at best because the terms of use do prohibit use for training rival models.
I don't think distillation as 'stealing IP' has any legal legs. They can probably claim violation of ToS at best because the terms of use do prohibit use for training rival models.
Not copyrightable IP, or at least it hasn't been challenged yet. I have experience with this: I made llama-dl, a way to download the original llama model. Meta issued a DMCA, I appealed to the HN community for funds, someone funded, and our lawyer successfully counterclaimed. Never heard from Meta again.
A lack of response to the counterclaim doesn't mean the issue is settled. But now there's legal precedent for people pushing back against companies that claim model weights are secret IP and therefore DMCA-able.
Precedent is a different thing -- that's a court decision that other courts could/should/must follow.
The idea is that selling, or giving away art grants the consumer to sell art of his own, not a clear copy of.
So the court rationally found Anthropic guilty for acquiring copies without paying for them. But didn't charge the training, deemed fair use.
If it isn't fair use, then we may need to sue all teachers for spitting out knowledge they ultimately acquired from someone's work.
Copyright, some would say is irrational. Humans learn, that is copying, distilling in fact.
Copyright though is pragmatic: it draws a line, art will be reproduced, let's just enforce that they can't be shameless (near) identical copies. Mix it up, derive the original enough so that you aren't competing with the original author.
The only argument to ban training on copyright data is that it unfairly compete with original authors. Which stance do you take? Neither would be irrational.
> If it isn't fair use, then we may need to sue all teachers for spitting out knowledge they ultimately acquired from someone's work.
Is such a lazy argument and always was. A human doing something is necessarily different than a machine doing it. We can be ok with a human doing a thing and simultaneously not ok with a machine doing it.
And in doubt, to pause and forbid commercialisation given the clear impact it's already having on people's financials, humans have invested in a given set of rules that ML are disrupting.
My argument is simply that the ruling for fair use isn't irrational. It's a rationality that you, me and many others would disagree with. Lazy? Yes. Also call it lazy, and biased towards VC's interest who can't squeezed much profit in status quo.
Either everyone was stealing all along, and they should all be sued out of existence, or nobody is stealing.
Personally, I think its all fair use and support all of the training, including distilations.
I suppose you think that model providers are not allowed to impose restrictions on the use of their model? Think carefully before answering.
They should have exactly the same ability to "restrict" it is as the writer of a book on those who read the book that was purchased.
And seeing as those model providers were not restricted from training on those books.... It would follow that other people would have the exact same right to train on the output of the models.
I simply demand that model providers are treated exactly the same as the data that they trained on. Either it was OK for them to train on other people's work, en mass, without permission, over the objection of the creator, and therefore its OK to do the same to them.
or none of its ok, and they should presumably be equally sued into oblivion, and equally shut down completely by the government.
Thats all. Take your pick. Either all of the training on either books and all the models, without permission, is ok or none of it is.
Additionally, the output of a model isn't even copyrightable. So actually there would be even less protections for that. Because of this, it seems that anyone could use it for anything.
EX: 3rd parties aren't bound by the TOS of the models. So someone could simply do a passthrough, and give the uncopyrightable output to someone else to distill, and since the distiller didn't sign the TOS they would be in the clear to train on non copyrightable info.
It seems you do not understand what distillation is.
I am saying that someone else would use that output to create or improve a different model. I think you could have figured out that this was the meaning of my statement instead of doing the irrelevant nitpick that you did.
Its also unrelated to my point, which is that person 1 could give the data from model A to person 2, and person 2 would be fine because they didn't sign any TOS contracts with model A, and the output from model AI is not copyrightable and therefore can be redistributed.
Additionally, it still doesn't address the point about how the original model trained on a bunch of other people's stuff without permission, so I don't see why the same shouldn't be done to their outputted content.
Do you have any substantive disagreements or are you just going to make a minute, incorrect nitpick and then not elaborate?
I don't think they are. Not copyright at least. There may be some "trade secret" stuff for them but they're not copyrightable.
Contrast this with Apple's case where it looks like they've got evidence of people walking out of the building with various physical artifacts on their way to an OpenAI interview.
Also, more seriously, there are plenty of pieces of information that are compelled to be published that don't lose their confidentiality. There actually is a general concept that confidentiality can be lost if you are negligent in its protection but it's a very nuanced thing.
But you can do an analysis of a bottle of Coke to figure out exactly what’s in it.
You can even make your own and sell it, if you call it a different name.
Ultimately it all hinges on the specifics of the training and how much human involvement was involved throughout the creation process. I suspect arguing this successfully in favor of upholding a copyright would be an uphill battle for many situations, but ultimately we'll need more court cases to know for sure.
The copyright office's recent statement on this matter were sensationalized when they really said nothing groundbreaking at all -- this was always the standard applied when any tools are used in creating a work.
Model training is expensive. Model distillation is cheaper. If you train models, and that lets your adversary distill them for much cheaper, that gives your adversary an advantage. Giving your adversary an advantage is bad, so you address it.
It doesn’t matter whether the advantage was gained legally, not even if it was fair. The advantage in and of itself is bad, and sufficient reason to act. You couch it in less nakedly power hungry terms, but this is the reasoning.
This is not too different from drug discovery where it's extremely difficult to come up with the molecule, but relatively easy to copy it. Similarly, it's really hard to create frontier models from scratch but much easier to distill them.
Assume that their model output is considered IP and that it's ruled illegal to train on that IP. I will offer to sell every content producer on earth an identity LLM that takes their content and outputs precisely identical content that they can then post. Good luck ever getting any training data for free ever again.