Can I use my Outputs to train an AI model?
support.claude.com
support.claude.com
We did so, please do not repeat it at home.
Claude and AI partners have taken all what they could from the Open Source projects without giving credit or respecting the licenses. They have increased the traffic on websites in an absolute disrespectful way increasing the hosting cost in inefficient and ridiculous ways.
They have taken all the important books and not asked permission from the authors.
Fair enough. Fair use.
Of course I would create a competitor software to Claude or any others if I could. Using Claude(and others) of course.
I am not paying you 200 dollars/month for you to tell me that I could not create code that competes with you. If you try to go to court in Europe with this you will lose.
It is just the same fair use you proclaim for taking the data from others.
Copyright infringement isn't piracy. Nor theft.
There might be a few exceptions, if you really guided your model in a certain particular way, but that is rare.
What /could/ be is that you violate their ToS. If that ToS is valid and legally and practically enforceable is a different question and will depend a lot on your jurisdiction.
It’s kind of like causing what Google Docs autocomplete sentences helps someone write is still.. theirs.
It's named "yours" at the time of payment, but actually only a license to use under certain conditions.
This basically makes their AI useless.
Originally it was not that way. They changed it because, as you suggested, their service has questionable utility without it.
The outstanding one is the “meaningful human input” part of copyright. The most recent ruling is that prompting alone does not count. If you write/rewrite sections, those are yours. Everything in between is a somewhat untested.
That sentence seems to be in a blurry line between outright “ownership” and licensing.
This type of explanation leans towards the reality being you don’t own the outputs from Claude.
This kind of an explanation is like trying to be half pregnant.
I don't see them as having a leg to stand on with this in court; they can only cut off your access for egregious violations.
Can you never tell anybody and just slap a license on it? Sure.
The issue here is not the license, but that you violate their ToS (if that is valid and enforceable is a different question of course). But if you publish the LLM output on github, and someone else takes it to train their LLM, and you did not actively encourage or help them, it's fine.
Compiler output isn't copyrightable? Is a compiler human?
Note that this is different from Gemeinfreiheit, where a copyrighted work essentially becomes public domain 70 years after the death of its creator.
but you can't copyright the output anyways. which either makes the restriction on training void or, it means the owners of the model own the copyright, and they only transfer some of the ownership to you. is there such an ownership transfer statement? i haven't seen one yet.
> Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software *without restriction*, [...]
> Safety controls may be lost – models trained on Claude's Outputs won't have our safety measures, potentially leading to harmful or dangerous AI systems. We also have no visibility into deployment, meaning we cannot monitor how these distilled models are used or prevent misuse.
So was that why your models caused three real-world security incidents?
No, that was just marketing :)
> we cannot monitor how these distilled models are used or prevent misuse
which is fine since they are none of your (their) business..
I think the only possible concern is using their models name - making it clear it's a new model merely trained in another one should fix that.
Has any individual somewhere around the world tested this in court by now? Sued Anthropic for copyright infringement because Claude can reproduce information that is only available on their website?
It shouldn't be that expensive, right? If you sue them for - say - $10000 then what would the costs of such a court case be?
Personally, I think "learning" is not a copyright violation. But if they themselves make it one, then they should also face the consequences, no?
if i publish that data, and someone else trains on it, am i liable? am i responsible to ensure that noone trains on that? how am i supposed to enforce that?
These agreements are trying to hold back the tide with their legal team's hands. This already happens at such massive scale that it's obvious they're completely helpless to stop it.
The author must be human, at least for now, in the US this precedent holds as far as I can tell:
https://law.justia.com/cases/federal/appellate-courts/cadc/2...
From my understanding distillation pretty much copies behaviour. If someone intends to distill from a Model, it can't extract unsafe behaviour, but would learn the same safety mechanisms, no?
> When Outputs are used to train new models without our oversight, additional risks emerge. Safety controls may be lost - ...
---
> What you can do with Outputs
> You can use Claude's Outputs to train models that don't compete with Anthropic's own models.
Why even include that bullshit at top? It is brazenly obvious to anyone with a brain that Anthropic and other frontier AI Labs have no real way to monetize or recoup their investment unless they strongly guard the usage of their models.
I’m not sure how an AI company feels that’s safer for their business. Specialization will always best generalization trying to do the same.
"We need to control the model to protect you from evil robots, except it's fine if the evil robots are not competing with our business" is hilariously hypocritical.
Now what.
It's also pretty wild to call this standard practice. It's not. I can grab any of the open weights models and train to my hearts content. So it's not standard is it. You'd like it to be standard because you don't want to compete.
And you don't trust us, but it is your company that's been going around telling us how excited you are that your model goes out onto the internet and hacking people.
This continual authoritarian bent from the least trustworthy people in the world is deeply problematic and the only saving grace is their absolute total and complete failure to enforce the restrictions they wish to place on us.
Your ability to build the God machine doesn't magically endow you with the moral authority or judgement to decide how it's used, and the fact that these people believe it does is a great indicator that they aren't to be trusted.
That's not even the case. No one "owns" the output. Raw generative AI outputs don't come with any new (copy)rights, which are currently only granted to human creative outputs.
Theoretically, this should be legal and ethical, when comparing to Anthropic's own behavior. That said, the reason you can't is Anthropic states in their terms that they don't want you to do this.
Anthropic's entire business model is skating on thin ice.
Or, what if I generate content with Claude/ChatGPT/Gemini, warp it in HTML using an open model, put this on my website conveniently dedicated to "Best practices in prompt and AI answers" for example, then train my own model that only scraps my website?
What they don't want is DeepSeek training their models with Claude output at scale. That's why they forbid it. Gives them a legal basis to cut off accounts doing that.
Not that it's effective because it's being done anyway.
And yes it's super hypocritical but that's another issue IMO.
So it would be supporting third party's attempts...
Then again. I suppose lot of scrapping could happen by accident and scrappers might ignore such files as allowable-use-cases-for-site-content-must-followed.txt instructing against use in training.
If such is ignored wouldn't be our fault right?
Or you can publish them on the web and countless others will do it.
The LLM companies are Pygmalion; the LLMs themselves are golems or Frankenstein's monster; we ourselves are merely Igor.
Or -- maybe better -- just fsck the TOS.
It's amazing how quickly you can spin from that to ownership with limitations.
Obviously it's bullshit. They can't say it's theirs, licensed to you because that would stop people using it commercially... So we have this Gordian knot of logic to explain that we should pay to used their models and infrastructure, and not use the output how we like because their models and infrastructure are theirs and exploiting that would be really mean. The argument makes them look like a petulant child.
But it's just a weak EULA, a flimsy non-compete. They say output is yours? It's yours. Do whatever you like with it. But don't be surprised if they limit access if Anthropic decide it's unsavoury.
Hell, arguably they should release their weights (or at least the weights of their older models), since they trained them on the concentrated knowledge of humankind.
Except there’s no “we” since their models aren’t open.
ah yes, this long-running, time-honored traditional industry, where we totally didnt write the rules ourselves
I noticed a similar thing for the Antrophic’s previous announcement on open weight models.
I don't care about your policies, go fuck yourself.
Wait, we're already doing it. /s