No. Both legally and practically, you absolutely do not.
The only thing copyright law gives you is an exclusive right to sell it for a limited period of time, as a whole in its original form or similar -- and to transfer that right.
Regardless of your desires, anyone can reuse it under the conditions of fair use. They can copy parts of it for parody purposes. If they're not selling anything or taking away from your sales*, they can reproduce it verbatim for private purposes. And even if they are selling something, they can summarize it, quote from it, rephrase it, and so forth.
And you don't actually get to decide any of that.
* Edit: added "or..."
> I wasn't asked and I don't really care to donate work to large corporations like that... I do get to decide what happens with it.
And I said:
> No. Both legally and practically, you absolutely do not.
You think you get to decide whether large corporations can train on your work. I'm saying the the law suggests you very much don't get to decide that.
Send some links if you see some definitive case law sorting this stuff out.
If ChatGPT regurgitates verbatim or nearly verbatim, something it slurped up from OP's blog, is that not plagiarism? Where do you draw the line? Where would a reasonable person draw the line?
Often rather than claiming human aspects to the machine, they are going further, and claiming machine aspects to the human.
Using mechanistic analogies for explaining the human body or mind isn't new, but as machines become better and better at imitating humans, those analogies become more seductive.
That's my rant; the danger with 'AI' isn't so much that humans are enslaved by machines, but that we enslave each other -- or dehumanize each other -- with machines.
You are entitled to control it's distribution and use. You are not entitled to control it's influence and effects.
AIs are not massive repositories of harvested data. The models are relatively small (<20GB).
https://www.pinsentmasons.com/out-law/news/google-thumbnails...
> A US court ruled this week that Google's creation and display of thumbnail images does not infringe copyright. It also said that Google was not responsible for the copyright violations of other sites which it frames and links to.
> The Court said that Google did claim fair use, and that whether or not use was fair depended on four factors: the purpose and character of the use, including whether such use is of a commercial nature or is for non-profit educational purposes; the nature of the copyrighted work; the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and the effect of the use upon the potential market for or value of the copyrighted work.
Taking copyrighted material and using it to train a model is not a copyright infringement - it is sufficiently transformative and has a different use than the original images.
Note that AI models can be used for different things. A model trained to identify objects in an image has never had uproar about the output of "squirrel" showing up in the output text.
The model also, as a purely mathematical transformation on the original source material does not get a copyright. If it needs to be protected, trade secrets are the tools to use to protect it. A model is no more copyright worthy than tanking an image and applying `gray = .299 red + .587 green + .114 blue` to it.
The output of a model is ineligible for copyright protection (in the US - and most other places).
The output of a model may fall into being a derivative work of the original content used to train the model.
It is up to the human, with agency in asking the model to generate certain output to be responsible for verifying that it does not infringe upon other works if it is published.
Note that the responsibility of the human publishing the work is not anything new with an AI model. It is the same responsibility if they were to copy something from Stack Overflow or commission a random person on Fiverr... its just that those we've overlooked for a long time - but it is similarly quite possible for the material on those sources to be copyrighted by and/or licensed to some other entity and the human doing the copying into the final product is responsible for any copyright infringements.
Saying "I copied this from Stack Overflow" or "I found this on the web" as a defense is just as good as "Copilot generated this for me" or "Stable diffusion generated this when I asked for a mouse wearing red pants" and represents a similar dereliction on part of the person publishing this content.
Would that satisfy you?
But we do? Open sourcing something with caveats is common. This code is public BUT not for commercial use. This code is public BUT you must display attribution etc.
Sure, blogposts are unlicensed (that I know) but the idea of something publicly available being held to restrictions is nothing new.
But telling me it is illegal to share what I learnt because the original source is copyrighted... doesn't sit right with me.
What protects particular solutions is patents. For example if someone were to obtain a patent for computing GCD of large integers the usual fast way, well then everyone else would have to use a different solution.
This analogy to someone reading a book, perhaps peppered with lots of legalese to the point of being hardly recognizable, will definitely be used in courts at some point. And I can't see how it wouldn't stand as a valid argument.
If you pick up a book and learn a fact, then yeah, you’re allowed to share that fact.
It’s weird that this topic keeps devolving into a form of “so what, it’s illegal for me to learn things?” Because: no, it’s not. And: You and a piece of software are treated differently under the law. You have a different set of rights than ChatGPT.
Gods, no. Where did you get that from?
Those might not be a problem regarding this specific case, but the case can easily be made that it ought to be.
I don't understand your point. Do you think it makes any difference whether I use my laptop, or a pen, or ChatGPT to violate copyright?
On the other hand, it's completely feasible to make a license that stops someone from training their model with some piece of info, is it not?
And regardless -- the problem now is that expectations of how content can be consumed are now fundamentally violated by automation of content ingestion. People put stuff up on the Internet with the expectation of its consumption by human minds, which have inherent limitations on the speed and scale on which they can learn from and reproduce things, and those humans are also legally liable, socially/ethically obligated, etc.
Now we have machines which skirt the limits of legality, and are able to do so on massive scale and without responsibility to society as a whole.
Different game now.
Then people obviously aren’t aware that bots have been indexing web pages and showing summarized information without going to the web page for three decades.
Further, almost every site has had an e.g. robots.txt which has permitted content harvesting only for certain accepted purposes for a couple decades now. So clearly people already had a sense of how they wanted their content harvested and for what purposes.
So you’re okay with Google making money off of your content. But not OpenAI?