I found the OpenAI bot scraping my blog recently. Assuming they used that data, when will they attribute me?
I found the OpenAI bot scraping my blog recently. Assuming they used that data, when will they attribute me?
Given that Google successfully used a fair use defense in Authors Guild, Inc. v. Google, Inc., I think it's likely OpenAI and the others will also win in court.
I do think it's possible for specific uses of the output of LLMs to be copyright infringement. That's why it's interesting to see Microsoft to indemnify customers of their commercial products in the event a case is brought against the customer. This is smart on Microsoft's part; the risk probably isn't very high and by making it a non-issue for their customers, many more will feel comfortable using their LLM-based features and services.
Like githubs servers host AGPL code as data, without having to be open-source
The perceived problem there, is if their model generates an exact copy of some AGPL code, and you use it in your project unknowingly, and then you get can sued
Note: i'm declaring my comment license as https://creativecommons.org/licenses/by-sa/4.0/
So if you remix or transform my comment by responding it, please attribute to me your response.
profit vs non profit also makes a difference
>I found [logs of users from Paramount's writers offices reading] my blog recently. Assuming they used that data, when will they attribute me?
To see that the idea on the face is silly. OP has no evidence that any of their work was used at all, or even that what was used could even be covered under the license in the first place.
I believe that this fact is and will be exploited to strip copyright and effectively transfer ownership using cleanroom/firewall techniques.
That's the key part. You haven't yet proved they have actually used your content for anything (other than, potentially, read the license to decide if they should include or discard from their training set).
But in practice we'll never know for sure if they are respecting the terms of licenses until 1) this is tested in court, or 2) there's some internal leak that points into either direction.
I would think OpenAI wants the thornier legal issues actually settled so that the whole ecosystem can grow within those terms & they can lobby for the legal changes they need/want?
.. wants the thornier issues to be debated and re-tried ad infinitum, as long as they generate cash flow and build their moat(s).. more likely
This behaviour seems more consistent with wanting is sorted out than stalling for time.