I'm hoping for this article to show up again on the frontpage on a EU timezone where I can read some civilized discussion about it.
I'm hoping for this article to show up again on the frontpage on a EU timezone where I can read some civilized discussion about it.
Plus, many technologists believe that copyright should be outright abolished, again because they disagree with it. No matter that around 40%† of the US' GDP is generated by industries which make use of copyright protection.
† https://www.uspto.gov/ip-policy/economic-research/intellectu...
As ever, I'll be a pedant and point out that "stealing copyrighted data" is not a thing.
More substantively, we don't know whether training is copyright infringement or not. The courts have yet to weigh in in any jurisdiction I'm aware of (i.e. the EU or the US).
As others pointed out, if LLM startups have to go broke after not being able to steal any more data then that will just be the reality of it.
I’m sure that if you somehow got access to ChatGPT weights and started selling them, OpenAI would be happy to call it stealing.