OpenAI, Microsoft sued for copyright infringement in class-action lawsuit
semafor.com
semafor.com
This is evidence of absolutely nothing as GPT cannot reflect over what is or is not in it’s training set. I’m shocked this is reported uncritically - here the impartiality of ‘just the facts’ distorts more than it enlightens.
The legality doesn't go away because I write software to train it's creativity. It's not copying. If you asked GPT to recite page 27, it couldn't.
Something like this:
https://www.reddit.com/r/ChatGPT/comments/17prrqe/who_needs_...
Have you read the whole lawsuit to determine that they did not dive deeper?
I can’t help but think the “proprietary” secret nature of the training data is just an attempt to sidestep the copyright issue. Like the “national security secret” line that the US government likes to use.
Crap training data = crap model.
Only licensed training data = base model not as good as open models.
Stolen training data plus RLHF to suppress this fact = great model!