> The model doesn't know anything - people personify LLMs too much.
Perhaps reading people's posts little less literal would help the conversation. I obviously know that the model weights don't 'know' something in the sense that a human knows something, but the model does store information. That information happens to (primarily?) be the statistical likelihood of one token following another token. What it doesn't store is a string of tokens that represents Sarah Silverman's latest book.
From what I can tell, all this angst comes from 3 or 4 related, but different, issues.
1. Did companies break copyright laws when assembling and using training for these models?
2. Does the model represent some form of copyright infringement in and of itself?
3. Does a model's ability to output chunks of copyrighted work have some implication of the legality of the model itself? (using said copyrighted chunks is already a solved issue)
4. Do we as a society owe it to humans benefitting from copyright the continued ability to create copyrightable content without competition from ML models?
I think comingling all of those points is doing everyone a disservice.
My assertion was only about #2 and none of the others. I feel like it's a clearly demonstrable situation that these models _don't_ infringe copyright directly. That being said, I am obviously not a lawyer, and my opinion is just that, an opinion.
FWIW, my general feeling on all of the points is: 1. Quite Likely (but fair-use is a fickle thing), 2. No, 3. It shouldn't, and 4. No, but we need to think through the long-term societal implications of ML decreasing the amount of human labor needed across all markets and come up with a plan that doesn't involve our fingers in our ears.