I'd like to see it happening but it sounds unrealistic.
I'd like to see it happening but it sounds unrealistic.
If GPT is blameless doing some things because it's a deterministic model, not an agent, then the "it would be okay if a person was taught like this" defence doesn't apply in other areas
More importantly, though, most judges are not philosophers.
That said, the outcome is unlikely - we have trained AI for more than a decade as ‘fair use’ at this point, it’s the application of the technology that is shifting the perspective, nor the act of training.
Every computer vision system in the world is trained on mostly public data for example.
Furthermore, the LLMs purpose is not to generate news so NYT will have to argue about the value of archive data. Many jurisdictions have thresholds of how much of an original work contributes to the derivative before it would be considered not fair use or plagiarism. Given the size of the datasets - good luck.
Fair use is about use. Spellchecking ML, search engine ML, etc. all different than ML that produces content.