Training on copyrighted data should be legally allowed
- of course exact reproduction of protected content is a no-no
- but learning is ok, as long as it is transformative. User prompts and responses are pushing the model outside its training distribution anyway
- users add their own intent, making usage transformative
- when LLMs synthesize from multiple sources, the result is transformative
- if you try to protect expression it is meaningless now, but if you protect abstract ideas it kneecaps creativity
- the problems of copyright started with the apparition of internet, not with AI
- revenues from royalty are almost zero today, as each new content competes against an unbounded list of other works that have been accumulating for decades online
- because royalties are shit, creatives now focus on ads, and this leads to enshittification, attention grabbing junk everywhere, attention is scarce content is post-scarcity
- we actually like interactive participation more than passive consumption; we now edit Wikipedia, contribute to open source, have papers published for free on arXiv, use social networks where our comments are shared with the world, play games instead of reading books - it is another age, the interactive age
- AI is actually more than an infringement tool, it is useful for many legit purposes
- and AI is the worst possible infringement tool, it can hallucinate details, get thins wrong; By comparison copying is free and easy and precise to the letter
So the idea that training is infringement is pretty abusive, it tries to make copyright be about abstractions which is wrong. We can't return to 1990s, so we have to live with its demise. It's been dying for 3 decades already.