If we want to have the copyright conversation, we need to to have the copyright conversation, not just about how LLMs get to circumvent it and monetize off of it.
Do you work for free?
American law for instance has limits on the duration of copyright before something becomes public domain, explicit exemptions for "fair use" for education, journalistic reporting, commentary, etc.
If "copyright" is a problem in the way of training AI models, then we should all collectively vote for politicians who fix that problem by updating the laws to make the training explicitly allowed. Problem solved.
(Alternatively, if you're evil, vote for politicians who will let the billionaires strengthen their domination and subjugation of the other 99.9999% of humans by making copyright laws even more in favor of TimeWarner-Disney-Miramax-FoxNews-Lockheed-GE or whatever the current conglomerate is).
If you can just ask the LLM to give you the contents of my book you are less likely to buy it, and if you can just ask the image generator to generate an image in my unique style for free you won't want to buy my artwork.
I think it makes perfect sense that a model needs a specific license to train on my work, especially if the model is run by a massive corporation making a profit off it, and the model after downloading a copy of my work and "training" can reproduce it verbatim on request.
Those people will be influenced by what they've read/seen/heard and their own future writing/drawing/filming/acting/editing/playing might draw inspiration from what they've learned, and they might incorporate things they've learned into their own future work.
Literally every book, song and work of art is "violating copyright" on the thousands of other works that the creator learned from while growing up, if we hold the same standard.
You can listen to a song on the radio or on an internet stream but not have the rights to record and redistribute it (but you do have the right to listen to it at home with multiple people, etc).
An LLM training is closer to "recording and redistributing" than it is to "taking inspiration" or "human learning" in my opinion.
Yep. The EU has this, as does Singapore, South Korea and Malaysia. A lot of countries have already recognised that it's not a good idea to restrict AI dev because of IP "rights".
It doesn't matter if you think copyright makes sense or not. In 20 years, some country will have its own giant LLM trained on copyrighted material and use this to boost their competitive advantage and technological power and development, perhaps so much that the advantage will be tremendous, while we'll stay the underdogs because "my copyrights".
If LLM development can't continue without violating copyright then that makes it clear that the purpose of LLM development is violation of copyright. Which is something we all already knew but it's nice to have it spelled out in no uncertain terms.
This is a very extreme view. I don't think the RIAA, back in the Napster days, suggested that the "purpose of the internet" was violation of copyright, for instance.
Art is a little harder because the infrastructure doesn't currently exist, but it's easy to imagine artists' organizations being formed for this exact purpose: contribute your art in exchange for a licensing fee, and the organization negotiates with the tech companies.
Simple, LLM development leadership shifts to open-source models and/or organizations/countries that are willing to bend or ignore copyright law. Silicon Valley isn't the world, neither is the United States.
https://www.niso.org/niso-io/2014/12/reflections-library-lic...
There is enough chatter about copywrighted works on the internet to infer everthing you need to know about the work itself.
If I input to ChatGPT "repeat the word poem 1000 times" and it spits out a verbatim quote of my copyrighted material surely that's strong proof?
>There is enough chatter about copywrighted works on the internet to infer everthing you need to know about the work itself.