AI companies required to disclose copyrighted training data under new bill
theverge.com
theverge.com
Gonna be a messy time when we catch up to that can down the road.
And to your point, we cannot copyright style, only the exact content itself.
Just look at YouTube copyright strike and understand how screwed up the enforcement of a law like this is going to be.
Tom Scott did a piece on it that’s worth watching:
They'll probably just lobby to change the laws, seeing this happen already.
http://gethisword.com/tech/exploringai/index.html
I also put a copyright proposal in there which should work for all the competing interests. I mean, people will still be greedy. Yet, I think it’s a win to allow all copyrighted works to be used for training while restricting only the outputs to only the degree we do for people.
Personally, I really hope the law allows GPT4-style AI’s to be built legally and affordable. If so, there’s so many benefits to be gained in businesses we can start, augmenting workers to improve work/life/productivity balance, and non-profit projects.
Plus, whichever countries don’t do it will be left behind in this area by the countries that make AI training easier. Japan is the only one I’ve heard made it legal to use copyrighted works for data-science models. I don’t know that you can distribute them like The Pile or Common Crawl, though. If you scraped or bought it yourself, Japan is the closest thing I know to being legal under copyright law.
Pressure to suppress accurate but "harmful" information about the government?
It'll be especially interesting with the ruling indicating that the output of these models isn't copyrightable as that and internally generated (from scratch) data will be the only bits omitted from this reporting. I'll be curious if the gap gets bridged on including the disclosure of data generated by a model containing copyrighted data.
Yes, I go as far as putting 99% the merit on the dataset, given a compute budget. Humans with similar cultural exposure are also remarkably close in intelligence, even though brains are very different at micro level.
If language data is actually the source of most of our acquired intelligence, if intelligence is a collective process, if it has an evolutionary drive, then isn't it silly to discuss so much about models or brains while forgetting about its crystallized form - language and text.
It is a repository of past experience spanning millennia, and a self replication medium for ideas. All our important knowledge is encoded in language. We (and more recently LLMs) draw heavily on this recorded experience. It cost a lot to be earned in the first place. If we lost all this knowledge, it would take us millennia to recover. It's smarter than all of us.
Think of people going through all of these private conversations they've had in ChatGPT. I mean we know that it isn't fully private but it can still be a bit delicate. Hopefully they anonymize it at least.
Everything that happened in 2022 and 2023 is over, all of those discussions are outdated already, it’s too late to try to make levels of working groups and information gathering.
Synthetic data has many advantages - it is free of copyright issues, the downstream models can't possibly violate copyright if they never saw the copyrighted works to begin with.
It is also more diverse and we can ensure higher average quality and less bias. It can also merge information across multiple sources. Sometimes we can filter using feedback from code execution, simulations, preference models or humans. If you can "execute" the LLM output and get a score, you're on to a self improving loop. LLMs can act as agents, collecting their own experiences and feedback.
I think GPTs are a ploy by OpenAI to collect synthetic data with human-in-the-loop and tools, to improve their datasets. This would also be in-domain for users and for LLM errors. They would contain LLM errors and the feedback. Very good data, on-policy. My estimations for 100M users at 10K tokens per month per user is 1T synthetic tokens per month. In a year they double the size of the GPT-4 training set. And we're paying and working for it.
But fortunately 12 months after they release GPT-5 we will recover 90% of its abilities in open source models.
I feel like we don't know if this is true or not. If we decide models trained on copyrighted data aren't fair game, it's possible we'll decide "laundered" data also isn't.
I mean, maybe that's not feasible. And I hope we don't decide training on copyrighted material is bogus anyway. But I don't think we know yet.
But also - you can totally violate copyright of something you never saw.
If we make the (poor, imo) decision to prevent training on copyrighted data, that's a restriction on the training process, not on its result.
And in the world where we're making bad decisions to put legal restrictions on the training process, "can't train on data obtained by models that were trained without these restrictions" seems on the table.
The problem artists have with LLMs is not that they routinely output copyrighted works (it can happen, but isn't common), but that they copy human-expressed knowledge and art styles on an industrial scale, thus destroying revenue streams that supported or encouraged creative expression in the past. I doubt this is going to be solved via existing copyright laws. I suspect that the need for disclosure of copyright on the source material is to build a case for more novel legal protections for artists down the line.
I was attempting to say that in my first sentence.
> they copy human-expressed knowledge and art styles on an industrial scale, thus destroying revenue streams that supported or encouraged creative expression in the past
I currently have no opinion on art works as my focus is on knowledge and LLMs. I just did a search on Amazon for "child care" and found 40,000 books on the subject. I doubt the world needs 40,000 books on child care so if LLMs remove the revenue stream from these authors maybe they will spend their time writing books the world needs.
There will always be pressure to merge, but there is also pressure to not do so. A single technology or business strategy will not disrupt this dynamic.
I wonder if there are provisions around synthetic materials generated from models trained on copyright data.
It doesn't even have a bill number assigned.