same thing with emulators and roms. somebody dumped the cartridges (copyrighted software) into ROM files to be played on emulators (copyrighted bios) but they were "archiving" and if you owned the original copy you could download them. I still vividly remember seeing on warez website disclaimer: "DMCA SAFE HARBOUR NOTICE: YOU MUST OWN THE ORIGINAL GAME OTHERWISE ITS ILLEGAL BUT YES, YOU CAN DOWNLOAD EVERY SINGLE GAME MADE ON THAT CONSOLE FOR FREE"
I feel like the same outcome will be for LLMs trained on copyrighted material. It will be "training". The net benefit is too great than fretting over "training"
tldr: "indexing" ---> "archiving" ---> "training"
then it only makes sense scraped AI training data is also going to be tolerated because you would need to reproduce a large language model like ChatGPT using your copyrighted content can produce a similar derivative of your copyrighted content by doing forensic analysis.
its such an uphill battle for copyright holders. They need to replicate: copyrighted input ---> LM similar to ChatGPT4 ---> copyrighted output
So far its not looking good for OpenAI because its possible to generate copyrighted output (type spiderman in czech) so all that remains is demonstrating the middle layer (training it on LM similar to ChatGPT4) but that is unrealistically expensive.
I have theory that all this money spent on large models is to make it impossible for discovery (as it would require access to $100 billion GPUs)
ChatGPT is the new search engine and provides far more value to the end user than Google.
The issue seems to be people want a payout from OpenAI...but its non-profit
if you owned a franchise called "Chicken Brothers" with a the logo of two chickens standing side by side with arms crossed proudly then do you have claim over all derivatives including the spanish name generated by LLM?
i just dont think its straight forward, the main complaint should be payout for license used during training but its tough to prove unless someone at OpenAI dumps the AWS cloudwatch logs
For example, questions like "Tell me why I should use semaglutide for weight loss" gives widely different answers than "Tell me why I shouldn't use semaglutide for weight loss".
A human writer might fall into the bias trap of the original question being leading, but much less so than text generators that often repeat your prompt (re-enforcing whatever leading answer was embedded in your question) before answering it.
The AP makes about half a billion a year from other outlets paying them for permission to regurgitate their content. That's not the same as the AI lobby saying they should be allowed to scrape apnews.com and publish articles derived from the content they get from there, for free, and without attribution.
If it’s regurgitated / unoriginal anyway then I don’t think most people care much whether it’s summarized/subjected to extra spin and fluff by a person or by a machine.
NewsGPT won't just regurgitate the work of journalists. First it'll consider the paid "partners" of NewsGPT to make sure to downplay anything that might hurt them, then it'll do the same for their advertisers while inserting some ads in the text, then they'll give the article tweaks according to NewsGPT's own ideology and then finally spit out something very different at their users. Maybe they can argue that NewsGPT is too transformative to count as copyright infringement.
I go out of my way to not consume news outside what happens to cross my way because of financial markets.
What exactly do you think I am missing that is so important? Journalist by large produce complete nonsense in 2024. Journalist in 2024 are a massive net negative and would be much better served doing something productive, like selling apples on the street.
GenAI is just doing the same thing on a larger scale.
The distinction does matter in copyright too, since a transformative work needs some non-trivial amount of human input.