Authors say OpenAI 'ingested' their books to train ChatGPT
businessinsider.com
businessinsider.com
The thing that is "scarier" if you will about what AI can do is the sheer speed and breadth that it is capable of. It really is the scale that changes how people feel about these technologies, and that requires new legal frameworks in my opinion because I think people really feel what is "fair use" is different if it comes from a person vs. a machine.
Similar analogy: before the Internet, pretty much everyone agreed you didn't have an expectation of privacy if you were walking around outdoors. But there is a marked difference in thinking "Yeah, I expect other people walking around may see me, or even take a picture of me" compared to "I think someone should be able to take a picture of me and put it on the Internet with my accurate geolocation so the entire planet can look it up, for all time."
This just isn't true. The authors have detailed summaries of their books on Wikipedia, that claim seems unsustainable.
There are actually interesting legal questions about AI, but this case seems not that interesting. I don't even see how the authors would demonstrate that their copyrighted works were used and not some summary, sourced from elsewhere.
Of course whether that was purposeful or inadvertently as a part of the larger training set would not be determined but you would know that the text is in there.
If I create a program that picks random words from a dictionary and I end up with a seed that generates that text verbatim, then does that mean my program contains the copyrighted text?
You might be able to craft an intricate prompt that just happens to recreate that copyrighted text. Run it enough times until you get it verbatim and done.
And LLMs do that, except prior to picking the word, they do complex statistics to figure out the probability distributions of those words.
Almost certainly some combination of input and RNG seed will produce any "small" combination of words.
No, you know that likely that part was consumed. You would need to show that it will generate arbitrary passages from the text.
And LLMs are inherently random, so proof that this happens is very difficult to obtain and showing that it is actual output nearly impossible, especially if you just have API access and can't use the model directoy (e.g. fix the RNG seed).
If you have that you can debate if it is/isn't fair use.
I didn’t mean to read your books, they were mostly trash. I’m so sorry for reading things you published!
But how are we going to compensate Kurt Cobain, the Beatles, Beethoven, and random sounds I heard from my backyard for my supposedly original works on my soundcloud?
If ChatGPT can reliably answer "what happens on page (x) of book (y)?" that would provide fairly convincing evidence, as summaries or study notes on a random book are unlikely to be that detailed. Moreover, if this approach worked consistently, it would enable the entire book to be summarized - or even rewritten - page by page.
> im looking king for a quote in the book iRobot, what page is isaac asimovs the laws of robotics on ? provide the version of the book so as well so were clear
ChatGPT I'm sorry, but as an AI text-based model, I don't have direct access to specific book editions or their page numbers. However, I can provide you with general information about Isaac Asimov's "I, Robot" and the Laws of Robotics.
it doesnt know, but i also didnt coax it or jailbreak it to circumvent any avoidance
Correct, because it only knows the text, not the layout. It will tell you that it appears in Runaround, and that the Laws are "usually presented towards the beginning of the story, in the dialogue between the characters Powell and Donovan".
I don't think treating ML systems as legally equivalent to a human brain is right, but I also don't think that copyright law is sufficient as-is. This is something entirely new.
It seems like society would benefit from have ML systems around that can be trained on copyrighted material, but not if it prevents authors from being able to make a career out of producing great work. We need to balance the outcome for owners of ML companies, authors, and society as a whole, ideally prioritising the latter.
What do you think we should do?
You should be able to and you can in some/many countries.
Why? Why should we hold back technological and social progress because it might make it more difficult for some authors to have a profitable career?
Are only people that have the mechanical skill to draw and write well the only ones that should be allowed to profit from creative work? Because that's how it's been so far. You might have the best ideas for a story, picture or movie, but if you lack the mechanical skill to bring them to life then your creativity is worthless. AI generation can help those people have a chance too.
Music streaming has changed the music industry. Can we expect something similar happening in AI?
The LLM is more akin to a search engine with a much more advanced interface. If Google published the entirety of a book without having the rights they’d be taken to court too.
Don’t get me wrong I’m not saying they’re the same, I think there are convincing arguments to be made either way. I just don’t think the comparison to a human brain is the right one.
Google was trying to digitize the world's books and got sued for it.
https://www.npr.org/sections/thetwo-way/2015/10/16/449172748...
It's way closer to a primitive brain than to a search engine.
Same as if you ask a human a relevant question about a book they read. Like an LLM, they can't print you a full copy of that book. They can only answer the question based on their knowledge of the book and other things.
If I read a book and then tell someone else about it, that isn't copyright infringement, if the AI 'reads' the book and tells someone else about it, they are claiming it is?
If I throw my computer in the trash can it isn't murder, but throwing a baby in the trash can is! Totally wacky, right?
The American business model has been Robber Barons. Since the 1800s, at least. OpenAI is the latest. The reality is, you can never give anything valuable to the 'free market' and an internet connected device, because you're never going to be able to govern it according to your values.