If a software service had legal protections like that, sure, I could build one that returns you any book you request and say that the service had integrated it into its worldview. Who can check, eh?
* Actually, in some countries you could be in trouble for reading a book and incorporating it into your worldview, to say nothing about quoting it, but let’s set that aside.
Not a relevant factor when it comes to copyright law. Fair use (the law that's most applicable here) applies regardless if you're a student using incorporating news articles into your work, or google making thumbnails and displaying them on their search results.
Furthermore:
> Examples of fair use in United States copyright law include commentary, search engines, criticism, parody, news reporting, research, and scholarship.
I do not see “automated generation of derivative works of arbitrary nature” in it.
The “automated” isn’t really key. If you read a book, and learn from it, and are able to use that knowledge in other contexts, should you pay a licensing fee? It doesn’t matter if “you” is a human or machine.
The point isn't that AI training is legal because it's like generating thumbnails. That is being argued in the courts right now. The point is that fair use exemptions isn't limited to "being a conscious human being enjoying human rights", as google generating thumnails and snippets using computers shows.
https://en.wikipedia.org/wiki/Perfect_10,_Inc._v._Amazon.com....
> Examples of fair use in United States copyright law include commentary, search engines, criticism, parody, news reporting, research, and scholarship.
Those are examples, not an exhaustive list. It's not even something that Judges are supposed to compare against when deciding whether something is fair use or not, see: https://en.wikipedia.org/wiki/Fair_use#U.S._fair_use_factors
Sure. However, my point is that this is not fair use*, so other principles need to be applied. Whether legal systems in various countries find that fair use applies here or not, I agree we are yet to see.
* At least in cases where it’s an LLM operated at scale for profit (which I suppose would not hold for Meta’s models if they were truly open, but that’s not the case if they require obtaining a license in some conditions).
This isn't a complete argument. Most of AI companies' argument relies on the fact that AI models are "transformative". That's a plausible claim, and as Perfect 10 v. Google, and Authors Guild, Inc. v. Google, Inc. has shown, being a for-profit company is hardly a disqualification from getting fair protection.
But sure, the “transformative” argument is the one that could apply (and even I believe Google used it to argue its case), if it can be shown that an LLM can not verbatim reproduce a given work (which, incidentally, is something that you, a warm-blooded fleshy human with agency who has the freedom to read books, cannot do, but LLMs were shown to do).
That said, relevant laws existed before LLMs, and may are outdated. If the goal is to balance reasonable uses while protecting original output of authors that ultimately drives innovation and creativity, I am not sure if the preexisting laws are continuing to fulfil their function, but that’s my opinion.
You have to try pretty hard to get LLMs to reproduce a work verbatim, especially any lengthy passages that aren't famous (and thus re-quoted on the internet a bazillion times). Moreover just because LLMs can reproduce a work verbatim if you try hard enough doesn't mean it's not transformative. Google search snippets and google book search has been ruled "transformative" by the courts, but if you tried hard enough you can use them to extract the entire work.
>That said, relevant laws existed before LLMs, and may are outdater. If the goal is to balance reasonable uses while protecting original output of authors that ultimately drives innovation and creativity, I am not sure if the preexisting laws are continuing to fulfil their function, but that’s my opinion.
AFAIK the era of mining the public internet or published works for AI training data is over, or at least coming to an end. Everything that could be mined, has already been mined, and besides, the internet is getting increasingly polluted by AI output. Private training data is where it's at now, whether it's sourcing document troves from companies (eg. emails, documentation, source code, etc.), or paying "AI annotators" to produce training data for you. If the argument is that human authors should get a cut of AI profits because their works were "stolen" to train the models, this is going to be a increasingly losing argument, because it doesn't have a leg to stand on for private training data.
The argument can be made that LLMs could not be created without expropriating the original works of all the authors they were trained on, and that argument would in fact be true and have quite sturdy legs as far as I’m concerned.
It’s not a historical instance of forgotten times, it started less than half a decade ago and I would be surprised if it’s not still ongoing (your argument about synthetic training data is forward-looking).
That makes as much sense as "American industry was built on the backs of British inventors (back it the day it was the "China" when it came to IP), so Britain should get perpetual (?) royalties from the US economy".
You are arguing that doing something that is legal if being done by humans is not ok if it is done on computers running an LLM.
I see no difference with cryptotokens here, the human has freedoms to do things and the human is responsible for them if those things are bad. (Just unlike LLMs, theft of property and all that is kinda always a crime, unlike reading a book in a shop without buying.)
Now, would that be a fair use of the books?
Ok, it doesn't tell much about AI and fair use, but I find it funny that your thought experiment is actually something you can do in real life.
There's no reason for Harry Potter for example being 10000 times more valuable than a book on quantum mechanics only because the former is more popular and the latter is on a more obscure topic.
You have bought the text so you have the readright, but you do not the copyright.
You do however, have the right to make derivative works based on the contents of the book. You reading a physics textbook doesn't mean you can't write a blog post about gravity or whatever, and you reading harry potter doesn't mean you can't write a series of fantasy books involving a young wizard trying to fight an evil wizard.
> The application was denied because, based on the applicant’s representations in the application, the examiner found that the work contained no human authorship. After a series of administrative appeals, the Office’s Review Board issued a final determination affirming that the work could not be registered because it was made “without any creative contribution from a human actor.”
That just means whatever they produce can't be copyrighted, not that they can't produce derivative works. Courts have upheld the right for google to produce thumbnails of copyrighted works, even though the procedure for producing thumbnails is done by a computer and thus can't be copyrighted.
Maybe we will end up agreeing that we just want to stick with those same laws for machine consumption and creativity. But maybe we won't since they are quite different things.
That... doesn't make it okay...
> A lot of them I didn't even pay for, I borrowed them from libraries or friends.
This 2nd sentence doesn't fit your first. What is your message?
So I expect to see that either you are no longer allowed to own computer software
Or a return of slavery.
Also if we find indecent portrayal of minors in a data centre I expect that we treat it as a strict liability crime and the entire data centre or corporation that owns it gets a long prison sentence, just like a human would. However that is suppose to work.
Could you clearly speak your point?
Those people who conflate them deserve it. You and me don't.
> there's really not much you, I or anyone else can do
We can make our own community. And Bluesky is very much not it.
I agree with you until this part. There comes a time where I don't think I deserve to get my eyes poked out just because other people find that fashionable.
Extrapolated out into some new future a hundred years from now when we have embodied AI humanoids walking alongside us, would it be weird if those humanoids were barred from buying a new book or charged a different rate than the humans they coexist with?
I’m still deciding how I feel about some of this too.
I'm not even against this to a point. The issue is what comes after. The monetization. The enshitification. The derivatives in place of real creativity.
The only way to prevent the things you are worried about is to let anyone train a model, and then compete to make the best product.
The enshittifying monopolies and big copyright holders are the only ones that would benefit from locking down training data or regulating AI compute.
They're not being charged, that would be a vast improvement over reality.
But if they had done that, I bet they would have been sued anyway.
“Because you wouldn’t have sold it to me. Or even if you would have, you would have put such onerous terms on it that I’d rather take this path.”
(Not saying this makes it right)
So just buying copies wouldn’t have helped them.
For a more direct counterexample, I can memorize something and type it back out, but if it is copyrighted the law doesn’t make an exception just because it passed through my head.
I do agree that we should encourage human creativity. But if AI isn't making copies, and the output of AI isn't awarded copyright (as is currently the case) then I think humans still have sufficient reward.
There will be a lot to figure out over the coming years.
On the other hand, they torrented books and then open sourced LLM weights. No punishment is too severe for that!
If you still don’t understand, I strongly suggest watching Max Headroom, “Lessons”, which you can get here: