2,844 karma · joined October 23, 2019
In the real world, never once has the march of technology been successfully stopped, no matter how many jobs it crushed. Even the luddites who are enjoying a historical "re-evaluation" today still lost. And we've been through society spanning destruction of livelihoods before too and didn't stop it then either, just look at how much of the population used to be employed in farming and agriculture and compare it to now. Automation came and destroyed all of those livelihoods too.
In the real world, we can't even stop countries from pursuing, developing and improving nuclear weaponry despite the fact that we know it is an existential threat, and nuclear weaponry has no uses at all outside of nuclear annihilation. If we can't stop the proliferation of nuclear weapons, what makes you think we have even a snowball's chance in hell of stopping AI development when AI actually has non-world ending uses and benefits to people?
I'm happy to talk about the "real world" and in the real world, we're repeating the same mistakes of the past thinking "surely this time will be different" instead of learning from those mistakes and trying something new.
How so? His reasoning to me seemed consistent with decades of case law with respect to technology and copyright infringement. Certainly it was consistent with prior rulings finding time shifting to be legal (Tivo), or format shifting (VHS, audio cassettes and CD ripping and burning, also Google Scholar). It was consistent with rulings finding that copies made for commercial purposes, even when that purposed might have an impact on the market for the item being copied or incidental markets can still be fair use (Connectix Virtual Game Station). And it was certainly consistent with rulings finding that fair use can still open you up to extensive and possibly threatening levels of competition (the Java API and Google's use thereof in Android). Honestly, I'm not sure how anyone who has followed the general direction of copyright and IP rulings over the last few decades was surprised at all that Anthropic succeeded on "training an AI model is fair use".
> The assumption with a retro game decompilation is that you take the game apart and learn how it works. The journey is as important as the destination, but when you use AI for this you skip the journey and make the destination pointless.
That might be your goal with a decompilation. But some people might just want to play their old games on different hardware. I doubt the goal of Connectix was to "learn" how the playstation bios worked when they made their copies for the Virtual Game Station. And I'd wager good money that well over 90% of the people that use any of the retro emulation software and hardware out there have almost no interest in the "journey" of building the emulator and learning how the game hardware worked except to the extent that they need to know that information to play whatever game they want to place.
> There is a parallel set of concerns being voiced by artists, too: that the artistic process is as valuable if not moreso than the actual work product.
A similar argument was made during the early backlash (and indeed is sometimes still made in the modern incarnation of the backlash) against CGI and computer animation. Yet today, it's clear that CGI and computer animation is its entire, different but equally skilled, branch of artistic endeavor. Despite early attempts to treat it as "push a button and the computer does it for you" the end result has been simply higher levels of expectation and demand that has birthed its own branch of heretofore unknown skills. Likewise, plenty of skeptics thought synthesizers and drum machines took all the skill out of making music. And yet, the music industry and artists continue to generate plenty of art, even when using these labor and process "destroying" tools.
So the question becomes, why do you think LLMs and other AI tools are unique among technological advancements and will not be subjected to the same "opening new frontiers" effects that prior advancements have all had, despite their detractors?
[1]: https://walt-disney-animation-studios.fandom.com/wiki/The_Li...
[2]: https://jhmoviecollection.fandom.com/wiki/Encanto_(film)/Cre...
We aren't exactly consistent with how we approach this either. For example, I agree that reading "GPT-5 will autocomplete [GRRM's] series" feels "wrong" in the same way that if Terry Pratchett's "Disc World" series was "continued" by some non-approved author that would also feel wrong.
But contrast that with something like Star Wars, when George Lucas dies, I don't think anyone is going to feel any strong discomfort with some random person writing new Star Wars stories, even if those stories use the canonical characters. I suppose Disney might have a problem with it, but as a society, I don't see very many people losing sleep over someone not licensed by the Disney corporation writing more Star Wars. Likewise Star Trek. Gene Roddenberry is long gone, and while Paramount has ownership of the IP, if someone wrote their own Star Trek stories, no one is going to feel like they don't "comprehend the distinction between art and fan fiction".
And for further contrast, consider IPs that are well and truly part of the public domain. Cthulhu was the work of one author but since his death the lore has been expanded by multitudes of people, and no one finds that distasteful or tone deaf. No one thinks the "Hades" or "God of War" series of games don't qualify for "art" because the characters and lore being written about were the works of dead authors and the new material is certainly not "authorized" by those authors or their descendants.
Obviously time and distance plays a part of this, but as a more contemporary example, I wonder how many people would lose sleep or feel any significant discomfort over unauthorized or even AI generated Harry Potter works, either now or post JK Rowling's death. At the very least, if this quote read that Gogineni would "rest easy knowing that even though JK Rowling has [lost the plot/is an awful person/pick your reason for disliking her or her later work], GPT-5 will autocomplete her series." would that make people as equally uncomfortable? In my estimation, I would guess it would fall somewhere between the GRRM version and something like new Star Wars material for most people.
Further, I'm not sure how they've "captured the commons". By definition the "commons" belongs to us all. Nothing prevents someone else from doing the same thing. That is, unless the Authors Guild succeeds in splitting the courts over the fair use of AI training and the resolution of that split finds that training isn't fair use. Then the commons can only be used by companies or people with pockets deep enough to license the material in perpetuity.
I think there's a pretty good argument to be made that every single robot is built on the labor, creative and technical knowledge and advancements of the workers that robot replaced. Robots after all, much more than LLMs, are incapable of creative output. Some human (probably a laborer) figured out how to stamp the steel in just the right ways, or how to cut the patterns in just the right ways so that the product could be manufactured. Then a robot company came in and stole that creative output, or more likely was sold that creative output by the company owners who stole/bought (depending on your point of view about labor and the ownership of labor's creative outputs in the current US legal system) to produce a robot that then displaced the laborer who created the process in the first place.
This order grants summary judgment for Anthropic that the training use was a fair use.
And, it grants that the print-to-digital format change was a fair use for a different reason. But it
denies summary judgment for Anthropic that the pirated library copies must be treated as
training copies.
The document you linked was written a month before the Bartz decision was reached. It's also worth noting even the document you linked says this in its conclusion: Various uses of copyrighted works in AI training are likely to be transformative. The
extent to which they are fair, however, will depend on what works were used, from what
source, for what purpose, and with what controls on the outputs—all of which can affect the
market.
[1]: https://copyrightalliance.org/wp-content/uploads/2025/06/Bar...Without copyright, the GPL would be unnecessary.
By comparison, the model isn't "watching" the show, as so many people are quick to point out that the "learning" analogy for what AIs are doing is flawed. There was never an intent by the creators that the show would be used to generate mathematical probabilities and weights in a statistical model and no one is deriving entertainment from making the statistical model. I suppose perhaps someone derives entertainment from AI training, but I suspect the number is small enough that "no one" is a reasonable approximation. So using the show to do so at least has an argument towards fair use. Or if the copy used for training was legally purchased, at least in the US it has the actual legal designation as fair use so far.
Don't get me wrong, I'm not saying that we should be returning to the days of the RIAA suing teenagers for their college education funds. But it seems pretty obvious that "pirating copyrighted material to explicitly use that material in the way that the creators of the material envisioned selling to you" is similar to, but arguably worse than "using copyrighted material (pirated or not) in a way not envisioned by the creator of that material to create a wholly different product". In both cases, the livelihood of the creator is possibly being affected, but one of them is a direct 1 for 1 loss of income while the other (again, if not specifically pirated) is an indirect impact.
This is a timeline that can not and will never exist if the Authors Guild and most of the other anti-AI lawsuits succeed. Because if they do succeed, the only people who will be able to provide an AI model will be companies with enough resources to license all the training data in perpetuity. The Authors Guild isn't mad because the authors can't train and use their own AIs, they're mad because the AI companies are making money and they're not getting what they perceive as their fair cut of that money.
> Imagine OpenAI, Anthropic &co having to compete by hiring [thousands] of their own talent to help training their commercial models.
If that's what the AI companies would need to do, how then does "each creative" compete? What authors or artists do you know that can afford to hire "thousands" of people to help them build bespoke AI models?
It seems to me that we should be figuring out how to make public models and datasets that can be used by anyone, not further strengthening copyright so that AI models can only be produced by companies with the resources to hire thousands of people.
There's no real reason to think that AI is unique among all the livelihood destroying technologies that have come before in that it will destroy jobs without opening new ones in their place. It seems to me fighting the march of technology is akin to fighting the tides. Yes, you can build sea walls, and in the modern age, we literally can hold back the tides. But at the cost of ever increasing expenditures of time, resources and/or money. We could spend our resources trying to preserve livelihoods whose time has come and gone, but why specifically these jobs and not say, coal miners, farriers, lamp lighters, whalers, operators, computers, thread spinners or any of the multitudes of other jobs that technology is obsoleted? If we think on all the jobs that have been lost to technology over the decades, how many of those jobs do we think would have left the world in a net better place had we eliminated the technology that obsoleted the job and kept all the people employed in those positions where they were. Are people scared for their livelihoods because the technology will destroy the job, or because societies tend to spend all their energy trying to fight the technology instead of helping people find new jobs and livelihoods and the technology always eventually wins?
> ...
> If the model can encode the copyrighted material with sufficient fidelity to reproduce them on command, it stops being fair use or should anyway.
That may be so, IF you could actually do that. Yet in over 200 individual allegations in the Authors Guild vs. Open AI case complaint[1], not a single one of them alleges that you are able to do this. They allege that you could at one point get detailed verbatim quotations, but also note that the models have been explicitly blocked from doing this. Instead, the vast majority of the actual complaints in the case are about generating "summaries" that contain information not in other publicly available summaries, and generating "detailed outlines" of supposed future installments of the copyrighted works, using the characters and details of the story. In other words, all the actual alleged infringements are about either:
A) the copies made in order to train the model
B) infringing derivative works generated by prompting the AI
C) the verbatim copies made from sources to which OpenAI did not have rights to
You would think if any of the authors at all had been able to print out verbatim copies of their works without having to take knowing and direct action to circumvent the blocks in place to prevent that from happening, those would have been some of the top complaints in the case. The same held true for Bartz vs. Anthropic, where the judge even noted in his ruling that while generating verbatim copies might indeed be infringement, the plaintiffs never alleged that had happened or was possible.
And since Bartz vs. Anthropic has (reasonably IMO) found the training to be sufficiently transformative as to be fair use, the complaints for point A are unlikely to succeed here. The complaints for C almost certainly will succeed, for the same reasons they succeeded against Anthropic.
That leaves B. The questions would be:
1) Are such "detailed" summaries infringing just because they can include things other summaries have yet to include? Personally, I doubt they're going to get much traction here unless the courts split and they win on point A. The fact that other summarizers have left certain details out does not inherently make a new summary with other details an infringing work. If we imagine a world where Empire Strikes Back is a new movie, if none of the public reviews of the movie reveal the twist, but you can ask an AI model to summarize the movie and the AI model reveals the twist, that might be disappointing, but I don't think there's any argument to be made that it is copyright infringement.
2) Are speculative outlines of future unpublished work based on the information in a published work infringing just because they have been created?Is a model that CAN be used by a user to intentionally create an infringing derivative work itself an infringing product? Again without splitting the courts and winning on "training is infringing therefore all outputs are also infringing" I just don't see how they can win here. A speculative outline of future works is something people have been doing forever (see also any fan site on the internet). While attempting to publish that outline commercially or produce a new work from that outline might itself be infringement, that infringement is the result of explicit and knowing actions of the user akin to putting a book on a xerox machine and producing a cut and paste fan edit from the work. Again the xerox machine is not itself the infringement, and the individual page copies probably are also not infringement until they are used in a specifically infringing way.
3) Does the ability of the model to theoretically produce verbatim copies of the training material if OpenAI were to re-program the model to remove the blocks they have put in place to do that mean the models are themselves infringing. This is perhaps the most "up in the air" question of the 3, but the law generally doesn't award damages on the potential for copyright infringement, only on actual acts of infringement. Handbrake and various DVD copying tools do not ship with the keys necessary to defeat the DVD protection schemes, yet they know how to use those keys and accept such keys provided by the users. As far as I know, no court cases have been brought or succeeded against any distributors of DVD ripping software despite the fact that evading the "anti-infringement" blocks in the software is both trivial and exposed to the end user. Given that evading the "anti-infringement" blocks of OpenAI's models is neither trivial nor exposed to the end user, I'm fairly comfortable saying that again without splitting the courts and winning on point A, the authors guild isn't likely to win here either.
[1]: https://authorsguild.org/app/uploads/2023/12/Authors-Guild-O...
Strong IP advocates have argued for years that devices that can be used to infringe copyright are themselves infringement of copyright. So far that hasn’t held up to court analysis provided that device can be and is also used for non-infringing purposes. Given that so far the courts have found that training an AI model is sufficiently transformative to qualify as fair use, it doesn’t seem likely that distributing a model counts as distributing copyrighted material.
Is it not Privacy & Security -> Apple Advertising -> Personalized Ads?
And of course the referenced settings despite being under the Screen Time top level settings are still within a sub-section there called specifically "Content & Privacy". So the quick answer is obviously not the correct answer.
Should your computer automatically build search indexes for your files? For most people, fast search is more useful than the disk space this takes up, but it would be doing something you didn't explicitly invoke.
Should your computer turn on global spell check (and even basic auto correct) for text fields? Again, you didn't invoke it explicitly, but most people probably think it's a good idea.
So you might say "well anything that just runs locally is ok" but even that's thorny. Obviously much of the AI stuff apple has been implementing is running locally and some of it has been around for a while (image classification in Photos for example). But not everyone wants the AI, even with local models.
And then there's things like "should your computer default to using NTP time?", again most people would probably say yes, especially given how important accurate time is for modern computer security. But it requires your computer doing network requests to make it work.
"Should your computer automatically check for updates?", again for most people the answer is probably "yes it should". But again it requires your computer to do network requests to make it all work.
"Should your computer automatically scan for malware?", probably yes it should, but keeping those malware lists up to date is going to require network requests and access.
There's always a tension between "being helpful and useful out of the box" and "doing stuff the OS vendor wants".
Anyone who wants an OS that does nothing out of the box is always free to grab any x86 box and install ArchLinux on it, but most people don't do that for the same reason that Spacemacs / Doom Emacs is more popular than a raw emacs install, and things like Sublime, or Notepad++ are more popular than those. Defaults matter a lot.
Probably not what you had in mind though...