Why are you turning him into perpetuum mobile in his grave?
Do you really believe Aaron would be arguing against AI companies and for publishing / recording guilds on the grounds of intellectual property claims?
No, it's the tech community that did a sudden about-face, and is now all "friendship ended with free access to information and technologies enabling people; now RIAA is my best friend", and this move is as dumb as that meme (https://imgflip.com/memegenerator/137501417/Friendship-ended).
That's not the point. All the rules and laws are enforced when its you and me but when it's big tech the laws are treated by these companies as mere instructions.
> friendship ended with free access to information and technologies enabling people; now RIAA is my best friend
Big tech will enable access to free information and will help people reach new heights. Do you see how wrong that sounds?
A: "Information is free"
B: "Information is not free"
C: "Information is free only for the rich and not free for everyone else, giving the rich a material advantage over everyone else that not only entrenches but accelerates wealth inequality and impedes class mobility"
You, or Swartz, are an advocate for A. Why, exactly, do you think that obliges you/Swartz to prefer C over B while A is not true?
We're literally talking in a thread about someone who committed suicide because the US government was hellbent on ruining his life with a felony conviction for piracy.
As we can see in the good article, the copyright system might almost have cost us a lot of AI capabilities. The damage it has done in cases where the lawyers got ahead of the builders is incalculable.
1. Unauthorised network access
2. Intent to distribute licensed material
The labs aren’t doing this. They are doing something similar to you and I downloading torrents. Look, I also think laws should apply somewhat equally to individuals and companies. But this is different.
According to the people trying to ruin his life. This is an accused thought crime, not something he actually did. Also, supposing it actually was something he did, your argument is that building a trillion dollar business on stolen licensed material is legally permissible but giving it away for free is worthy of your life being ruined. Wonderfully coherent world view, that is. Piracy is fine, but only if you hoard it to yourself and profit from it!
That doesn't seem different to me. Copyright, inherits.
If its fair use, like the companies currently claim, maybe it is different. But I suspect that the insane push to create a copyright carveout in countries around the world, is not unrelated to a judgement that it probably isn't fair use.
Song lyrics for example
It would be like if I sold you a pizza, but when you opened the box it contained all the source code for the latest GTA. The pizza wasn't copyrighted by anyone, but the what was inside the box was.
Strong IP advocates have argued for years that devices that can be used to infringe copyright are themselves infringement of copyright. So far that hasn’t held up to court analysis provided that device can be and is also used for non-infringing purposes. Given that so far the courts have found that training an AI model is sufficiently transformative to qualify as fair use, it doesn’t seem likely that distributing a model counts as distributing copyrighted material.
>training an AI model is sufficiently transformative to qualify as fair use
This is the key question and that courts have decided this way so far doesn't mean its the correct decision. If the model can encode the copyrighted material with sufficient fidelity to reproduce them on command, it stops being fair use or should anyway.
> ...
> If the model can encode the copyrighted material with sufficient fidelity to reproduce them on command, it stops being fair use or should anyway.
That may be so, IF you could actually do that. Yet in over 200 individual allegations in the Authors Guild vs. Open AI case complaint[1], not a single one of them alleges that you are able to do this. They allege that you could at one point get detailed verbatim quotations, but also note that the models have been explicitly blocked from doing this. Instead, the vast majority of the actual complaints in the case are about generating "summaries" that contain information not in other publicly available summaries, and generating "detailed outlines" of supposed future installments of the copyrighted works, using the characters and details of the story. In other words, all the actual alleged infringements are about either:
A) the copies made in order to train the model
B) infringing derivative works generated by prompting the AI
C) the verbatim copies made from sources to which OpenAI did not have rights to
You would think if any of the authors at all had been able to print out verbatim copies of their works without having to take knowing and direct action to circumvent the blocks in place to prevent that from happening, those would have been some of the top complaints in the case. The same held true for Bartz vs. Anthropic, where the judge even noted in his ruling that while generating verbatim copies might indeed be infringement, the plaintiffs never alleged that had happened or was possible.
And since Bartz vs. Anthropic has (reasonably IMO) found the training to be sufficiently transformative as to be fair use, the complaints for point A are unlikely to succeed here. The complaints for C almost certainly will succeed, for the same reasons they succeeded against Anthropic.
That leaves B. The questions would be:
1) Are such "detailed" summaries infringing just because they can include things other summaries have yet to include? Personally, I doubt they're going to get much traction here unless the courts split and they win on point A. The fact that other summarizers have left certain details out does not inherently make a new summary with other details an infringing work. If we imagine a world where Empire Strikes Back is a new movie, if none of the public reviews of the movie reveal the twist, but you can ask an AI model to summarize the movie and the AI model reveals the twist, that might be disappointing, but I don't think there's any argument to be made that it is copyright infringement.
2) Are speculative outlines of future unpublished work based on the information in a published work infringing just because they have been created?Is a model that CAN be used by a user to intentionally create an infringing derivative work itself an infringing product? Again without splitting the courts and winning on "training is infringing therefore all outputs are also infringing" I just don't see how they can win here. A speculative outline of future works is something people have been doing forever (see also any fan site on the internet). While attempting to publish that outline commercially or produce a new work from that outline might itself be infringement, that infringement is the result of explicit and knowing actions of the user akin to putting a book on a xerox machine and producing a cut and paste fan edit from the work. Again the xerox machine is not itself the infringement, and the individual page copies probably are also not infringement until they are used in a specifically infringing way.
3) Does the ability of the model to theoretically produce verbatim copies of the training material if OpenAI were to re-program the model to remove the blocks they have put in place to do that mean the models are themselves infringing. This is perhaps the most "up in the air" question of the 3, but the law generally doesn't award damages on the potential for copyright infringement, only on actual acts of infringement. Handbrake and various DVD copying tools do not ship with the keys necessary to defeat the DVD protection schemes, yet they know how to use those keys and accept such keys provided by the users. As far as I know, no court cases have been brought or succeeded against any distributors of DVD ripping software despite the fact that evading the "anti-infringement" blocks in the software is both trivial and exposed to the end user. Given that evading the "anti-infringement" blocks of OpenAI's models is neither trivial nor exposed to the end user, I'm fairly comfortable saying that again without splitting the courts and winning on point A, the authors guild isn't likely to win here either.
[1]: https://authorsguild.org/app/uploads/2023/12/Authors-Guild-O...
By comparison, the model isn't "watching" the show, as so many people are quick to point out that the "learning" analogy for what AIs are doing is flawed. There was never an intent by the creators that the show would be used to generate mathematical probabilities and weights in a statistical model and no one is deriving entertainment from making the statistical model. I suppose perhaps someone derives entertainment from AI training, but I suspect the number is small enough that "no one" is a reasonable approximation. So using the show to do so at least has an argument towards fair use. Or if the copy used for training was legally purchased, at least in the US it has the actual legal designation as fair use so far.
Don't get me wrong, I'm not saying that we should be returning to the days of the RIAA suing teenagers for their college education funds. But it seems pretty obvious that "pirating copyrighted material to explicitly use that material in the way that the creators of the material envisioned selling to you" is similar to, but arguably worse than "using copyrighted material (pirated or not) in a way not envisioned by the creator of that material to create a wholly different product". In both cases, the livelihood of the creator is possibly being affected, but one of them is a direct 1 for 1 loss of income while the other (again, if not specifically pirated) is an indirect impact.
The idea being (before the rise of online peer-to-peer piracy) to prosecute the people making bootleg VHSes rather than the people buying them.
With the rise of these AI behemoths, it seems that rule is now inverted: You can download all the pirated ebooks you want, as long as it's for large-scale for-profit commercial use.
It's literally the opposite. Anthropic paid a $1.5 billion settlement. Litigation against OpenAi is still ongoing. Meanwhile, no one has ever been punished just for consuming pirated media.
It would be shocking and outrageous if this was a case of this being 'the cost of doing business' for the big guy and a life-ending judgement for the little guy. Luckily our justice system is clearly allowing individual citizens the pleasure and honor of eating cake.
Even if AI companies were literally doing the exact same thing as Aaron Swartz but not getting punished for it, that still doesn't make Swartz's punishment retroactively their fault. If you have a problem with powerful people being powerful, then I suggest you direct your complaints to the heavens.
Fun fact: Kim Dotcom is still fighting extradition while these drama queens (I.e Dario) are lecturing us about how much access the peasants should get to AI models fed and trained with stolen IP.
I strongly believe Aaron would oppose the appropriation of content. The problem with AI (in this context) is not that the AI companies gain access to information that regular people can't freely access. The problem is that AI erases the information about who originally created a piece of work.
When people want to freely share their work, then they usually reach for the Creative Commons licenses and not for Public Domain, because the latter doesn't protect authorship.
Imagine OpenAI, Anthropic &co having to compete by hiring [thousands] of their own talent to help training their commercial models.
Another thing is attribution. Even a book that was re-published illegally can be easily attributed to its author. What happened is the exact opposite: no attribution, not even a notice, just obfuscation that strips away any traces of the original ideas and original work. Imagine piracy websites and trackers just dropping first few pages that name their authors, and publishing "the book you are looking for". This is exactly what happened.
I like the vision, but the model would need more than just the author's work. The situation you are imagining would require models like we have as a base.
This is a timeline that can not and will never exist if the Authors Guild and most of the other anti-AI lawsuits succeed. Because if they do succeed, the only people who will be able to provide an AI model will be companies with enough resources to license all the training data in perpetuity. The Authors Guild isn't mad because the authors can't train and use their own AIs, they're mad because the AI companies are making money and they're not getting what they perceive as their fair cut of that money.
> Imagine OpenAI, Anthropic &co having to compete by hiring [thousands] of their own talent to help training their commercial models.
If that's what the AI companies would need to do, how then does "each creative" compete? What authors or artists do you know that can afford to hire "thousands" of people to help them build bespoke AI models?
It seems to me that we should be figuring out how to make public models and datasets that can be used by anyone, not further strengthening copyright so that AI models can only be produced by companies with the resources to hire thousands of people.
What big AI companies have done is illegally hoovering up copyrighted creative output of individuals and creating a situation where the wages that normally would be paid to those individuals instead go to that one company (that stole their work) which now becomes disproportionally rich and powerful.
In both cases companies obtain money and power by hoarding information obtained through dubious means (in the former case most academics willingly participate while at the same tone they often don't really have a choice). Exactly what Swartz was fighting against.
I think what Swartz did was moral, and his prosecution was unjust.
I think what OpenAI did was moral, and them getting sued for it is unjust.
Why do you have one position for Swartz and a different one for OpenAI?
(Aaron Swartz was a mailing-list friend of mine, so I do have some bias here. But in part we knew each other because our moral position on this was similar)
I think many people on HN dont mind OpenAI use of copyrighted material, but do not like how they try to do regulatory capture of a market and try to say their own copyright is now somehow more important.
Like how OpenAI trying to make "distillation" illegal while it exactly what they did with whole intetnet, books, everything.
But it is actually possible to agree with one thing a company does and disagree with others.
Does (b) make things more just, as certain possible unjust acts don't happen?
Or does it compound the injustice by creating a double-standard, and perpetuate it by hiding the problem from any with the power to bring an end to it?
On the one hand, there is a notion of consistent principles regarding the legal handling of the topic, which is being appealed to by some people such as yourself.
The other side seems to be drawing attention to the fact that these principles are not applied consistently by society. They're questioning the moral validity of holding to principle in a circumstance where it's guaranteed to be applied with very specific biases that are rarely explicitly stated.
I've noticed this talking-past happening in other subjects too. For example American drug laws. There's the principle of drug laws, and then there's the practice of which kinds of people gets the laws applied to them. One group of people focus on the one, and another focus on the other, and they just kind of talk past each other.
[1] broadly the same laws: I do understand one was criminal and one is civil and yes I agree that injustice. To me neither should be criminal.
"Why do you have one position for an activist and another for a eight-hundred and fifty-two billion dollar, for-profit corporation?"
"Why to you have one position for someone who wanted to grow the intellectual commons and another for a corporation trying to enclose it?"
"Why do you have one position for someone who gave his work away for free and another for a company that charges for access to proprietary tech?"
What’s the motive? For openAI its profit.
I stand by my moral position.
Then, who is "we" here giving the appearance of a consensus opinion here and in mainstream? It's a vocal minority, it's the powerful, it's the causes they fund and put resources behind to continue to preserve their causes. And now today, it's astroturfing, fake AI-LLM-bots almost indistinguishable from you and I. Don't mistake artificial consensus for reality.
> Why is big tech getting away with so much more?
IMO Because the majority of people are passive, standing by, tolerating abuse and trickery by the minority. This is a perpetual cycle in humanity: those minority use their power and leverage and abuse their positions until they are ousted. We have tolerated this because we haven't stood up yet and acted to change things and demand equal enforcement of the laws that appear to apply to us but not to them. If history says anything, they are afraid and panicking and will continue to be more abusive until they push our buttons more and more, and usually it explodes in their face because they still need us (which is why mainstream rich people push robotics and automation and AI down our throats so aggressively because they know all this) and yet never have minority humans won that approach before. Leadership always changes. Life always changes, and no force can stay dominant for ever.
There are several multi billion dollar companies where the founding thesis was “what if we just ignore the law?”
The more interesting question is IMO if AI training actually falls into one of these cases. You can read a book and also copy it, but you do not do because of the law. However, you have the ability to do so. Is having the ability to do something already forbidden?
It's like if you read a plumbing book and then made YouTube videos on how to fix a sink.
The law they broke was pirating the materials, not training per se, even though training is what so many people object to: the judge ruled that actually training a model, when the materials you used were ones you otherwise had lawful access to, was not a breach of law.
IMO, the laws need to change to reflect what tech can now do. This wouldn't be the first time, copyright law has had to shift several times before as new means of reproduction are created.
* the Anthropic one
How? The current system enables the GPL. The GPL protects many open source projects.
Why has restrictive Linux succeeded far more than any BSD ever has?
You still haven't articulated how that's restrictive.
I did not bring up Linux, much less called it bad, yet you made up a strawman and started attacking it. Good day.
A license in isolation isn't interesting. The practical results of projects under a license is interesting.
Linux is a very successful practical result under the GPL and copyright makes the GPL work. Without copyright the GPL would be unenforceable.
Without copyright, the GPL would be unnecessary.
It really is one of those "without law you cannot have freedom" things.
No, they were responding to the post defending OpenAI that you wrote. If you meant to communicate something other than “criticism of OpenAI in this context is unwarranted” then it looks like you forgot to do that and wrote something else instead
Is quoting “criticism of OpenAI” without the rest of the post a way of saying “checkmate”?
uhhh... capitalism