We have no way to reconstruct memories from a preserved brain (yet). The exact ways in which humans form memories and store information isn't even known yet; we're still drilling into the specifics from higher-level concepts.
Modeling the human brain like nodes with weights ignores a lot of biological processes. Blood/oxygen flow, hormones, neurotransmitter decay, physical locality, chemical delays and interference from things like myelin sheaths, and other physical processes affect the synapses that are partially mirrored by computer simulations of neural networks. Unlike neural networks, human brains also don't work based on a single clock signal triggering input and output from all notes in instant steps.
Human memories are also not just "data in, weights out". They are heavily modified by things like mood, concentration, language(s) spoken, context, and emotional triggers. There's no way to feed a dictionary into a brain. Memory preservation consists of multiple stages, with differing memory types, involving various brain segments with dedicated functionality that can actually grow back due to neuroplasticity in some cases.
Efforts are being made to emulate living cells on computers, but LLMs aren't that. Inversely, efforts are also made to feed brain cells artificial signals and train them to play video games, which results in different behaviour compared to the systems we use for LLMs or other AI systems.
It doesn’t seem clear to me that it does.
Is the argument that sufficient complexity in how an “intelligence” processes this copyrighted data leads to the output being transformative vs not transformative in a more simple mind/model?
What if the output is exactly the same, or comparable enough, regardless of the degree of complexity of the mind/model?
This would be a separate argument, though, from the notion I responded to above that the difference in the processes of a human mind and LLM are the reason why "learning" from copyrighted material is a violation of copyright in one case and not the other.
In my view the biggest issue to raise is the effect of the use upon the potential market for or value of the copyrighted work (https://en.wikipedia.org/wiki/Fair_use). The most spectacular example of damage would be stackoverflow, thought stackoverflow content is not copyrighted. I think there is little doubt that LLM's drive attention from original sources. That might be deemed damaging, especially in the long run.
There's no damage to the potential market or value of the works because they're given away for free by their owners.
There may be, because technically copy/pasting SO code is governed by CC BY-SA 4.0 license which requires attribution and things aren't so obvious especially for commercial purpose.
I’m sure there’ll be a human on the books, on paper.
If you turn an LLM into a person, you may have a ethical and legal basis for treating that LLM like a person. There's no law about artificial intelligence being sentient or not, but law applies only to people, so that'd be the supreme court case of the century. I remember the Star Trek TNG episode about this topic and while the answer was perhaps more obvious with mister Data, the best arguments for and against synthetic consciousnesses have all been made in that episode.
The law doesn't care for how human-like programmers may think their program is, and neither should it in my opinion. What matters to the law is that the output is a result of an automated process, which comes with a completely separate set of rules and conditions compared to fair use.
The complexity of the program isn't a very good legal defence in my opinion because there's no clear line when the complexity would be enough to be considered human like. You could, for example, also claim that a computer is just very good at doing imitations, just like a person can be good at doing imitations on stage, and that an mp3 file is just an elaborate imitation act.
Even with a digital system identical to a physical system I don't think you can state human-ness as an argument. A tape recorder is just a sophisticated, automated way of sending an electric field through a magnet, similar to what a human can do with a dynamo and a spool of tape; a sort of delayed-action theremin, which would turn it into a musical signal. In turn, a neural network can be solved by human brain power if you pay enough people to work on a single iteration for an entire year. Almost everything a computer can do is just a sophisticated way of doing what humans are already doing, so I don't see why this is different when it comes to this topic.
There's no obvious "this is human like now" threshold and I doubt there will be until we know exactly how the human brain works.
When it comes to AI generated works, we don't currently know where the line between copyright violation and copyright exemption lies. If this goes through, it's the third major lawsuit of its type, the other two being actions against Stable Diffusion by artists to prevent it from producing derived works from their art.
IANA but I think you can assume that the "but computers are just like digital humans" approach won't fly in court. That's not really important, though; both sides of the coin are already having clever lawyers write up legal defences for their points of view, and it's more than likely that some other factors ("is a model a derived work" (probably) and "is the output of a model a derived work" (who knows!)) will decide the future of AI and copyright. The interesting thing is that academic research is essentially exempt from copyright law, so nobody can demand takedown of their content from research data sets, but whether the commercial branch of AI companies can use their academically generated models to serve their customers?
As an upside, I don't think the death of ChatGPT is the end of AI. This whole scenario could've easily been avoided if AI companies paid for their data set or had restricted themselves to works they had the license to (public domain, CC0, etc.), and OpenAI in particular has been pretty brazen in their "we'll see about it if it ever comes up" approach. Companies like Github are probably in the best place legally, where their users are already signing off on "we can take your content and do whatever the fuck we want" terms and conditions.
I'm pretty upset at companies using our personal data to make gobs of money off of. I'm also upset that they're now using our knowledge work to make even more money off of. We don't exist as computational nodes for them, a free resource to exhaust. It is a completely one way street with no consent. So I am in favor of all of these companies getting a reality check.
The fact of the matter here is that parties, such as OpenAI, are benefiting from others' knowledge work, protected or not, in a completely one-sided way and all for free. And I don't feel sorry for companies that need to build Trojan horse products, such as OpenAI and Google, in order to survive off of other people's data that they never compensate for.
Someone else made the analogy - if you read a NYT article and then go do a stock trade based on what you read, didn’t you do that based on value generated by them, and why wouldn’t they own that too?
Like if you don’t want to talk analogies then talk principles, and humans are diffusion machines. When you write a term paper from sources you are simply diffusing those words into a new arrangement, but it’s still fundamentally someone else’s work. Why do you get to profit off the model that results from someone else’s work?
Copyright has mutated into this bizarre chimaera where people (like NYT) are essentially claiming ownership of ideas (and derivative works fall into a similar space) and that’s inherently in conflict with a system that is supposed to promote the creation of works. But it has turned into this bizarre shibboleth that if you came up with an idea it’s yours for life+75 years, completely yours and nobody else can work off it or remix it without crediting you. And that’s an unusual state, humanity hasn’t existed like this forever, the Berne convention is only 50 years old and already falling apart from unintended consequences.
Anyway there’s no proof that copyright benefits the small guy more than corporations. Disney squashing someone for writing a Star Wars fan fiction happens a lot more than Disney ripping off someone’s fanfic character for their series. Like patents there’s this mythos of it benefiting the small guy and that’s absolutely not how it works in the real world.
Also, NYT is a particularly egregious plaintiff here because they’re essentially just factual reporting of occurrences, which (like a phone book) are not really copyrightable in itself. You can copy a phone book without infringing copyright and you can train an AI model on a phonebook, and training an AI model on NYT in particular is basically doing that but for historical facts and occurrences. The fact that this costs money for NYT to generate is irrelevant, this is the “sweat of the brow” doctrine and was already swept aside by the phone book case. Just because you spent time/money making it doesn’t mean it’s copyrightable. A large amount of NYT comment is factual observation and tabulation and probably is not copyrightable in the first place. But separating that out is of course going to be challenging for NYT’s lawyers!
I think historically it's about copying wholesale and redistributing for profit. That doesn't seem to be what's happening here.
If I read a publicly available article, and I create a summary, is that covered by copyright? If I get an AI to do that, is that somehow a 'special' type of summary that is covered?
Can a provider of content somehow say "you may not use this for summarization"? Or apply other terms to my consumption if they are making it publicly available?
I think the comparison is that if you ask a person something like:
"Describe a power station?"
They will likely have never been in a power station. They will be leaning heavily pieces of content they have consumed over the years, that presumably were copyrighted. If you create an article that describes a power station, have you breached copyright and should be sued by all the people over the years that have produced content about power stations?
Is it suddenly different when a computer does it?
The answer is yes.
Everyone here knows where this is heading, and yet people will sit here and defend these companies as if they're on some righteous path. Humans are already becoming disposable statistics and the engines of compute, all for free and all for the benefit of corporations who didn't pay for any of it and don't even contribute back taxes.
So you'd be fine with it if the model only ingested a humanly plausible amount of data? I suspect that would only make their legal issues worse, since the LLM would be much more likely to repeat tokens from the training set verbatim.
I wouldn't be surprised if the avenue of attack is that fair use laws are for humans, not robots, and if an AI has been trained on copyrighted data, that's not fair use.
Also, don't forget that in reality what's happened is that a bunch of copyrighted text is encoded in the LLM in a way a human can't understand, but that the LLM essentially CAN understand.
Still a copyright abolitionist though. Maybe now more people will join the fight?
The weights are not a reproduction of the content. They are capable of it but so is a photocopier a lot more and we didn’t ban those either despite them technically being a lot more useful for violation.
Nah, this is expansionist doctrine and agenda for copyright - these companies are trying to copyright style and locations in latent space now.
It would be different if I trained my own AI, for my personal use.
Again, LLMs don’t copy so it’s not a good metaphor.
You brought up photocopiers and you never said it wasn't a good metaphor so I don't know why you pre-pended "Again".
If it's not a good metaphor, that invalidates your point that:
"we didn’t ban [photocopiers] either despite them technically being a lot more useful for violation"
So now you're arguing against yourself.
I'm going to bypass this question a bit and say, who cares?
Why do we need to treat these things the same way we treat humans? Why can we not say that it's okay if a human does it, and not okay if it's a computer? There's nothing that requires us to establish 'fair' as treating them the same as people.
Because it would be absurd for it to be legal to do X, but illegal to do so with an efficient tool. Especially when the activity X in question is "learning".
Did we somehow stealthily develop a neural interface that lets us feed the 'learning' that 'AI' is doing into a human brain? Have we actually figured out how to do that?
No, we haven't. So humans are still learning the same way, but with a new tool to condense and summarize some information. Kinda like a textbook in school. But we don't treat those as human beings with human rights do we?
No such thing, "learning" in a vacuum describes nothing here. Might as well ask why I am allowed to make noise, e.g. speak, but when I install 5000 watt speakers on every square meter of the planet suddenly it's a problem, and roll my eyes at the inconsistency of not being allowed to "do X more efficiently with a tool".
It's legal in most places in the U.S. to own a semi-automatic weapon that fires x bullets/minute, but not a fully-automatic one that fires 10*x bullets/minute. It's legal for me to send water into my sewer by flushing the toilet, but not for me to pump the larger volume of water in my sump pump into that same sewer.
It's legal for a child to throw a handful of sand from the beach into the ocean. Do you think it should be legal for anyone to build and operate a sand throwing machine that methodically throws all of the sand on the beach into the ocean?
We've been living in a world where people read news articles, then wrote almost the same article on their own website to sell ads. It's been a standard business practice for a decade now, what's so special about LLM based blogspam? The end impact is still the same, people reading the blogspam instead of the source
This isn't as well-defined for humans as you might think, so can't be well-defined for LLM techniques by comparing them to human agents.
My argument against the current hoovering up of data under various licences for AI training, which they claim can#'t reproduce anything verbatim, is CoPilot. If there is no risk, then why did they only use public repositories and none of their own private ones? Surely they think their code contains good training material, unless they think their own code is gobbledygook. Or back in terms of licences: if it can't breach the GPL family, then it can't breach their own commercial licensing arrangements.
Would OpenAI have a problem with humans for using some of their code/documents/other in this way?
So I can write my own article about the sky being shown to appear blue much of the time, but I can't copy someone else's article about the same subject.
Like yes if you copy a NYT article verbatim it’s like copying a phone book ads and all, and that’s infringement. But that’s not what a LLM does, NYT doesn’t like their content being used and summarized at all, even in a rearranged form that merely relies on the factual information included in the article. That’s what they want to get paid for, and unfortunately that’s not copyrightable and OpenAI is correct they don’t have to pay for that. NYT disagrees but again, they are kinda attempting to claim copyright on the factual information because they wrote some connective sentences between.
The traditional trick there is to include some small amount of fake data in the directory. You know someone has copied your collection of facts instead of compiling their own because it includes your fake facts. Mapmakers have used the method for at least as long as cartography has been part of our recorded history, see https://en.wikipedia.org/wiki/Trap_street for details. As noted in that page, the legal status of this, like many IP related issues, depends upon jurisdiction.
> But that’s not what a LLM does,
What does it do that means it is only summarizing factual information? While NYT effectively trying to claim copyright on facts is wrong, OpenAI claiming it can't reproduce copyrightable information while it can reproduce/summarize facts found within the same training set seems at best disingenuous.
> Trap streets are not copyrightable under the federal law of the United States. In Nester's Map & Guide Corp. v. Hagstrom Map Co. (1992),[3][4] a United States federal court found that copyright traps are not themselves protectable by copyright. There, the court stated: "[t]o treat 'false' facts interspersed among actual facts and represented as actual facts as fiction would mean that no one could ever reproduce or copy actual facts without risk of reproducing a false fact and thereby violating a copyright ... If such were the law, information could never be reproduced or widely disseminated." (Id. at 733)
And yes the EU has the concept of “database rights” but notionally there is still supposed to be a creative step required in the selection or arrangement of records. So just a raw copy of the numbers in a telephone directory is theoretically not copyrightable, but a telephone book might be because of the creative/transformational step. It’s possible this might be such a low bar that it’s impossible to fail to clear, but, at least on paper you can’t copyright mere facts and figures either.
But either way it’s generally true that simple facts and figures are not protected and trap streets are a discredited and clumsy attempt to work around this.
US law is not the only law.
Current US law has not been as it is for the entire existence of the US.
Trap streets and other such devices have existed much longer than the US.
As well as copyright law the trick can help detect beaches in contractual agreements that cover use of information from services. Action based on such breaches do not necessarily end up in a court of law at all.
Out of court settlements do not necessarily rely on the letter (nor intent) of the law, but often instead the expense (money directly, time, potential reputational risk) of defending a position even if the law is on your side. The threat of action is often enough to make the other party cave and such actions will usually happen well out of public view (I know of one instance involving a list of phone numbers, that I won't go into in detail because while there is no NDA or such in force this discussion is not worth irritating people I have the confidence of!).
> clumsy attempt to work around this.
That the trick is clumsy does not mean it isn't still commonly used (it absolutely is) or that it has not been used in successful cases between map makers and such (it has, one significant example is given in the paragraph directly following the one you selected to quote from the page linked in my previous post).
Irrespective of copyright issues, the question is how to avoid creating a new class of large rent seekers in the LLM space.