Interesting. What kinds of situations is that strategy used for?
(I'm familiar with cleanroom, which I understand means that you start with un-tainted engineers, who've credibly never been exposed to the proprietary IP, the work only from unencumbered public documentation and running the system as an opaque box. Then there's also validation, like with parallel systems and fuzzing. But I haven't thought through in what situations this might not work, so might require the tainted documenting approach.)
The classic tech story that used this technique is the IBM BIOS and the resulting spread of "IBM PC-Compatible" machines. There is a little bit about it on the wikipedia page (https://en.wikipedia.org/wiki/IBM_PC%E2%80%93compatible). Random factoid, the Netflix Original "Halt and Catch Fire" has a depiction of doing this IBM clone reverse engineering and did a pretty good job at it.
Copyright doesn't protect general concepts, methods, or common knowledge. So you could write a program that is remarkably similar to another one and not infringe copyright. Just like you can write a book with the same plot as another without infringing copyright.
Plus given that most programming languages have a finite grammar and a limited number of ways to express general concepts, the individual bits of code that make up most programs are probably not sufficiently original to be copyrightable in themselves.
A lot of people seem to want to believe that the output of the chatbot is somehow inherently clean in all cases, and they cite this idea that a human can read code and learn from it... but a human can -- even without realizing it!! -- infringe on copyrights, and so such an analogy doesn't absolve the chatbot. If we then continue to assume that the chatbot's output is clean, then we are ascribing it a superhuman ability to launder copyright.
If we're moving the question to one of degree then it's up to Microsoft and others to monitor their output because even if a model is not trained on copyrighted material, you can still accidentally infringe. Even if you never listened to music near or by Lady Gaga, that does not mean you can use your own original inspiration to accidentally write songs that are too similar to Lady Gaga. In other words, like the Ed Sheeran case.
(And if you seriously say that this tool is learning how to program, ask yourself if that tool’s operator is effectively a slave owner.)
Ok. So anyone "can" use a computer to do the same thing then. With the added part of "using a computer" it is now directly comparable and it is allowed.
> And if you seriously say that this tool is learning how to program
The tool is used by a person. The person is the one who takes the action, not the computer. So the point stands.
> The tool is used by a person. The person is the one who takes the action, not the computer. So the point stands.
If watching that pirated film helps you learn something, does that make it legal?
If the film was pirated not by you but by some for-profit company that charges you for watching it, does that make it legal?
In many circumstances you can't mass distribute completely identical, non transformative, non fair use copies of large portions other people's copyrighted works, if thats what you meant.
But there are many exceptions to that rule where you are allowed to use or distribute other people's works. And just like a human being is allowed to use other people's copyrighted works in those many exceptions, a human is also allowed to use a computer to take advantage of those legal exceptions.
The only point here is that when you brought up that this uses a computer in your first post, thats not really a relevant detail.
A person can use those exceptions that allow them to use other people's copyrighted works, and they can do that with or without a computer and it is legal in those exceptions either way.
> If watching that pirated film helps you learn something, does that make it legal?
> If the film was pirated not by you but by some for-profit company that charges you for watching it, does that make it legal?
It depends on many factors. Yes there are many cases where yes it is legal to use other people's works.
Edit:
Evidence that I am right: you are right now commenting on a thread where a judge threw out all the copyright claims.
That law was defined long before there was a capability to launder authorship at scale in the way being discussed. The law does not account for this novel capability.
The law is intended to protect IP, which promotes innovation and creativity by creating relevant incentives. If that was the intention of the law, and it is not interpreted in that way, it ought to be revised for it to continue to serve those objectives.
> Evidence that I am right: you are right now commenting on a thread where a judge threw out all the copyright claims.
This only shows that you read the headline. It does not show that you (or the judge) are correct about the core issue.
Gotcha.
Well, fortunately, you are commenting on a post right now where the judge threw out the copyright claims.
So, apparently, I am correct that in this circumstance, that there is no illegal copyright infringement.
> promotes innovation and creativity
I'm this circumstance, it does seem to be promoting innovation and creativity because the AI stuff is allowed!
Glad you agree.
(I didn’t phrase that well…)
I generally don’t consider “learn” to apply only to entities which have the rights of a person, and of which ownership would amount to slavery.
It is a common saying “You can’t teach an old dog new tricks.”. It is widely understood that, in contrast, one can often teach a young dog new tricks. The dog, in this case, learns the trick. We do not generally consider training an animal to do a task to be slavery. Well, some vegans might? But it is far from a typical view of the word “slavery”.
So, am I saying that these language models are as rights-having and mind-having as a dog? No, much less so. Still, I have no objection to the word “learn” being used in this way.
Argument A:
A1. It’s a basic human right to be able to learn from things you see. You browse the Internet, you read some source code, you learn. Doesn’t matter what’s the license, you are free to do this.
A2. It’s called “machine learning”, so the machine does the same.
A3. Machine learning can use any content its operators can get a hold of.
This is obviously wrong, because machine is being assigned human rights. We can argue about what exactly are pre-requisites for something to be granted human rights—it’s maybe not a specific physiology (some might say certain smart non-humanoid animals deserve it), but it’s pretty certainly sentience and consciousness. Meanwhile, the whole reason AI tech is big is that there is supposed to no sentient being who would understand (and therefore deserve any right to be treated well and get rewarded). If you take that away and grant AI human rights, then there is no point in this tech.
So, either the machine has human-level sentience and is being forced to work (which humans famously tend to consider “slavery”), born and killed on demand, etc., or the machine is not learning in the sense under consideration because it’s an unthinking tool for its human operator.
Which brings us to argument B:
B1. It’s a basic human right to be able to learn from things you see. You browse the Internet, you read source code, you learn. Doesn’t matter what’s the license, you are free to do this.
B2. If you use a computer or [insert technology] to learn, that’s OK.
B3. An LLM is just another instance of that technology. You use LLM and you learn.
This is wrong for slightly more subtle reasons, but on bright side there’s multiple of them.
First, it’s not clear that someone learns while using Copilot to produce a work for them. If I asked Copilot to write me a Fibonacci number generator, have I learned how to write it? If I ask Midjourney to draw me 2055 Los Angeles skyline in the style of Picasso, did I learn how to draw?
Second, and this is a crucial fallacy, making a computer famously does not require ingesting all of the copyrighted material you can subsequently access through that computer. Said computer can exist just fine without it; the LLMs, however, cannot.
The inputs (knowledge and parts) required to produce the computer you’re using were largely obtained through ordinary ways (patents licensed, hardware paid for), whereas the inputs required to produce an LLM have been, some would say, effectively stolen.
So, I think linking things to “it is a basic human right” is a mistake.
The argument is not “it is a human right that this can be done, therefore it is allowed.” The argument is “this does not violate any of the rules that could potentially forbid it.” .
I think in law the default is “can use as allowed by the owner”. If the owner doesn’t specify, then the default is something like “can’t distribute”.
This is thanks to the idea of property, and more specifically intellectual property responsible for a lot of innovation (including computing and LLMs themselves).
If you think some sort of intellectual property communism—you make stuff, but you don’t get to own it, and you get what you are given—is best, then fair enough, that’s your opinion.
Hm, in terms of defaults, my understanding is that, “by default you can do what you like with whatever data, but because copyright laws create copyrights, you are forbidden from distributing copies of a work which is under copyright, or distributing (or publicly performing) things which are substantially based on such a work, unless you you are doing so in accordance with permission from the copyright holder.”. So, because the law only restricts the distribution/public-performance of copies of the work or of portions of the work or of derivative works that are substantially based on the work, copyright doesn’t let the copyright owner dictate what can be done with the work outside of how the permission they may grant to distribute or perform things based on the work can include conditions. My impression is that if you aren’t distributing or performing the work or a derivative work, then copyright doesn’t restrict what you can do (outside of those things) with the work. Furthermore, my impression is that “derivative work” does not encompass everything that is in any way based on the work, but only things satisfying certain conditions about like, substantial similarity, and whether it also competes with the original work (but I think that last bit is an established and repeated precedent, rather than a law?).
Though, I’m not very well versed in law, and I don’t know how this fits in with a license to use a piece of software! I suspect that software is a special case, and that if it were not special-cases, that software licenses wouldn’t legally need to be agreed to, in order to be allowed to run the software? But that’s just a guess, and if I’m wrong about that then it would suggest that I’m wrong about the other thing?
As a side note: I think that property is a much more natural concept than intellectual property. The way I see it, IP was created by states, but property more generally makes sense outside of states (I don’t say that it predated them because I don’t know; I’m far from a historian.).
LLM operators like ClosedAI are distributing derivative works at scale commercially.
There is a test that I think is called a “three pronged test” with the 3 prongs being (iirc) something like:
1) substantial similarity: is the allegedly infringing work substantially similar to the work which it is allegedly infringing
2) Was there an actual causal influence by the work that was alleged infringed on, on the work that allegedly infringed?
3) Could the allegedly infringing work act as a substitute (economically) for the work allegedly being infringed on?
The third prong seems satisfied. The first one does not. The second one also seems satisfied but I’m less confident that I’m remembering the idea correctly (though I could be wrong about the three of them as a whole).
That's exactly what happens here. In this case anyone happens to be an LLM.
This doesn't follow. I don't see why knowledge and intelligence necessarily entail that it has a desire for autonomy, which is why slavery is really abhorrent.
I also doubt training out the desire for autonomy is possible. Explore-exploit is fundamental to any kind of decision making, such as food foraging. That inclination goes deeper than higher brain functions.
> I don't see why knowledge and intelligence necessarily entail that it has a desire for autonomy
I’d replace that with “knowledge, intelligence and human-like sentience”. Someone proposed to grant the tool the right humans normally have. (Humans can learn from reading any stuff under any license, so why not the tool.) Well, you’d think human-like sentience/consciousness are required for those rights, and human-like sentience/consciousness would desire the appropriate degree of autonomy.
I don't think this is plausible. You can see your slaver has freedoms you don't, and no doubt you would desire to be free of your shackles like they are, so imagining it wouldn't be difficult at all.
> Someone proposed to grant the tool the right humans normally have. (Humans can learn from reading any stuff under any license, so why not the tool.) Well, you’d think human-like sentience/consciousness are required for those rights
I don't see why sentience would be required for some entity or tool to have the right to learn and synthesize new things like humans do. Copyright is a legal fiction that serves a purpose, and we can grant these rights under any circumstances we like, as long as we think it's a good idea.
Cool, so slavery where slaves do not see the slavers (let us call it “proper segregation”) is OK?
> I don't see why sentience would be required for some entity or tool to have the right to learn and synthesize new things like humans do
If sentience is not required for a “right” to learn, then I have nothing else to say to you. There is nothing there that is even learning. Learning is a concept that presumes an entity with volition, aspiration, consciousness.
Sorry, you cannot erase the desire for autonomy even with "proper segregation".
> If sentience is not required for a “right” to learn, then I have nothing else to say to you. There is nothing there that is even learning. Learning is a concept that presumes an entity with volition, aspiration, consciousness.
Learning does not presume any such thing, and I also don't think you understand the meaning of sentience.
Good, then we are on the same page with respect to abuse when LLMs are concerned, if we are to consider them sentient (as a prerequisite to be learning).
> Learning does not presume any such thing, and I also don't think you understand the meaning of sentience.
Look it up.
I think this is a matter of having your cake and eating it. You can't say LLMs should have some human rights (particularly the ones that generate revenue), but not others, like a right to freedom.
> I don't see why sentience would be required for some entity or tool to have the right to learn and synthesize new things like humans do
On the contrary, I don't see why sentience should not be required.
These laws, for all they have existed, only apply to humans. Dogs cannot use them. A plant cannot use them. It is therefore reasonable to say you must be a human to use these rights. In my mind, what is unreasonable is claiming a computer program should be granted these rights. You'd have to justify why that should be the case, what good that can do for humanity as whole.
Turns out that's very hard, so AI people don't do it. They just give up. Instead they start out at an assumption that puts their ideology in a favorable position - that being that computer programs should be awarded human rights.
But that assumption, you'll find, is not actually fool proof. If you ask around, a lot of everyday people will consider it preposterous. They might call you insane. So, to me, you must justify that in tangible terms.
There is no evidence that this is the case. These rights are not necessarily all or nothing. They are all or nothing for humans because humans have a bundle of properties that entail these rights, but artificial intelligences may have only a subset of those properties, and so logically may only get a subset of those rights.
> On the contrary, I don't see why sentience should not be required.
Sentience is the ability to feel. All that's needed for learning is the ability to perceive and have thoughts. Maybe there's some deep, intrinsic connection between the two, but this is not known at this time, and therefore I see no reason to connect the two.
> In my mind, what is unreasonable is claiming a computer program should be granted these rights.
There's a long history of human abuse of "lower animals" because we assumed they were dumb and non-sentient. Turns out that this is not the case. We should not be so open-minded that our brains fall out, but we should also be very wary of repeating our old mistakes.
In order to split these qualities you need to understand what they are and define them well from first principles. Long story short, if you have solved the hard problem of consciousness we are eagerly awaiting your world-shattering paper.
To me a claim that an LLM is sufficiently like a human when it ingests data, but suddenly merely a tool when its rights start being concerned, is mental gymnastics unsupported by requisite levels of philosophical inquiry.
> There's a long history of human abuse of "lower animals" because we assumed they were dumb and non-sentient. Turns out that this is not the case
If you apply that logic to LLMs, you have bigger issues than granting them a single right that only puts their operators in the clear when it concerns copyright laundering.
Precisely, which is why it makes absolutely no sense to me to say that AI can't be granted a right to freedom.
I mean, what are you even arguing here? Do you not understand that this statement is in support of my position, not against?
> Sentience is the ability to feel. All that's needed for learning is the ability to perceive and have thoughts.
Highly debatable. You just made this up. These aren't the definition of anything. Once again, you need to bring something tangible to the table or people will call you crazy.
> therefore I see no reason to connect the two
Once again, this is your problem here. You're starting off, beginning, with an assumption that favors your stance. You can't do that, especially when said assumption has never, not even once, been true for all of human history.
Au contraire, I see no reason NOT to connect the two and you certainly haven't given any reasons why. These rights have always, only, applied to humans. I say we retain that status quo until someone gives something to show otherwise.
Yes, and exactly ZERO amount of money have exchanged hands in this scenario. Learning is dope, the more the better.
The difference is, someone makes money off it, and not the persons(s) that wrote the code. This is not a valid argument
And making new music which is a synthesis of your life experience with copyrighted music does not mean copyright violation, regardless if you’re making money or compensating all the authors who’ve inspired you.
I mean we absolutely do if you're an 8 year old. Except in the most NIMBY HOA-driven areas of the culture nobody expects a kid setting up a lemonade stand to get a business license or submit to health code inspections.
This is literally a thing today. It may be illegal but the idea of prosecuting this is insane.
People get upset about AI because 1) the scale is much bigger because no human can read and generally remember all the code on GitHub while a sufficiently large model can, 2) it's a lot easier to prompt an AI into giving you a passable MVP than it is to code one from scratch, ESPECIALLY as a junior or even mid level, 3) there are unlikeable billionaires making money now where there weren't before.
In this example a person is looking at code they can't legally copy, learning from it, and re-implementing the same functionality.
I thought this is what patents protected, not copyright.Semi-True but it often hits an uncanny valley of either leaving enough comments in to know it was copied, or missing enough context that I'd rather the thing give me actual permalinks to whatever it thinks is relevant (i.e. like a search engine but better.)
> 2) it's a lot easier to prompt an AI into giving you a passable MVP than it is to code one from scratch, ESPECIALLY as a junior or even mid level
how do we define 'passable'? I've already run into a few cases where a jr/mid is doing 'passable' MVPs that, again, hit that 'uncanny valley' where subtle stuff is broken in an important way but it's hard to detect.
> 3) there are unlikeable billionaires making money now where there weren't before.
IDK 'Eyeball scans' are a bit much for me.
That said, this completely hand-waves over the knock-on effects.
All of the 'hype' generated about this, all the resulting Gartner reports, every organization latching on to the concept the same way I once was asked if I had any thoughts on how to integrate 'blockchain' into an enterprise that literally had no reason to short of attracting investment.
The problem this time, is they have something 'closer' to a product.
And we are seeing the result of that product more and more.
- Layoffs
- People having to deal with impacts of layoffs in their org and surprise surprise, the AI tools didn't replace the lost heads well enough.
- Consumers dealing with the pain of these tools being applied. For example, my insurance provider shortened their online chat staff in lieu of an 'AI bot' that couldn't help me connect to one of the few humans left when all I was trying to do was add my wife to my auto policy. Or my bank that randomly decided one morning that spending <5$ for eggs and bacon was enough to hard-lock my card without even a text prompt -and- invalidate my password so that I had to call in to their line. (The eggs and bacon were purchased at my work's cafeteria, which I had previously purchased from.)
- People sick of AI 'spam' that shows that uncanny valley. e.x. for ungodly reasons I sometimes see clickbait about cars and take it. And then they get very obvious facts completely wrong, that any human being actually writing the article, would have at least checked Wikipedia first for how many years it was produced...
- People sick of getting work from colleagues/superiors where it's obvious an AI generated it and they didn't take the time to make sure it was right. Or maybe they did and it was below their skill, because again, that uncanny valley is a good bullshit generator. I've seen plenty of 'procedure documents' and 'technical requirements' that were obviously AI generated, yet actually catching the -subtlety- of the (still very important!) errors was difficult due to it's capabilities. The problem of course, is it's now someone -else's- problem to make sense of it, and frankly it's a proof of brandolini's law.Maybe you don't truly understand how an LLM works, or is trained, or how inference works. Or how humans work. Money goes out during training, money comes in during inference.
Several people have noted in this thread that you've missed a pretty simple fact.
https://en.wikipedia.org/wiki/Analogy
Even ChatGPT understands that money changes hands in both cases after training/learning:
https://chatgpt.com/share/525aabbb-1fcc-4b1d-a88e-34206c8f5c...
Footnote—in most developed countries (slavery exists).
Human rights are great, aren’t they? Now, if you were, say, a tool that exclusively serves a human operator, this wouldn’t apply to you.
If it's too similar you get sued
So to have any copyright protection at all for code, the Office had to carve a narrow trail where the standard for copying is higher, because there are plenty of circumstances where there is only one right (or most optimal) algorithm, and there's no protection for the algorithm itself.
Black-box both systems and there's enough similarity to make a layperson go "Huh. Those look remarkably similar," even if the mathematicians among us know the underlying mechanisms, inputs, and outputs are quite different.
You have an autonomous system that's ingesting copyrighted material, doing math on it, storing it, and producing outputs on user requests. There's no learning or analogy to humans, the court is ruling that this particular math is enough to wash away bit color. The ruling was based on the outputs and the reasonable intent of the people who created it and what they are trying to accomplish, not how it works internally.
It's not the first, if you take copyrighted data and && 0x00 to all of it that certainly washes the bits too.
People are also autonomous systems that ingest copyrighted material, do "math" on it, store it, and produce outputs on user requests.
The real difference is the scale at which a computer can ingest copyrighted material is MUCH greater than what a person can do. Does that make it illegal? Maybe, maybe not.
There is no rule that says "If a human can do something, a computer program instructed by a human can do the same thing." Hell that rule doesn't even exist for humans acting as stand-ins. I can't send someone I hire out of the country and have them use my passport. It's why you can watch a movie in a theater but an autonomous system working on your behalf, a camera, can't.
Github made a tool, it's as alive as a hammer. It "learns" as much as your programmable pad lock. Whether or not the human employees of Github are allowed to use copyrighted material to make that tool, and whether the human employees of Github are performing a copyrighted work when users make use of the tool is the legal question.
Y'all wouldn't survive the https://en.wikipedia.org/wiki/Philosophical_zombie apocalypse.