Once we have adopted the attitude that we can just copy as we please without attribution, it would be much more difficult to find motivated artists, and we would have failed as a society.
I asked if I can learn from that code, which obviously I can. There is no license that says "You cannot learn from this code and take the things you learn to become a better programmer".
That's exactly what I do and it's exactly what AI do.
Did you actually read the link you were given? Clean room design is because you may inadvertently plagiarize copyrighted works from your memory of reading it.
i.e. the act of reading may cause accidental infringement when implementing the "things you learn"
Surely you know this isn't the case right? Maybe you're confused because we're talking about programming and not a different creative artform?
Great artists read, watch and consume copyrighted works of art all day, if they didn't they wouldn't be great artists. And yet the content they produce is entirely there own, free from the copyright of the works they learned from.
What's the difference then in programming? Why can an artist be trusted not to reproduce the copyrighted works that they learned from but not the programmer?
Every author reads books
Every painter view paintings
Unless you're arguing that every single artist across every field of artistic expression is constantly being jeopardized by claims of copyright infringement, this is a nonsensical point to make.
They cant. which is why that quote "Good artists copy, great artists steal" exists.
AI has already been shown to be "accidentally" reproducing copyrighted work. You too, can do the same.
Its likely no-one (including yourself) will ever be aware of it - but strictly speaking it would still be copyright infringement. This is the relevance and context of the link you were given.
You're describing thought crime right now. It's not illegal to learn things.
That's completely different from reading and learning from code, which is what grondo described.
Clean room design relies on this, in a clean room design you have one party read and describe the protected work, and another party implement it. That first party reading the protected work is learning from closed-source IP.
AI (e.g. copilot) has already been shown to break copyright of material in its training set. Thats the context of this whole thread.
If an AI infringes on copyright then it infringes on copyright, that's unfortunate for the distributors of that code.
Humans accidentally infringe on copyright sometimes too. It's not a unique problem to machine learning. The potential to infringe on copyright has not made observing/learning/watching/reading copyright materials prohibited for humans, nor should it or (likely) will it become prohibited for machine learning algorithms.
Grondo said that AI should be given access to all code, including private and unlicensed code.
He was given a link to Clean Room Design demonstrating the problem with the same entity (the AI) reading and learning from the existing code and the risk of regurgitation when writing new code.
He goes on to say thats what he does, which doesn't change that fact.
> Humans accidentally infringe on copyright sometimes too.
Indeed we do, and its almost entirely unnoticed, even by the author.
> nor should it or (likely) will it become illegal for machine learning algorithms.
If those machine learning algorithms are taking in unlicensed material and then they later output unlicensed and/or copyrighted material, then they are a liability. Why would you want that when you can train it otherwise and be sure it NEVER infringes others IP? Its a no-brainer, surely. Or are you assuming there is some magic inherent in other peoples private code?
Because it could produce a better model that produces better code.
You're now arguing a heavily reduced point. That a model that trained on proprietary code is at higher risk of reproducing infringing code is not a point under contention. The clean room serves the same purpose, it is a risk mitigation strategy.
Risk mitigation is a choice, left up to individuals. Maybe you use a clean room design, maybe you don't. Maybe you use a model trained on closed-source IP, maybe you don't. There are risks associated with these choices, but that is up to individuals to make.
The choice to observe closed source IP and learn from it shouldn't be prohibited just because some won't want to assume that risk.
But fundamentally the problem is copyright, the copying of existing IP, not knowledge. grondo4 is completely correct that there is no legal framework that prevents learning from closed-source IP.
If such a framework existed, clean room design would not work. The initial spec-writers in a clean room design are reading the protected work.
Right. And they're only exposing elements presumably not covered by copyright to the developers writing the code. (Of course, this assumes they had legitimate access to the code in the first place.)
Clean room design isn't a requirement in the case of, say, writing a BIOS which may have been when this first came up. But it's a lot easier to defend against a copyright claim when it's documented that the people who wrote the code never saw the original.
Unlike with patents, independent creation isn't a copyright violation.
What are you contesting?
Did you acquire it illegally? That's illegal.
Was it publicly available? That's fine, so long as you aren't producing exact copies and violate normal copyright law.
Yes you are allowed to read closed-source, proprietary code and become a better programmer for it.
I've decompiled games to learn how they structure their code to improve the structure of games that I program. I had no right to that code and I used it to become a better programmer just like AI do.
That's not copyright infringement. You have a right to stop me from using your code, not learning from it.
So, yes: They have a right to stop you from "learning" from their code. If you want that right, see if they're willing to sell that right to you.
They absolutely do not, and as pedantic as it may be I think it's very important that you and everyone else in this thread know what their rights are.
If you sign a contract / EULA that says you cannot decompile someone's code than yes you are liable for any damages promised in that contract for violating it.
But who says that I ever signed a EULA for the games I decompiled? Who says I didn't find a copy on a hard drive I bought at a yard sale or someone sent me the decompiled binary themselves?
Those people may have violated the contract but I did not.
There is no law preventing you from learning from code, art, film or any other copyrighted media. Nor is there any law (or should there be any law IMO) that stops an AI from learning from copyrighted media.
Learning from each other regardless of intellectual property law is how the human race advances itself. The fact that we've managed to that automate human progress is incredible, and it's very good that our laws are the way they are that we can allow that to happen.
Your argument is based on the idea that you and AI should have the same rights?
I do not see how this works unless AI going to be entitled to minimum wage and paid leave?
Otherwise it is just a money grab