See Marvin Minsky’s comment regarding “suitcase words”.
Hash the content. If the content has been viewed by the user in the past, halve the numerical hash value. Then sort the list.
Doing this to a music list will create a list that is biased toward music the user has listened to before but will still seem random enough to look like intelligent suggestion that the system has learned to identify. It is just math that mimics learning.
The mathematics of artificial neural networks is math. It is only math. One can make it very complex or very simple but in the end it is just math and pointers.
With math, you can just describe everything, including the human brain.
My original argument was specifically about neural networks, that I don't really see the principal difference in how a human learns from reading code, and how Copilot has learned from reading code.
Right now it is impossible to accurate describe the human brain in the form of math. What we can do is write simplified models that either describe or mimics behavior for which, if we apply abstractions, we can call predictive models. Their effectiveness are quite poor but that has never stopped people from trying to use poor predictions models to predict the future.
My statement on human brain was: In principle, you can describe it with math. This doesn't mean that we know how to do that yet.
My statement on Copilot was: Comparing learning of the human brain to learning of the artificial neural network, both are still very similar, much more similar to other (machine or other) learning methods. Sure, there are differences. But my point is: Those differences, why are they relevant for the copyright question?
The mind that can acknowledge and appreciate your work in this scenario (Copilot) does literally nothing of its own free will except 1) take your code and 2) give it to me, possibly combined with someone else's code. This is the sole purpose of its entire existence and full range of its capabilities. Is this enough of a difference compared to a human mind when copyright is concerned?
It spares me from knowing that you exists, that you wrote a library that does this thing I need, that I can contribute to it, etc. In such a scenario, what is the motivation for you to make your library publicly available in the first place (other than generate revenue for Microsoft or whoever I pay for access to the network)? Does copyright have relevance to OSS now?
That's not even scratching the surface of differences
If I ask Stable Diffusion to create a picture of Elon Musk wielding lightnings and riding a giant blue sparrow over a desert during a storm, the result would be more creative than what could be produced by most humans. I believe that counts as a proof.
Also it's pretty clear you haven't worked with many artists from that statement.
Humans are also trained only on things produced by humans. The only exception is nature, but ML model can be trained on photos of nature, too. Also, you are missing the point.
> Also it's pretty clear you haven't worked with many artists from that statement.
1. I employed quite a few artists over the past 15 years.
2. I wasn't talking about artists. I was talking about regular humans. The vast majority of them are absolutely unable to create anything resembling Elon Musk riding on a giant sparrow.
And that is what copoilots AI mostly does.
It doesn't "understand the concepts and reproduce something alike" in the sense a human does. It might understand some concepts here and there but it also does a lot of heavy lifting my verbatim "remembering" (i.e. copy pasting) code.
This is also why some people argue that the cases for copilot and some of the image generation networks are different as some of the image generation networks get much closer to "understanding and reproducing a style". (Through potentially just by it being much easier to blend over copy-pasted snippets in images to a point its unrecognizable.)
One of the main problems GitHub has IMHO is that anyone who has studied such generative methods knows that:
1) they are prone to copy-pasting
2) you don't know what they remembered (i.e. stored copies of in a obscure human unreadable encoding, i.e. just distributing such a network can be a copyright infrigement)
3) you don't know when they copy past
4) the copy pasted code often is a bit obscured, ironically (and coincidentally) often comparable with how someone who knowingly commits copyright theft would obscure the code to avoid automated detection
Which means GitHub knowingly accepted and continued with tricking its copilote users into committing copyright infringement under the assumption that such infringement is most times obscured enough to evade automatic detection....
There is no equal sign between a person and a program.
There is also that thing called "scale" that is critical to the interpretation of the action.
Is eating meat fine? - maybe. Is eating all animals OK? - Hmm...
This argument is hardly less flawed than the one you are criticizing. And you statement that 'there is no equal sign ...' is also unconvincing, as we're not equating these two, but the process of learning, which is quite similar.
Thats the thing, there is no reason to think that they are similar.
It may well be so. I was arguing with respect to your previous post, in which you stated as a relevant difference merely the size of the job.
Ai/ml is not artificial general intelligence. It's a mathematical model.
1. People have certain rights, duties and prohibitions. Equating the right of George Lucas to use ideas he saw with rights of a machine to do that misses the point by the same measure as asserting that MS enslaves the copilot, but in the opposite direction.
2. Scale does matter. If I'm an ordinary person then the act of eating won't ruin the ecosystem. Now imagine a construct that operates under the same principle of eating, but its jaw, stomach and speed of eating is many magnitudes larger - do we apply same limitations to both, because the principle of eating is the same?
Also, since I'm spelling things out, the fact that I'm seeing the same argument many times over, and that it is so obviously flawed, makes me think that this is a symptom of astroturfing.
Hardly relevant, given that the machine has no rights, so no one is equating those with anything. The point is that the machine is doing automated learning on behalf of the developers who are training it, so what should be decided is whether those very human people have a right to train their model in that way.
And the thing is that the machine does not benefit from the same rights as a person, so we can't absolve MS from responsibility because "it does a similar thing to what people do".
So, to add one more point to spell out: the context matters! ;)
The question is whether a person is doing the same as Copilot for this particular case, i.e. reading source code to learn.
You have not really given any argument why this is not the case. Or maybe your reference to scale? So only because Copilot has read more code than a human possibly could, that makes it different? But why exactly is reading a bit of code fine w.r.t. copyright, but reading more code suddenly violates copyright?
Note that the reason why Copilot needs more code to learn is just because the learning currently is not as efficient as for humans.
In all fairness, the article mentions the Fair Use doctrine. You could make the argument that reading a bit of code is allowed, but doing it at a large scale would not be covered as an exemption to copyright.
Counting all the code I have read in my life, it's also quite a lot.
Even novelists do not sit all day long in a closed room reading other people's work and then do a collage of what they've read. Otherwise no books would have been written in the first place.
Cut the AI off humans' work, let it interact with the real world and see what it produces. It will be nothing.
Once (if ever?) an AI is capable of producing an actual original work, I'm fine with other AIs stealing from the first one. Please leave humans alone.
That's correct, but it misses the point.
This is about reconciling 1. being allowed but 2. not being allowed:
1. The human uses a machine, where the machine is an organic one it grew itself.
2. The human uses a machine, where the machine is one it made or acquired.
To a lot of us, there's no difference.
That "experiment" could just as well be done on humans, though, cut them off of any work that any human has done before and you may get simple cave paintings, if you're lucky.
Monkeys have evolved enough to start making their own tools [0].
[0]: https://www.scientificamerican.com/article/monkeys-make-ston...
... yet.
Please study the series of events that unfolded in the music industry after folk begun incorporating recordings made by other artists in their own work and proceeded to sell the result.
Spoiler: The deeply nuanced question of feeding a mechanical recording through a series of complex physical and mathematical apparatus and whether that constituted a transformational creative act did not come up during the proceedings or final judgements!
Is CoPilot just trained on OSS, or on private repos too?
There are GPL repositories which force you to open your code, which is one aspect, and there are "source available" repositories, which allows you to see the code, but forbids everything else.
There are a lot of blurry areas about this, and in my opinion, an AI learns like a human is not a solid basis for fair use.
On the other hand, if private repositories are crawled too, this would be very, very bad.
We just talked this with a couple of friends. I always cite what I got from where (it's just two occasions, but it's not zero), and always respect their licenses.
I'm worried about both ways of the permeation: GPL to closed and closed to open. Open source is a widely misunderstood concept and people (and companies) are using that misunderstanding to validate their blanket options. That's wrong on so many (legal to ethical, and everything in between) levels.
Emulator writers are afraid to read leaked console code, because any resemblance of their code to it means destruction of years (or decades) of reverse engineering and clean room development done in that domain. If code licensing is that important and crucial, why a court tested license (e.g. GPL) is so worthless? Is this fair, again in the same cross-section (legal to ethical)?
There's a lot to be discussed, and a lot of ideas to be re-learnt here. Open Source (or precisely Free / Copylefted software) doesn't mean free for all. We need to understand that.
Am I violating copyright? Yes
Imagine they change the character names in those paragraphs. Am I still violating copyright? Yes
At some point you can change enough of the text to not violate copyright. The grey area involves the courts.
It feels very simple to me so I might be missing something.
> It feels very simple to me so I might be missing something.
In my opinion, you are missing something subtle:
In continental Europe, there is a different law tradition - civil law (https://en.wikipedia.org/wiki/Civil_law_(legal_system) ) - that is different from the Anglo-American common law tradition. To quote from the wikipedia article:
"The civil law system is often contrasted with the common law system, which originated in medieval England, whose intellectual framework historically came from uncodified judge-made case law, and gives precedential authority to prior court decisions. [...] Conceptually, civil law proceeds from abstractions, formulates general principles, and distinguishes substantive rules from procedural rules. It holds case law secondary and subordinate to statutory law."
So if you are attached to the civil law system, you seriously want to avoid this grey area involving the courts (which is much more accepted in common law) and instead want to codify into laws what you mean by this grey area.
The simple layman's version of copyright is that copyright applies to a specific form of a thing and not about the ideas behind that thing.
So, no, George Lucas was not infringing anything. Nor is hip hop music making use of samples infringing anything. Or Andy Warhol integrating photos into his works. Nor is it illegal to paraphrase or refer other authors. And as Oracle found out by challenging it in court, trying to claim ownership over APIs to prevent third party implementations is also not going to work.
All of that falls under fair use. Fair use is what makes copyright useful. Without it you'd have to live in fear that legal copyright holders might come after you if you apply the ideas that you might have been exposed to via their copyrighted work. Fair use exists such that you can make use of information provided to you via a copyrighted work.
It's an interesting test of open source licensing because I'm not aware of any other area of copyright where works come with an explicit "if you use this somewhere else you must credit me as the initial author" in the implied/provided license.
Comparing music, literature, etc. to code is difficult because of both this difference and the existence of software patents. The manner in which infringement happens (and the scale) is often different as well.
It doesn't matter whether it's music, literature, or code. Fair use is fair use. And it's been challenged so often that no judge is going to make any exceptions just because we are now dealing with software.
End of story. No basis for any copyright infringement here. Not even worth trying out in a court because you'd be laughed away. The plaintiffs in this case clearly realized that and did not bother with even trying to prove otherwise.
Software patents are not part of this court case either for the obvious reason that the vast majority of copyright holders in this case don't actually hold any patents whatsoever. And if they would, it would not be Github's problem but the problem of those creating possibly infringing products without a license. Github just gives people access to (public) knowledge here. That's what a patent is: public knowledge. It's up to the user to decide if they are OK shipping products that include that. And it's their problem to do any due diligence.
I don't want Github or any other megacorp-backed entity abusing the open source community in the way micro$oft is here, it's as simple as that. If they wish to train it on entirely proprietary Microsoft code, then by all means go nuts, but to take the work of open source projects and to hide behind the pretense of the mathematical model behind the A"I" learning something is simply ridiculous to me.
I find it quite curious that they're not doing that (training it on their own codebase). Perhaps they're afraid of their little intelligence spitting out proprietary code verbatim like it's been shown to do many times with licensed open source code.
Next hypothetical.
These can be quite inventive works; nevertheless, no-one seriously argues that the video content does not breach the original animators' copyright.
The video content of an amv is a much better analogy for what copilot does to third parties' code than anything else I've seen in this post's discussion so far.