[1]: https://twitter.com/mitsuhiko/status/1410886329924194309
[1]: https://twitter.com/mitsuhiko/status/1410886329924194309
Google ‘fixed’ its racist algorithm by removing gorillas from its image-labeling tech - https://www.theverge.com/2018/1/12/16882408/google-racist-go...
(I make no claim that this is actually what is happening, it's just not incompatible with behavior in the generous case).
That's the real problem. The issue is that Microsoft looks bad. The problem was solved by making Microsoft not look as bad. Verbatim output of inputs is not a problem, it's an understood property of the model.
So wait, was the obvious copyright elephant in the room solved somehow?
> The issue is that Microsoft looks bad. The problem was solved by making Microsoft not look as bad.
Depending on who you ask, adding a hack like this in an attempt to make them not look as bad just makes it look worse. Especially when the hack is discovered.
Copilot's FAQ says this (under heading Who owns the code GitHub Copilot helps me write?): "GitHub Copilot is a tool, like a compiler or a pen. The suggestions GitHub Copilot generates, and the code you write with its help, belong to you, and you are responsible for it."
They are essentially affirming that the output is not covered by someone else's copyright, but that is far from clear.
And I think it is precisely the copyright issue that turned this into a PR issue. Verbatim copies are just a very obvious demonstration of the copyright issue. The issue isn't gone when they filter out this specific snippet; the people who are concerned about copyright issues are going to remain concerned.
> GitHub has the rights to use your code for training.
Sure, but that's not at all the same as saying that the output produced by the AI free of copyright issues. That's kind of orthogonal.
The direct analogy would be trying to stop it from outputting The Lord of the Rings by blocking the phrase "a long-expected party" (the title of its first chapter).
That said, at the very least it seems like it would be rude to include code in the training data if the developer has expressly said they don't want that.
[1] From their FAQ: "Training machine learning models on publicly available data is considered fair use across the machine learning community."
You could easily make the case that training is fair use, but that doesn't have to imply the model's output is non-infringing.
For example, it seems reasonable to train a model by feeding copyrighted texts and images, and that model could be useful for analyzing the content, finding facts, or detecting features. But we're in murky waters when the model also starts outputting the original content (be it verbatim or "derived").
Not all that different from human learning: you can study and learn from publicly available books but that doesn't grant you the right to recite their contents and claim it as your own, original work.
Put another way, copyright only applies to creative expressions, not functional expressions. It does not matter how creative the idea is. If the work of authorship is software that embodies the function (and no other expression), it is not copyrightable.
So where is the line between creative and functional expression in software? The law does not provide clear guidance. Ultimately, it’s up to a judge.
the problem is that their AI is insufficiently creative
Isn't it typical Microsoft, in some 90s and 2000s sense?
[1]
if (version.StartsWith(“Windows 9”))
{ /* 95 and 98 */
} else {
http://www.reddit.com/r/technology/comments/2hwlrk/new_windo...