> GitHub Copilot is a code synthesizer, not a search engine: the vast majority of the code that it suggests is uniquely generated and has never been seen before. We found that about 0.1% of the time, the suggestion may contain some snippets that are verbatim from the training set.
1. https://news.ycombinator.com/item?id=23093911
2. https://www.reddit.com/r/HobbyDrama/comments/gfam2y/furries_...
[1] https://docs.github.com/en/github/site-policy/github-terms-o...
Probabilities are fun!
e.g. https://lwn.net/Articles/61292/ and most likely only one opinion.
on the other hand, it would be interesting to learn about what the copyright implications are of
a) creating a utility like copilot (it is a software program) and contains a corpus based on copyrighted material (the database that has been trained)
b) using it to create code based on the corpus and resulting in software as a work under copyright.
I say dumb because I am, perhaps chauvinistically, assuming that no brilliant algorithmic insights will be transferred via the AI copilot, only work that might have been conceptually hard to a beginner but feels routine to a day laborer.
Then again that assumption suggests there'd be nothing to sue over.
Also, Oracle v Google opens the possibility of a fair-use defense in the event that Copilot does regurgitate a small snippet of code.
https://twitter.com/natfriedman/status/1409883713786241032
Basically they are building a system to find explicit copying and warn developers when the output is verbatim.