on edit: fixed typo
I would also be extremely surprised if most open source copyright holders didn't already expect their licensing terms to protect against this kind of code/authorship laundering. Speaking individually, I know that it certainly surprised me to hear that GitHub thinks that it's probably okay to regurgitate entire fragments of the training set without preserving the license.
It should be considered as fair-use of USA except we don't use Common Law system so we explicitly state what exempt from the copyright protection.
That said - so they would be able to sell some things in Japan that they couldn't other places.
It's still copyright laundering, if you ask me.
My understanding as a two-year student of ML is that you are allowed in the US to go download any old image, train on it, and then release the model as long as the outputs are "sufficiently transformative."
That last phrase is the key part, and has never been tested in court. It's entirely possible that either I'm mistaken here, or that the courts will soon say that I am mistaken here. https://www.youtube.com/watch?v=4FA_gt9w28o&ab_channel=guava...
Copilot seems ... well, less transformative. I'm still not sure how to feel.
The irony is that copilot won't suggest its own source code, just everyone else's. It is open source without the benefits.
https://twitter.com/luis_in_brief/status/1410985742268911631...
Yeah, don't post your code in the public for everyone to read. If I am musician, and I play my song publicly, people will hum it if they like it. I can't do anything to stop that, except not play the song to anyone.
Plus, to be honest I’m not even sure whether I’m for or against that, I am just wondering if one was against it but still wants to do open source work do they have any recourse?