Like if you can't figure out which works were used to create the AI just by looking, its hard to argue that they "copied" the work. Copyright is not a general prohibition on using the copyrighted work only the creativity contained within.
Like if you can't figure out which works were used to create the AI just by looking, its hard to argue that they "copied" the work. Copyright is not a general prohibition on using the copyrighted work only the creativity contained within.
It isn't difficult to show copyright infringement in these models. The assumption should be that copyright infringement has occured until proven otherwise.
Just the fact that they are indiscriminately web scraping proves that. Just because it is publicly and (monetarily) freely available doesn't mean it isnt copyrighted.
When I first tried Copilot, I asked it to write a simple algorithm like FizzBuzz and it ripped off a random repo verbatim, complete with Korean comments and typos. Image models will also happily generate near-identical [2] copies (usually with some added noise) of copyrighted images with the right prompt.
[1] https://bair.berkeley.edu/blog/2020/12/20/lmmem/
[2] https://www.theregister.com/2023/02/06/uh_oh_attackers_can_e...
A human reproducing a patagraph word for word in an educational context would probably not be considered copyright infringement (although lack of attribution might be problematic). In the US anyways. The US is sonewhat unique as having very broad fair use when it comes to material used in an educational context, much broader than most other countries.
Another factor is the effect on the market of the original product.
Non-attribution + commercial use + affecting the marketability of the original product (which is what LLMs do) seems unlikely to be considered fair use by any existing precedent.
That being said IANAL.