Yes, this is what the people who use the AI are doing. As well as people who release the AI model or a product that uses it.
Yes, this is what the people who use the AI are doing. As well as people who release the AI model or a product that uses it.
if the people are using the AI to replicate an exiting works, to try to loophole the copyright act, then they will just get sued.
> release the AI model or a product that uses it
I argue that the model itself cannot be construed as copyright violation. After all, the model is information. What if i released a table of all of the word frequencies from books, and published that table? The table of word frequencies does not violate the copyright of the books from which it was derived.
Just because you _could_ re-derive the original books from this dataset, doesn't mean the dataset violates copyright. It _could_, if the dataset cannot do anything else (e.g., i just zipped up the text of the books and released that). But the AI model does not _only_ output the original, but it could also generate new works.
Were they publishing copies, your harry potter analogy would hold, but that is categorically not the thing happening.
Diffusion is not compression, or copying. It's stylistic synthesis, which is more like someone reading all the Harry Potter books, then doing a podcast about Harry Potter fan fiction.
Which, incidentally, exists (1)