(Copied from a comment of mine written more than three years ago: <https://news.ycombinator.com/item?id=33582047>)
(Copied from a comment of mine written more than three years ago: <https://news.ycombinator.com/item?id=33582047>)
Arguments that make a case that NN training is copyright violation are much more compelling to me than this.
I'll prove it by induction: Imagine that I have a service where I "train" a model on a single image of Indiana Jones. Now you prompt it, and my model "generates" the same image. I sell you this service, and no money goes to the copyright holder of the original image. This is obviously infringment.
There's no reason why training on a billion images is any different, besides the fact that the lines are blurred by the model weights not being parseable
You gloss over this as if it's a given. I don't agree. I think you're doing a different thing when you're sampling billions of things equallly.
the model isn't the one infringing. It's the end user inputting the prompt.
The model itself is not a derivative work, in the same way that an artist and photoshop aren't a derivative work when they reproduce indiana jones's likeness.
A regulation that require restaurants to have a public bathroom is more akin to regulation that also require restaurants to check id when selling alcohol to young customers. Neither requirement has any relation with land rights, but is related to the right of operating a company that sell food to the public.
This is not the case in the US yet many places still have public restrooms, due to it benefiting the users themselves regardless of government.