> I'm familiar with the widespread practices of how AI models are trained, yes. There's an implicit "and we must be able to do this" in your argument, which is not at all evident.
Maybe it's not evident to non-practitioners in the field, but to every serious practitioner it is obvious that you can't train a state-of-art model without a lot of data (at least right now without some colossal breakthrough), and that actually licensing that data is not really practical (because you need terabytes of it), and it's definitely going to be impossible for anyone who isn't a megacorporation.
Can we agree on this point? If not can you please explain how do you think that e.g. a single individual like me will be able to train e.g. an image diffusion model (so I'd need a few terabytes of images) if I have to respect the licenses of every image in the training set?
Okay, so I hope we can agree that it won't be possible? So now here's the question: do we want such AI models to exist, or do we want to make them illegal (and maybe available only to huge megacorporations)? These are our only two choices, which logically follow from the requirement that we need a lot of data for training.
What I'm advocating for is that we should allow such models and that they're beneficial to us as a society, hence the "and we must be able to do this" in my argument.
I'm starting with the assumption that I want these models to exists and that everyone should have access to them, and then go backwards from that. What you're starting with is the assumption that the training data copyright should be respected, and you're going backwards from that. But these two graphs are not connected, which is why we can't agree.
Or in other words, what you're (indirectly) advocating for is to make those large models effectively illegal. This is, of course, a valid stance, and if you want to take it then you're free to do so. But that's objectively what you're proposing in practice, and personally I disagree with it.
> Do you think you'd get away with training an AI on a bunch of animated Disney movies, and asking it to generate new images in that style, and using the result in commercial endeavors?
Yes.
Just the same as if I'd draw an image in the style of an animated Disney movie by hand.
In both cases I'll be sued for trademark infringement if the image's of the Mickey Mouse though.
In many cases the current "inequality" of how law is applied to individuals and to megacorporations has little to do with the law itself, and everything to do with how rich the megacorporation is. Try to set up an apple orchard and pick an apple as a logo[1] and tell me how it goes. The law explicitly states that another company, say one which produces computers instead of actual apples, has no merit here, but alas they have deep pockets, so here we are.
[1]: https://www.wired.co.uk/article/apple-vs-apples-trademark-ba...
> Question your assumptions about the world that results from requiring AI to respect Open Source licensing and other small copyright holders such as independent artists or online comment/story authors. It's not a corporate dystopia. It's a level playing field.
Well, let's see, for the sake of argument let's assume that the current widely believed legal status quo is true. (That is, that you can train a model on any data regardless of copyright because it's fair use. Although in my country that's explicitly allowed by law so here we don't have to assume anything.) Right now OpenAI can scrape 1TB of data off the Internet and legally train a model. I can also scrape 1TB of data off the Internet and legally train a model. And it can be any data, not just open source programs and content produced by small copyright holders. Is this not a level playing field?
Are you seriously suggesting that having to pay billions of dollars to license the training data necessary to train a model is a level playing field? I guess if nobody will be able to do it then it will be, in a way, a level playing field; I just fear that entities with enough money will be able to license enough data anyway and then the rest of us will end up with nothing.