> Meta: We stole so much stuff that, actually, we didn't steal any stuff
> Google: If things were different, things would be totally different
> Andreeson Horowitz: But we already spent so much money
> Microsoft: Think about how this would hurt the little guys, like us
> Anthropic: Fucking shut up about it
> Hugging Face: It's super legal but honestly it might not be idk
> StabilityAI: It's legal in places that aren't here
It's worth noting that, back when it suddenly became technically feasible to copy all songs for free, the response from the rightsholders of music was, essentially, to demand that we all collectively pretend it isn't. It might make sense for the rightsholders of training data to try to do the same, if they could speak with one voice and if they had lobbyists, but I guess the fact that they can't and don't settles that.
To me, training a model is less like redistribution and more like reading something and having it influence your thoughts. You may be able to reproduce certain memorable sections of things you've read verbatim (common with poetry, for example) and you may well be able to reproduce a 'style', but you don't have the entirety of the source material memorized in a way that it could be redistributed in full (but if you did, that may be well an infringement at such a time).