Isn't this basically the entirety of the latest AI craze? They basically took a public good - the information available on the Internet - and hid behind some thin veneer of "we are not stealing, we just trained an AI on the information" and then they sell it. Note, I'm intentionally not writing "free information available on the Internet", because information is not free. Someone has to pay (in time or money) to generate it and host it. They might have provided it gratis to the public, but nobody asked them if an AI can come along, harvest it all and regurgitate it without a hint of reference to the original source.
Much of that information is not even free in the monetary sense, it is supported by ads. The AI will not only not click through the \ds, it won't even generate repeat traffic as once the information is harvested, there's no need to access the source anymore.
If you really think about it, it's a brilliant business model. It's a perfect theft, where the affected group is too diffuse and uncoordinated, it's extremely difficult to prove anything anyway, and the "thieves" are flush with investment capital so they sleep well at night.
LLMs have undoubtedly great utility as a research tool and I'm not at all against them. I think they (or a model similar in objectives) are the next step in accessing the knowledge humanity has amassed. However, there's a distinct danger that they will simply suck they sources dry and leave the internet itself even more of a wasteland than it has already become. I have no illusions that AI companies will simply regress to the lowest cost solution of simply not giving anything back to whoever created the information in the first place. The fact that they are cutting off the branch that they are sitting on is irrelevant for them, because the current crop of owners will be long gone with their billions by the time the branch snaps.