Well no, you're also paying them for having done the work to "acquire" that data. That acquisition arguably amounts to the greatest theft in history.
Because I hate to break it to you, they could have zero drop in quality by just not incorporating US data...
If we're talking about copyright, why are we somehow entitled to profits derived from stealing Taylor Swift's IP? Why do you get a cut of AI derivatives and not get half of her wealth, too, directly?
Microsoft has trained models entirely on synthetic and public data with SotA results.
This is so false and unsupportable it's comical. The same goes the other way, if you claim they would use no value by only incorporating American data.
Also, much of the data used to train LLMs are not strictly public domain. For example, copyrighted books and source code with attribution-requiring licenses feature heavily in many corpuses. There are still pending lawsuits against the labs here, yet they continue to push forward. It’s no surprise that there is popular demand for redistribution.
Don't fall for the great lie of intellectual "property".
It was 1 part of an observation of opposed forces. “On the one hand information wants to be expensive, because it's so valuable. The right information in the right place just changes your life. On the other hand, information wants to be free, because the cost of getting it out is getting lower and lower all the time. So you have these two fighting against each other.”[1]
Some people removed the context and used it to say most information should be available to all. LLMs are information.
You thought this question proved what?