They will either pay for it to be generated or get good enough at producing synthetic data that actually improves LLM quality.
Given how much data they need that will be pretty expensive, I mean really really expensive. How many people can write good training data and how much per day?
Doesn’t sound sustainable.