The tech scales, but accessing the training data is a real problem. It's not like scraping the whole internet. And most of it is unlabeled.
When I do write something up, it is usually very finalized at that time; the process of getting to that point is not recorded.
The models maybe need more naturalistic data and more data from working things out.
Scale is not always about trougput. You can be constrained by many things, in this case, data.