All the RL data are exactly public. There are huge amount of distilled data freely available, and that amount is more than enough to train a ~10T model.
Nope, because the big AI companies are paying billions for it. They wouldn't pay anything for public data.
Subscription engineering is a deep field. Neither OpenAI nor Anthropic have any technical advantage in this field.