The biggest problem OAI has is that they don't own a data source. Meta, Google, and X all have existing platforms for sourcing real time data at global scale. OAI has ChatGPT, which gives them some unique data, but it is tiny and very limited compared to what their competitors have.
LLMs trained on open data will regress because there is too much LLM generated slop polluting the corpus now. In order for models to improve and adapt to current events they need fresh human created data, which requires a mechanism to separate human from AI content, which requires owning a platform where content is created, so that you can deploy surveillance tools to correctly identify human created content.