They have a very valuable user base (all kinds of world leaders for example), so the data is not the only valuable thing they have.
There are hundreds of millions of people on Twitter, and a few of them are very smart. I don’t see how that helps here though.
Userbase and their social networks and interactions is the data.
They don’t have much value from advertising point of view anymore.
The training methods are nothing secret, right? The architecture is well known.
Expecting the entire training dataset to be fully open is delusional.
Right, because its not like the training dataset was built off comments posted by all of us in the first place.
How ungrateful we are, to demand the ability to access what was unconsentually built off our hard work in the first place.
"How was Grok trained?
Like most LLM's today, Grok-1 was pre-trained by xAI on a variety of text data from publicly available sources from the Internet up to Q3 2023 and data sets reviewed and curated by AI Tutors who are human reviewers. Grok-1 has not been pre-trained on X data (including public X posts)"