AFAIU neither of those are relevant to GPT-like architectures but it's not inconceivable to think there might be a model architecture in the future that takes advantage of those. Purely from a information theoretic POV, there's non-zero bits of information in the timestamp and relative ordering of tweets.
1) Facebook Posts/Comments, 2) Instagram Posts/Comments, 3) Youtube Comments, 4) Gmail content, 5) LinkedIn Comments, 6) TikTok contents / comments
X and Reddit are definitely valuable, but they're definitely not unique. I think Meta and Google have inherent advantages because their data is not accessible to LLM competitors and they have the actual capabilities to build great LLMs.
Unless X decides to tap AI talent in China, they're going to have a REALLY hard time spinning up a competitive LLM team compared to OpenAI, Google, and Meta, which I think are the top three LLM companies in that order.
I mean, could it be that it's just that the platforms you're familiar with are similar quality? There are major quality differences. Consider, for example, HN vs Instagram. Do you really see no difference in the quality of discourse, or do you just not use Instagram?
By bulk/raw data volume, I'd say that the vast majority of internet communication is the same quality, yeah-- I'll stick by that assertion. That's not at odds with acknowledging there exist locations where intelligent communication happens. My position is just that the signal to noise ratio is pretty bad in the majority of places.
Is probably LLM poison.
> 4) Gmail content
Is huge but also has enormous privacy issues. Most people by default assume their emails are reasonably private, whereas most people wouldn't assume their comments on these platforms are private.
https://www.youtube.com/@HyperspacePirate
https://www.youtube.com/@scottmanley
That's high quality content, timestamped and about current events.
There is very little content on Twitter that compared in quality to one will written news article.
I don't think Musk is the type of person to make the same mistake, so we'll either end up with a Twitter LLM that accurately represents the sum total of the Twitter firehose, and/or many derivative LLMs each having a set of, possibly orthogonal, biases. Honestly, I think the later is preferable and would represent the diversity of opinions in reality more accurately.
Given the data source, I think it will be important to be able to switch between LLM personalities in the future to get the "crowd truth".
We need an xkcd showing a conversation between twitter, reddit, and hacker news based LLMs. Political rage meets memes meets pedantry.
Current LLMs are trying to predict typical human prose from samples pulled from the internet. So it isn’t as if they are sacrificing quality for quantity. A bunch of text from the internet is a very good representation of typical human prose. Whether it is well written or the descriptions contained in the prose accurately represent, like, actual physical reality is another issue.
Maybe they want to predict something with, like, less dimensionality but more utility than a paragraph of fiction.
In which the distinction between "data" and "information" is crucial. Especially now that the "floodgates" have been re-opened regarding misinformation, bots, impersonators and the likes.
Data is crucial when in need of training body. But information is crucial when the training must be tuned, limited or just verified.
I guess X is harder to scrap without permission.
We introduce new datasets derived from the fol- lowing sources: PubMed Central, ArXiv, GitHub, the FreeLaw Project, Stack Exchange, the US Patent and Trademark Office, PubMed, Ubuntu IRC, HackerNews, YouTube, PhilPapers, and NIH ExPorter. We also introduce OpenWebText2 and BookCorpus2, which are extensions of the original OpenWebText (Gokaslan and Cohen, 2019) and BookCorpus (Zhu et al., 2015; Kobayashi, 2018) datasets, respectively.
From https://arxiv.org/abs/2101.00027 (The Pile: An 800GB Dataset of Diverse Text for Language Modeling)
And there's my incentive to stop posting on HN.
It's been a blast, guys. I'm going back to lurker mode.
Everybody smile for the camera, or we could just moon them, or both!
Its strength is freshness and volume, but I guess these can be achieved without Twitter if you have a strong web crawling infrastructure? Also, the current generation of LLM is not really capable of exploiting minute-level freshness... at least for now.
There’s not a lot of data in Twitter today resembling long-form content: essays, news articles, books, scientific papers, etc. That’s probably why Twitter/X expanded the tweet size limit, to be able to collect such data.