> Why would an AI company start scraping twitter html, instead of using an already existing archive?
I can think of a few possible reasons. They might want more up-to-date info, or they might have no real developers and the scraper was created by a business guru who prompted ChatGPT and didn't understand the code that came out.
Given what else Musk has asserted about Twitter, and how often former or current Twitter devs have contradicted him, it may not even be what Musk said.
> Twitter is not higher quality data than any other web page
Eh, depends how much you can infer from retweet, favourites, etc.
Won't be the only such site, but it's probably better training data than blog posts are these days.
But yeah, I absolutely agree that Twitter doing this caused a lot of damage to any orgs, corporate or government, which wanted to be public, anything from restaurants announcing special offers to governments issuing hurricane warnings. Twitter isn't big enough to assume everyone has an account, like Facebook is.