What kind of system architecture would be good to load & search through such dataset. We tried to explore Mongo Atlas but it is coming out very expensive. Other alternative is to throw away most metadata & just keep ID & Tweet.
By the way, you have to pay for this access or there are some other options available? It is interesting to ingest all this data by ourselves
Concerning ingestion costs, I think it took me ~10 hours of 4vCPU vm last time, when I loaded 1.5 TB data to ClickHouse from this dataset: http://toddwschneider.com/posts/analyzing-1-1-billion-nyc-ta... 1.1 billion rides became 3.5 billion already.
Tweets come with a lot more data than that.
Tweet object model: https://developer.twitter.com/en/docs/twitter-api/data-dicti...
Example payloads: https://developer.twitter.com/en/docs/twitter-api/data-dicti...
500 million times 1KB is... still one USB stick, not even $50 https://amazon.com/SanDisk-512GB-Ultra-Flash-Drive/dp/B083ZL...
Or going by the retweeted quote tweet that's 40% bigger zipped, it's still well within 1TB which is a normal size for a flash drive.
I'm not very fond of Musk but this is a weak criticism.
That's like saying it's easy to check if a picture shows a bird because it's only 100kB.
I’m boggled by the extent to which people mind-read Musk through his tweets. To me it’s obvious that his Twitter is an outlet for stress busting and banter, like smoke breaks.