If you have the firehose, couldn't you just only sample every 10th tweet and make a "decahose"[1] of your own to work with?
[1] Shouldn't this be "decihose" since it's 1/10 not 10x?
If you have the firehose, couldn't you just only sample every 10th tweet and make a "decahose"[1] of your own to work with?
[1] Shouldn't this be "decihose" since it's 1/10 not 10x?
Tweets come with a lot more data than that.
Tweet object model: https://developer.twitter.com/en/docs/twitter-api/data-dicti...
Example payloads: https://developer.twitter.com/en/docs/twitter-api/data-dicti...
I’m boggled by the extent to which people mind-read Musk through his tweets. To me it’s obvious that his Twitter is an outlet for stress busting and banter, like smoke breaks.
500 million times 1KB is... still one USB stick, not even $50 https://amazon.com/SanDisk-512GB-Ultra-Flash-Drive/dp/B083ZL...
Or going by the retweeted quote tweet that's 40% bigger zipped, it's still well within 1TB which is a normal size for a flash drive.
I'm not very fond of Musk but this is a weak criticism.
That's like saying it's easy to check if a picture shows a bird because it's only 100kB.
What kind of system architecture would be good to load & search through such dataset. We tried to explore Mongo Atlas but it is coming out very expensive. Other alternative is to throw away most metadata & just keep ID & Tweet.
By the way, you have to pay for this access or there are some other options available? It is interesting to ingest all this data by ourselves
Concerning ingestion costs, I think it took me ~10 hours of 4vCPU vm last time, when I loaded 1.5 TB data to ClickHouse from this dataset: http://toddwschneider.com/posts/analyzing-1-1-billion-nyc-ta... 1.1 billion rides became 3.5 billion already.