Genuine question: We signed up for the deci/decahose last month and for now making everyday dumps of about 50GB in flat files.
What kind of system architecture would be good to load & search through such dataset. We tried to explore Mongo Atlas but it is coming out very expensive. Other alternative is to throw away most metadata & just keep ID & Tweet.