S3 is basically a database. You can store everything there and interact with it via SQL with Athena, for example. I could see utilizing all of Amazon's "infinitely scalable" services to build very simple and powerful software really easily.
I start my data as .csv.gz but the first step is a CTAS to extract columns and convert to compressed parquet. This step basically costs the most but gives a 10x data size reduction to downstream steps.
Athena does not work at all if you perform large numbers of small indexed read queries, definitely use a traditional database for that.