Can't you do all this same research in just a few seconds for like pennies, just by sending a few sql queries to Google's BigQuery? I'm pretty sure we've had stories about that here.
https://medium.com/google-cloud/github-on-bigquery-analyze-a...
https://medium.com/google-cloud/github-on-bigquery-analyze-a...
Spark would have a been a simple option to do this kind of processing, with less lines of code, and could also run on "spare compute". Same goes for the "How does one process 10 million JSON files taking up just over 1 TB of disk space in an S3 bucket?": there are appropriate file formats for storing and querying big datasets, text/json is simply the least efficient option and likely the cause of the "$2.50 USD per query" number...