This. If you can fit it on your soldered-in SSD, I'm not sure it really counts as "Big Data".
aws s3 cp --recursive bunch-o-data s3://some-bucket/
spark-ec2 --region eu-west-1 --identity-file s.pem --key-pair=spark --instance-type m3.2xlarge --slaves 40 -v 1.5.2 launch my-cluster
Is significantly easier than making PG work easily at the 100GB scale in my experience. spark-ec2 is a script that ships with Spark to make it easy to set up a cluster in AWS.[0] http://spark.apache.org/releases/spark-release-2-0-0.html#re...