100GB can be handled without problems by mysql or postgre on a single node (yeah maybe not a laptop).
You need to tune your database and take the right design decisions, but it's still less work than setting up a distributed system.
aws s3 cp --recursive bunch-o-data s3://some-bucket/
spark-ec2 --region eu-west-1 --identity-file s.pem --key-pair=spark --instance-type m3.2xlarge --slaves 40 -v 1.5.2 launch my-cluster
Is significantly easier than making PG work easily at the 100GB scale in my experience. spark-ec2 is a script that ships with Spark to make it easy to set up a cluster in AWS.[0] http://spark.apache.org/releases/spark-release-2-0-0.html#re...