Yeah but can you store same amount of data on one ssd as on 10 disks?
Our HBase cluster had 6 * 2Tb disks: about 8.5Tb of usable storage (the other 3.5Tb accounts for data replication/duplication) per host in the cluster. However, you need about 200 bytes in memory per kb on disk and should assign only 32Gb of heap to HBase. That's 2.5Tb wasted, per host. Couldn't just plug those disks out and use them somewhere else: you need all the disks in parallel to overcome the IO/bandwith bottleneck.
aws s3 cp --recursive bunch-o-data s3://some-bucket/
spark-ec2 --region eu-west-1 --identity-file s.pem --key-pair=spark --instance-type m3.2xlarge --slaves 40 -v 1.5.2 launch my-cluster
Is significantly easier than making PG work easily at the 100GB scale in my experience. spark-ec2 is a script that ships with Spark to make it easy to set up a cluster in AWS.[0] http://spark.apache.org/releases/spark-release-2-0-0.html#re...