Setting a new world record in CloudSort with Apache Spark
databricks.com
databricks.com
This is a great task for a rigorous academic paper!
Disclaimer: I wrote the blog post.
Spark is a way of processing data, ideally stored in a system of record (Hive/HDFS/S3/MemSQL etc).
They're not the same.
For the kinds of processing both Spark and MemSQL do (e.g. join operation) is Spark faster than MemSQL?
Old record was 330 r3.4xlarge machines. Which actually cost very similar today. It looks like the old record used a LOT more RAM than this one. 40,260GB vs 3,152GB and similar CPU 5,280 vs 4,728 cores. Although having the RAM doesn't mean it was used, but unless they were using all the RAM I don't see why they would use the more expensive r3 instances.
Edit: Follow on, I think the cost savings is definitely in the low RAM usage. Can't really get that many cores without more than 3 times as much RAM in AWS.
- https://cloud.google.com/blog/big-data/2016/02/history-of-ma...
When I asked why BigQuery doesn't do these sorts, the answer came straight from the post "Nobody really wants a huge globally-sorted output. We haven’t found a single use case for the problem as stated."
These accomplishments are awesome nevertheless!
Disclaimer: I'm Felipe Hoffa, and I work for Google (http://twitter.com/felipehoffa).
:)
Also, seeing how expensive it is to sort 100TB ($144) you have to wonder why it wouldn't be better to do it on your own hardware.
Hence my confusion.