Also, very curious to know (from any Googlers browsing HN) if Dremel is still the state-of-the-art within Google, or if there is already a newer replacement.
Also, very curious to know (from any Googlers browsing HN) if Dremel is still the state-of-the-art within Google, or if there is already a newer replacement.
[0] http://research.google.com/pubs/pub41344.html [1] http://shark.cs.berkeley.edu [2] http://blinkdb.org/
Hadoop/Hive was not focused on speed but proving it is even possible to do queries reliably on such large datasets. Once this was achieved it immediately became clear the long waits for Hadoop/Hive batch jobs to finish made it impractical for many uses. Presto, Impala, Drill, RedShift, etc were all designed Primarily to address this problem and be much, much faster than Hadoop/Hive so that the data could be queried interactively.
All these new projects/products are in a very active competition to find the best way or ways to do this. You should compare RedShift to these other projects rather than Hadoop if speed is an issue for you. Each has it's pluses and minuses depending on the situation.
Impala, Drill, etc. avoid all those unnecessary reads and writes by implementing the querying logic directly, rather than by compiling to Map Reduce.
Shark [0] is an interesting counterpoint. It takes essentially the same approach as Hive but on Spark instead of Hadoop, and achieves similar or better performance than the more "direct" implementations.
For example as of today RedShift can hold a maximum of 256 terrabytes of compressed data while Facebook's Hadoop cluster was over 200 Petabytes in late 2012. RedShift only supports limited query and data types and a single index while Hadoop can theoretically handle arbitrary data processing. But if these constraints are acceptable then RedShift will likely be orders of magnitude faster in most cases.
Other projects/products will have different tradeoffs but they are almost always faster as this was almost always the primary goal.
Amazon Redshift enables you to start with as little
as a single 2TB XL node and scale up all the way to
a hundred 16TB 8XL nodes for 1.6PB of compressed user data.
from http://aws.amazon.com/redshift/features-and-benefits/The 16 node limit I was familiar with is not a hard limit. You can request more nodes.
http://aws.amazon.com/redshift/faqs/#0080
It would be interesting to see performance comparisons on these huge datasets. I would expect to see new and interesting problems at that scale.
Redshift looks to be order-of-magnitude faster than Impala or Shark in all the test. Does this mean that once RedShift supports user-defined functions, there is no competing solution that is any match? (Unless you want to avoid using the cloud)
Rcfile or parquet would be a more interesting benchmark.
In the "What's next?" section, they say they want to re-do the Impala tests using Parquet, which is a columnar format based on the Dremel whitepaper (http://parquet.io/).
Parquet was a joint effort between Cloudera and Twitter, and now it's being developed by many other companies. You can use it with Hive, Pig, MapReduce, Cascading, Crunch and I think Apache Drill's first milestone has adopted it as a columnar format as well. Parquet also allows you to use your Avro or Thrift schema (soon Protobuffs) to write Parquet data, too.
It's a separate project in the ecosystem and has its own roadmap (https://github.com/Parquet/parquet-mr).
"Here from HackerNews? This was originally posted several months ago. Check back in two weeks for an updated benchmark including newer versions of Hive, Impala, and Shark."
[0] http://research.microsoft.com/en-us/projects/dryad/ [1] http://research.microsoft.com/en-us/um/people/jrzhou/pub/Sco...