Besides, right now there's a huge need for analytical engines on top of HDFS (data lake or whatever you call it) - that's why there are all of these: (shamelessly stolen from an email from drill users):
- Apache Hive <https://hive.apache.org> (SQL-like, with interactive SQL thanks to the Stinger initiative) - Apache Drill <http://drill.apache.org> (ANSI SQL support) - Apache Spark <https://spark.apache.org> (Spark SQL <https://spark.apache.org/sql>, queries only, add data via Hive, RDD <https://spark.apache.org/docs/latest/api/scala/index.html#or... or Parquet <http://parquet.io/>) - Apache Phoenix <http://phoenix.apache.org> (built atop Apache HBase <http://hbase.apache.org>, lacks full transaction <http://en.wikipedia.org/wiki/Database_transaction> support, relational operators <http://en.wikipedia.org/wiki/Relational_operators> and some built-in functions) - Cloudera Impala <http://www.cloudera.com/content/cloudera/en/products-and-ser... (significant HiveQL support, some SQL language support, no support for indexes on its tables, importantly missing DELETE, UPDATE and INTERSECT; amongst others) - Presto <https://github.com/facebook/presto> from Facebook (can query Hive, Cassandra <http://cassandra.apache.org>, relational DBs &etc. Doesn't seem to be designed for low-latency responses across small clusters, or support UPDATE operations. It is optimized for data warehousing or analytics¹ <http://prestodb.io/docs/current/overview/use-cases.html>) - SQL-Hadoop <https://www.mapr.com/why-hadoop/sql-hadoop> via MapR community edition <https://www.mapr.com/products/hadoop-download> (seems to be a packaging of Hive, HP Vertica <http://www.vertica.com/hp-vertica-products/sqlonhadoop>, SparkSQL, Drill and a native ODBC wrapper <http://package.mapr.com/tools/MapR-ODBC/MapR_ODBC>) - Apache Kylin <http://www.kylin.io> from Ebay (provides an SQL interface and multi-dimensional analysis [OLAP <http://en.wikipedia.org/wiki/OLAP>], "… offers ANSI SQL on Hadoop and supports most ANSI SQL query functions". It depends on HDFS, MapReduce, Hive and HBase; and seems targeted at very large data-sets though maintains low query latency) - Apache Tajo <http://tajo.apache.org> (ANSI/ISO SQL standard compliance with JDBC <http://en.wikipedia.org/wiki/JDBC> driver support [benchmarks against Hive and Impala <http://blogs.gartner.com/nick-heudecker/apache-tajo-enters-t... ]) - Cascading <http://en.wikipedia.org/wiki/Cascading_%28software%29>'s Lingual <http://docs.cascading.org/lingual/1.0/>² <http://docs.cascading.org/lingual/1.0/#sql-support> ("Lingual provides JDBC Drivers, a SQL command shell, and a catalog manager for publishing files [or any resource] as schemas and tables.")