Apache Drill's Future
mail-archives.apache.org
mail-archives.apache.org
However, the competion was fierce and each Big Data vendor (MapR, Cloudera and HortonWorks) was pushing its own solution: Drill, Impala and Hive on Tez. Competion is always a good thing, but it fragmented the user base too much so no clear winner emerged.
At the same time, Spark SQL got sufficiently better to replace these tools in most use cases and Presto (from Facebook) got the traction and the user base that none of these projects had by being vendor agnostic (and its adoption by AWS in Athena and EMR also helped boost its popularity).
The big thing it adds is that it isn't stuck to any storage format. It has a connector that lets you load data into it from basically anything, from mysql dbs to hdfs files to whatever. So you can do cross database joins and just not care about where the data lives. You can also output to almost any database too.
Impala, Presto etc don't fit that model at all - they follow the Volcano model.
In this mode, they are not pure functional - if a task fails, there is no way to reproduce the output of that task.
The func within the map() was guaranteed to produce the same output for the same input across multiple attempts on failure or concurrently (for speculation).
Because of this, they can be faster as they do not wait for a task to be complete to run a subsequent stage & can pipeline better, but at the cost of failing all queries running on a node during a crash.
There are no retries for anything. This was deemed acceptable, if your hardware is reliable and the response to a failed query is just to "run it again", rather than per-node query recovery.
The reason for proper node failure tolerance for Spark/Tez/Flink etc are because they follow the functional model as closely as possible with exceptions for non-deterministic functions (say, UUID() in a SQL call).
The advantage of the failure tolerance is that these tools can push the whole cluster towards a single query performance when it is otherwise idle, because preemption can recover capacity out of a running system, if a higher priority query enters the system at a later point in time.
Edit: thinking about it more, how could Presto accomplish joins and other SQL operations without a shuffle? Seems very similar to Spark SQL, which is just syntactic sugar for multi-stage MapReduces.
Then it does a bunch of techniques, one of witch is map reduce, to do computation.
Its like you took a distributed SQL server and made it never have its own storage outside of memory. Instead it operates on wherever the data is living.
TIL that Presto is available in EMR.
Presto does not manage storage itself, but instead focuses on fronting those data sources with a single access point, with the option to federate (join) different sources in a single query.
Hive makes a terrible data warehouse no matter how SQL compatible it is.
Unfortunately to date, distributed queries will fail if the paths _and_ files are not symmetric in name - all file paths and names must exist on all nodes - therefore the "in situ" approach is not available. It appears the project focused on querying distributed file systems like HDFS and S3 and therefore had a lot of competition.
I hope some group picks up where HPE orphaned Drill after the MapR acquisition and pivots to a pure distributed worker approach. Running a drillbit on nodes where the data originates could be useful, the original example was SQL over http logs directly from webservers.
[1] https://mapr.com/blog/drill-your-big-data-today-apache-drill...
The reality is that nowadays both SparkSQL and Presto are way behind Hive, in terms of both speed and maturity. Hive made tremendous progress since 2015 (with the introduction of LLAP), while SparkSQL still has the issue of stability of fault tolerance and shuffling. (Presto does not support fault tolerance.) So, IMO, SparkSQL is nowhere near ready to replace Hive.
If you are curious about the performance of these systems, see [1] and [2] which compare Hive, SparkSQL, and Presto. Disclaimer: We are developing MR3 mentioned in the articles. However, we tried to make a fair comparison in the performance evalaution.
[1] https://mr3.postech.ac.kr/blog/2019/11/07/sparksql2.3.2-0.10... [2] https://mr3.postech.ac.kr/blog/2019/08/22/comparison-presto3...
On the other hand, our Presto cluster runs pretty much anything we throw at it, and when it fails, the failures are easier to anticipate and mitigate. It's also quite simple to deploy and operate.
I think that now, maintaining that compatibility is less of a need for Spark and Hive has introduced a lot of goodies in the meantime, so there might not be a need for the SQL flavors to be in lockstep anymore.
Apart from adding new features (e.g., ACID support), a lot of effort is still put into optimizing queries and runtime. In essence, Hive is a tool specialized for SQL, so it tries to implement all the optimizations you can think of in the context of executing SQL queries. For example, Hive implements vectorized execution, whose counterpart in Spark was implemented only in a later version (with the introduction on Tungsten IIRC). Hive even supports query re-execution: if a query fails with a fatal error like OOM, Hive re-generates a new query after analyzing the runtime statistics collected by then. The second query usually runs much faster, and you can also update the column statistics in Metastore.
In contrast, Spark is a general-purpose execution engine where SparkSQL is just an application. I remember someone comparing Spark to Swiss army knife, which enables you to do a lot of things easily, but is no match against specialized tools developed for a particular task. My (opinionated) guess is that SparkSQL will be replaced by Hive and Presto, and Spark streaming will be replaced by Flink.
Hadoop 1.x (i.e. "MapReduce" execution engine):
* Apache Pig
* Apache Hive
* Apache Drill
* Cloudera Impala
If I recall correctly, neither Drill nor Impala actually used Hadoop 1.x MapReduce as the execution engine, and were mostly bundled to read data commonly stored in the HDFS cluster.
Hadoop 2.x (i.e. the MR2 / "YARN" era):
* Apache Pig
* Apache Hive
* Apache Tez (technically a substitute execution engine for MR2), built to allow containers to persist, optimize coalesce operation/task stages, amongst other things to reduce overall job latency
* Apache Spark (technically a substitute execution engine for MR2)
Spark entered the Hadoop ecosystem, as many people were storing their data in HDFS, and the Hadoop 2 YARN resource model/containers provided the compute resources to run Spark as an execution engine, in lieu of MR2. You could and can also run a separate Spark-dedicated cluster, but many people were already running Hadoop and storing their data in HDFS clusters.
"Shark" became SparkSQL somewhere around Spark 1.3-1.4x? and Schema-ed RDDs evolved to DataFrames and better enabled people to reason and interface with their data in a table-like manner. Python/PySpark performance also rapidly improved from things like Project Tungsten and DataFrames.
https://databricks.com/blog/2015/02/17/introducing-dataframe...
Post-Hadoop / MR2:
* Hive
* Spark
* Presto
Tez was very much backed by Hortonworks, as part of their HDP Hadoop distribution, and motivated improve the performance of existing Apache Pig and Hive tools (major contributors from Yahoo, Microsoft, Hortonworks). Hortonworks later incorporated Spark as part of their distribution.
Spark was adopted by Cloudera as part of their CDH Hadoop distribution, and coexisted with Impala.
Tn the post-Hadoop / post-Spark world, both Hortonworks and Cloudera merged as well:
https://www.cloudera.com/about/news-and-blogs/press-releases...
Also since we're talking MPP withSQL/SQL-like dialects, we may as well mention that Greenplum, ParAccel/Redshift also coexisted with all of these.
https://stackoverflow.com/questions/53533506/what-is-the-dif...
https://github.com/apache/arrow/commit/e6905effbb9383afd2423...
And the platform/tools side of Drill now lives on as Dremio, which uses Apache Arrow.
https://github.com/dremio/dremio-oss
So the essence of Drill still lives, but it became half Apache project and half vendor controlled and supported, and the root of that split is now orphaned.
As far as I can tell, the implication is that there are now fewer than three people interested enough to participate in code reviews, and ASF rules require at least three +1 votes for basically anything to happen.
What the ASF won't let you do if you can't muster the votes is actually make a release — that takes 3 votes from people on the Drill PMC (Project Management Committee). If you can't get 3 PMC votes, the project cannot even make security releases and must be retired.
As to who can commit code, from the ASF's standpoint any person with commit rights can do so at any time. However, the project may impose additional constraints, such as requiring a code review.
Projects can amend their bylaws around this, but that also goes to the PMC - lazy consensus can apply, but yeah you can't go through the source release process without 3 binding votes.
Eventually, you realize that this sort of software is not anybody's hobby project and takes real dollars to pay for development.
Doesn't help that Dremio would profit from Drill collapsing and leaving that space open.
Drill: https://en.wikipedia.org/wiki/Apache_Drill
MapR sold to Hewlett-Packard Enterprise (HPE): https://en.wikipedia.org/wiki/MapR
If you're interested in Drill, check out OctoSQL[0]. It shares the same vision of querying multiple datasources using pure SQL, and pushing down as much operations as possible to the underlying datasource.
Moreover, there's been a huge rewrite under way this past year, ready to use on the master branch, yet unreleased however (will be available soon).
It adds Kafka and Parquet support and most importantly first-class unbounded stream support. Including temporal SQL extensions for working with event time metadata (instead of system time) so you can use stuff like live updated time window aggregations on incoming kafka streams. It also now uses on-disk badger storage as the primary way to store its state, so you can do Group Bys / Joins with lots of keys, and restarts of OctoSQL won't alter the final result (exactly-once semantics).
Make sure to check it out, it's also very simple to get going locally!
Disclosure: I'm one of the main contributors.
There's no distributed execution yet.
Easy to get going locally.
I'm not sure but I think Presto is in-memory? We're optimizing for SSD disks to easily support big states, and achieve durability this way. SSD disks are still plenty fast.
Definitely better streaming support. We're very much concentrating on good streaming ergonomics. OctoSQL isn't a batch execution engine at its heart, it's a streaming one.
Firstly, to paraphrase Monty Python, "Drill's not dead yet." There are some efforts to get corporate sponsorship, but it will take time. HPE's withdrawal was expected, but disappointing none the less. With that said, we are still gearing up to release Drill 1.18 which will have a considerable amount of enhancements, including new formats (SPSS, HDF5, possibly SAS) as well as new storage plugins to enable Drill to connect to Druid as well as REST APIs.
Personally, I've always felt that Drill was marketed to the wrong audience. As everyone notes, there are many competitors in the big data analytics space: Spark, ES, Splunk, Presto, Impala to name a few. Where I see Drill as filling a rather unique niche in the market is the small-to medium size data analytics with complex data.
For instance, if you have a CSV file, you use Excel. If you have 100 CSVs you have to code. Or what if you have an Excel spreadsheet and you want to pull data from your corporate reference API, which happens to return JSON and uses OAUTH authentication? Drill is the kind of tool that can bridge this gap and allow analysts to rapidly get value out of these situations.
As an example, here's a demo of me building a COVID dashboard from a REST API and spreadsheet in about 15 min with zero data prep: https://youtu.be/oEOhFWm3D9A
Another example: incident response. Let's say you have a PCAP file, use Wireshark. What if you have 10GB of PCAP files? Most likely you'll have to code up some solution. With Drill however, you can query that w/o coding, which means that you can get the value out of this data faster.
I know there are people using Drill. I know when I demo Drill to analysts, they love it. (Ok.. I'm biased here but I would say that's an accurate representation of the response to my presentations) What Drill lost was an active developer community. I hope over the next few months, that Drill's users will step up a bit and contribute to code reviews and/or actual code contributions. I've been thinking about creating a security focused fork of Drill as well so we'll see what happens.
If you have ideas/comments/questions, you can email me at cgivre@apache.org.
https://projects.apache.org/projects.html?category#big-data
Were you looking for a single project for all your big-data needs?
What you see on apache.org is what gets put in presentations at say, the Hadoop Summit.
edit: On a more helpful note, Apache Spark is probably as close as you can get to a single project for all your big data needs if that is what one wants out-of-the-box from an open-source project. It includes a SQL framework, streaming framework, either bundles or improves upon more general work done in Hadoop, etc. It can be pretty vendor-controlled at times, but it's birth was in academia, making it pretty different from the other projects that were mostly born as components in already established commercial platforms. There are pros and cons to that, of course.
Isn't awk enough? :)
For anyone curious about Apache Drill, it was inspired by Dremel which was used internally at Google and once the engine beneath GCP's BigQuery.
https://static.googleusercontent.com/media/research.google.c...
In the era of Hadoop, it was most closely aligned with the MapR distribution, while Hortonworks aligned with Hive, and Cloudera offered Impala as their solution.
History of Apache Drill: http://radar.oreilly.com/2015/09/apache-drill-tracking-its-h...
Is there a white paper describing what replaced Dremel in BigQuery's architecture?
Here's some interesting reads on BigQuery: https://cloud.google.com/files/BigQueryTechnicalWP.pdf
Also, Happy 10th Birthday, BigQuery! https://cloud.google.com/blog/products/data-analytics/bigque...
import pandas as pd
df = pd.read_csv('data/us_presidents.csv')
df.to_parquet('tmp/us_presidents.parquet')
The only faster/better CSV parser I've used with any frequency is the fread function in the R package data.table. When I used R in the past, Parquet was much less popular, but I think now the arrow package supports writing to Parquet.
This article explains how to convert CSVs to Parquet with Go and Scala: https://mungingdata.com/go/csv-to-parquet/
If you're converting hundreds/thousands of files, Spark/PySpark is probably the best tool for the job. For fewer files, Python or Go is just fine.
https://research.google/pubs/pub36632/
Wondering if maybe Spark REPL or Apache Zeppelin might be a decent replacement for Drill.
https://spark.apache.org/docs/latest/sql-distributed-sql-eng...
From an analytics perspective, a lot of people like the ability to connect to a JDBC or ODBC source.