eg. importing a SAP feed into a database, or loading a bunch of csv files, or like processing a bunch of images...
...anything where you have to convert data from some source through a series of steps (typically a DAG) into some useful output.
However, its often misused.
For example, if you have a trivial amount of data, or trivial process an ETL is over engineered cruft where a simple script would do.
So there are many (rubbish) ‘simple ETL frameworks’ which offer zero value and technical complexity for no benefit.
Unless you need multiple servers processing data through multiple steps and you need the auditing and process control... you probably don’t need an ETL, just a simple script.
+1
> Unless you need multiple servers processing data through multiple steps and you need the auditing and process control
I'll stress the "multiple servers" part. You can add in a substantial amount of multiple, sequential steps and auditing and process control in a simple script. The part that adds orders of magnitude worth of complexity and operational overhead and points of failure is being able to distribute it to multiple servers. Distributed architectures are operationally and architecturally expensive. And far more often than not, completely unnecessary for a given use case.
While I love both Spark and Dataflow, both of them are incredibly complex distributed systems with very high operational costs. Someone, somewhere is paying a lot of money to have an operational resource maintain that complexity. Whether you have an internal devops resource doing so or you're using a managed service, you're paying for that complexity somehow. And, for a lot of workloads, you aren't actually getting any more value than you would from standing up a ~$50/month standard Debian/Ubuntu server and a set of simple scripts on it.
Not necessary. ETL these days can be streamed, realtime, etc.
The ETL frameworks (Airflow, Luigi, now Mara) help with this, allowing you to build dependency graphs in code, determine which dependencies are already satisfied, and process those which are not. They'll usually contain helper code for common ETL tasks, such as interacting with a database, writing to/reading from S3, or running shell scripts.
Mara data integration is indeed a glorified version of Make (with cost based scheduling and lots of visualizations)
An « ETL » framework is for me something that will allow you to express and run maintainable data pipelines, as a coder.
Kiba is exactly that, in the sense that it does not include built-in sources/transforms/destinations, but rather defines a set of guidelines you can follow to easily implement reusable ETL components with high quality & unit tests.
More information in the README: https://github.com/thbar/kiba/blob/master/README.md
Cron is awful for cloud-based deployments trying to keep data from multiple sources in sync and compute business-specific metrics over them.
ETL frameworks seek to address this pain by building a large stack of software around scheduling the running of and unifying the writing of these data munging tasks.
As such, they tend to be as awful as the sum of their dependencies minus some small constant quality factor. So, more awful than any one part but less awful than the storm you'd be howling into if you didn't take some sort of infrastructural solution.
Getting data from one place to another - usually the end place being a structured database (or it ought to be).
If you have been in the industry over 10 years you probably used Perl or Python scripts.