Databricks Orchestration pipeline, AWS Step Functions - good examples of DAGs as a configuration.
Databricks Orchestration pipeline, AWS Step Functions - good examples of DAGs as a configuration.
v1 of the system built before I joined was Step Functions and the like. It gets hairy just as you say.
v2 I built and designed with the lead data engineer, we called it Coriaria originally. We're hoping/planning to open source it eventually, although it's a little wrapped around our company's internal needs & systems.
It chooses neither "config" strictly speaking nor "code" for the DAG, instead the primary representation/state is all in the PostgreSQL database which tracks the dataset dependencies and how each dataset is built. It's a DAG in PostgreSQL as well.
To make dataset creation and management easier, I also wrote a custom Terraform provider for Coriaria. This made migrating datasets into the new system dramatically faster. The provider is really nice, supports `terraform import` and all that. Currently we have it setup so that there are separate roles/accounts that can modify an existing dataset, but reading state only requires authentication, not authorization. This enables one team to depend on another team's dataset as an upstream data source for their datasets without granting permission to modify it or create a potentially stale copy of the dataset. Terraform's internal DAG representation of the resource dependencies is leveraged because "parent_datasets" references the upstream datasets directly, including the ones we don't build.
We're able to depend on datasets we don't build ourselves because the system has support for Glue catalog backends to track and register partition availability.
Currently, it builds most of the datasets using AWS Athena & S3, however this is abstracted over a single step function. There's no DAG of step functions, it's just a convenient wrapper for the Athena query execution.
The system also explicitly understands dataset builds and validations as separate steps. The dashboard makes it easy to trace the DAG and see which datasets are blocking a dataset build.
We're adding more integrations to it soon so that other ways of kicking off dataset builds and validations are available.
If people are interested in this I can begin lobbying for open sourcing the system. My colleague wanted to open source it as well.
All else fails, I'll rebuild it from scratch because I don't like the existing solutions for managing datasets. We've been calling it a data-flow orchestration system or ETL orchestration system, not sure what would be most meaningful to people.
I think the main caveat to this system is that I'm not sure how much use it'd be for streaming data pipelines, but it could manage the discretization of streaming into validated partitions wherever streamed data is sunk into. Our operating assumptions are that you want validated datasets to drive business decisions, not raw event data streamed in from Kafka. Making sure the right data is located in each daily (or hourly) partition is part of that validation.
Quick google shows that others have done things like this already...
[1] https://noise.getoto.net/2021/10/14/using-jsonpath-effective...
[2] https://aws.amazon.com/about-aws/whats-new/2022/01/aws-step-...
[3] https://docs.aws.amazon.com/step-functions/latest/dg/sfn-loc...
Makes me wonder if Amazon internally was using Step Functions, ran into issues trying to scale to larger graphs, realized multiple teams were using Airflow, and created the Managed Airflow service.
We use it for a complex pricing process that invokes 30-40 micro services that can take up to minutes per step.
Our core is open source https://github.com/MagnivOrg/magniv-core
We can set you up with our hosted if you would like to poke around!
I am part of one such team. We were using Windows Task Scheduler on a windows VM to run jobs, we figured it would be a nice idea to (dramatically) modernise and move to airflow, but we grossly underestimated the complexity, learning curve, and surrounding tools it requires. In the end we (data science team) didn't get a single production task up and running. The data engineers had much more success with it thought, probably because they dedicated much more time to it.
Will look forward to trying AWS Step Functions.