Lessons learned from running Apache Airflow at scale
shopify.engineering
shopify.engineering
Upgrades have been an absolute nightmare and so disruptive. The scalability improvements in airflow 2 were a boon for our runtimes since before we would often have 5-15 minutes of overhead between task scheduling, but man it was a bear of an upgrade. We've since tried multiple times to upgrade past the 2.0 release and hit issues every time, so we are just done with it. We'll stay at 2.0 until we eventually move off airflow altogether.
I stood up a prefect deployment for a hackathon and I found that it solved a ton of the issues with airflow (sane deployment options, not the insane file-based polling that airflow does). We looked into it ~1 year ago or so, I haven't heard a lot about it lately, I wonder if anyone has had success with it at scale.
Prefect very cleanly written, good design and flexible. IMHO it is a platform that will be the next big thing in the area.
How I know, I deployed prefect as a static config gathering system across 4000 servers, both Linux and Windows. No other software stack came close, as one of the core concepts of prefect is 'expect to fail'. Things like Ansible Tower die really quick with large clusters due to the normal number of failures and the incorrect assumption that most things will work (as you can for a small cluster).
I wish I got to use it in my current work but there is no use case. Yet.
Interesting use case, I use prefect for data pipelines, never thought about that kind of use case.
With prefect I created a task 'collect machine details for windows', another 'collect machine details for Linux', another 'collect software inventory'.
I have a list of machines in a database so I create a task to get them. That task is an sqlalchemy query so I can pass the task a filter.
I get a list of linux machines and pass that to a task to run. I get a list of windows machines and pass that to a task.
Note that the above don't depend on each other.
I have a task that filters good results from bad. I have another task that writes a list to a database.
Other tasks have credentials.
Another task puts errors to an error table, the machines that failed get filtered from the results and run into this task.
I plumb the above up with a Prefect flow and it builds a DAG that runs the flow. Everything that can be run in parallel does so, everything that has some other input waits for the input.
Tasks that fail can be retried by Prefect automatically. Intermediate results cached. And, I get a nice gui for everything. I can even schedule it in the gui.
Prefect 1.0 feels like second-gen Airflow--less boilerplate, easy dynamic DAGs, better execution defaults, great local dev, etc etc. It's more sane but you still feel the impedance mismatch from working with an orchestrator.
Prefect 2.0 is a first-principles rewrite that removes most of the friction from interacting with an orchestrator in the first place. Finally, your code can breathe.
The primary reason we run Airflow is because it can execute Python code natively, or other programs via Bash. It's very rare that a DAG I write is entirely SQL-based.
(T)transform is decoupled and would be served in set-based operations managed by dbt.
Luigi doesn't force you into using a central orchestrator for executing and tracking the workflows. Tracking and updating tasks state is open functions left to the programmer to fill in.
It's probably geared for more expert programmers who work close to the metal that don't care about GUIs as much as high degrees of control and flexibility.
It's one of those frameworks where the code that is not written is sort of a killer feature in itself. But definitely not for everyone.
The other orchestrator besides Toil to check out is Cromwell, but that uses WDL instead of Python for defining the DAG, and it's not a super powerful language, even if it hits exactly the needs for 99% of uses and does exactly the right sort of environment containment.
I'm also hugely underwhelmed by k8s and Mesos and all those "cloud" allocation schemes. I think that a big, dynamically sized Slurm cluster would probably serve a lot of people far better.
It essentially abstracts workflows as custom resources. Works phenomenally well and quite stable.
Feel free to ping at xavier.marcos@astronomer.io if you're open to a chat.
- Configuration as code. Configuration should be a way to change an application's behavior _without_ changing the code. Make me write a workflow as JSON or XML. If I need a for-loop, I'll write my own script to generate the JSON.
- It's complicated. You almost need a dedicated Airflow expert to handle minor version upgrades or figure out why a task isn't running when you think it should.
- Operators often just add an API layer over top of existing ones. For example, to start a transaction on Spanner, Google has a Python SDK with methods to call their API. But with Airflow, you need to figure out what _Airflow_ operator and method wraps the Google SDK method you're trying to call. Sometimes the operator author makes "helpful" (opinionated) changes that refactor or rewrite the native API.
I would love a framework that just orchestrates tasks (defined as a command + an image) according to a schedule, or based on the outcome of other tasks, and gives me a UI to view those outcomes and restart tasks, etc. And as configuration, not code!
We take inspiration from the Kubeflow project and run tasks as containers. With a GUI for editing pipelines and managing scheduled runs we come pretty close to what you’re asking for (bring an image and run a command). And it’s OSS, of course.
That's how you end with extreme mess in logs, UI and metrics.
> But with Airflow, you need to figure out what _Airflow_ operator and method wraps the Google SDK method you're trying to call.
Or you can use PythonOperator with hooks, that generally integrate external APIs with Airflow connection system.
https://github.com/apache/airflow/blob/main/airflow/provider...
I think the bigger problem with Airflow's Operator concept is N*M problem of integrating multiple systems. That's how you end with GoogleCloudStorageToS3Operator and stuff like that.
Not that I recommend it. It's quite lovely in principle, but really flawed in practice. It's YAML hell on top of Kubernetes hell (and I say that as someone who loves Kubernetes and uses it for everything every day).
Having worked with some of these tools, what I've started to wish for is a system where pipelines are written in just plain code. I'd like to run and debug my pipeline as a normal, compiled program that I can run on my own machine using the tools I already use to build software, including things like debuggers and unit testing tools. Then, when I'm really to put it into production, I want a super scalable scheduler to take my program and run it across dozens of autoscaling nodes in Kubernetes or whatever.
The only thing I've come across that uses this model is Temporal, but it's got a rather different execution model than a straightforward pipeline scheduler.
We use Argo at my work, because we made the “unit of work” just a container that’s run by K8s. There isn’t a Python runtime around to run this stuff, and that’s a good thing.
Same as with Pulumi, I’ve seen what devs do when you let them write arbitrary code anywhere-you end up with code that lives everywhere, confusingly cuts across different infrastructural layers with wild-abandon. It becomes difficult to on-board, difficult to maintain, the blast radius becomes unknown, and just because it’s Python, it will randomly break in 6 months because nobody pinned a transient dependency.
YAML is awful. Argo could do with some trimming in its configs, but I’ll take it, and other “constrained” alternatives over arbitrary-code-in-Python-that-controls-resources-that-run-more-arbitrary code.
OK, so you are using some service that expects JSON to config. You have a script to generate the JSON. Now you gotta make sure to run the script on changes. You also need to check in the JSON into your VCS despite it being a generate artifact. This leads to the perpetual "somebody fixes an issue in the artifact, and it gets wiped after someone uses the script" issues.
What you're saying is theoretically very good! But it doesn't play well with the combo of assumptions made everywhere in hosted services and some PL infrastructure:
- every project is 1 git repository, and vice versa
- commit security is an all or nothing proposition (so it's hard to have permissions like "this file can only be edited by this user" or "this commit can be made directly to `main`, _if_ it only touches this file")
- configuration as configuration
- version numbers have to be written into code
And obviously this is a problem, because everyone keeps on inventing YAML-based programming idioms for each of their services! Because static configuration doesn't map to even the simplest of setups very well.
People come and them mention stuff like Dhall. But really people want to be able to write programs that read files, make HTTP requests, and then make decisions about what to do based on that. And things like "if this test suite fails, do this other orchestration". The ideal CI flow is not staticly definable at the start of execution!
There is a killer protocol out there, and it's ... something like Skylark, with a standard library that gives file reading, checking of statuses, probably HTTP and the like.... and an executable that allows for configuring permissions in certain ways. The bonus would be just re-doing Bazel but with less guardrails. If someone showed up with CI/CD that was more in this vein they would become the dominate CD host very quickly in my opinion.
IMO I think GUI variables are a big weakness of Airflow.
I, simply, never want to be responsible for a typo in a prod config variable. Ever.
I'm all for version controlling every, single, variable.
What I've found while building out Shipyard, a hosted lightweight orchestration platform, is that teams want something that "just works". Servers that "just scale". Observability that doesn't require digging. Notifications and retries that work automatically. Workflows that don't mix business logic with platform logic. Code and workflows that sync with git. Deployment that only takes a few minutes.
For the straightforward use cases, where you need to run tasks A -> G daily, with a bit of branching logic, Airflow is overkill. Yes, Airflow has a lot of great complex functionality that can help you down the road. But Airflow keeps getting suggested to everyone even if it's not best suited to their use case, resulting in lots of lost time and engineering overhead.
While I have definitely have bias, there are a lot of other high quality alternatives out there to explore nowadays!
I’d also like to throw Matillion into the mix of “horror story tools”. I have far too many matillion-induced bad stories.
* Teammates making inscrutable ETL flows because it lets you make nice/visual based flows, and instead of laying them out sanely, they draw pictures with them. One was particularly fond of making flowers and fish, which-while amusing-was entirely unhelpful when there were pressing prod issues and you’re trying to follow the flow around some flower petals.
* nigh incomprehensible generated SQL
* a design that let users mix orchestration and transformation, so prior teammates created jobs that would stomp on each other’s outputs, because they didn’t orchestrate them sanely
* sometimes the runtime would just…stop for reasons we were still unable to resolve.
* lets you run arbitrary Python/bash/etc script, so people would put all sorts of wild scripts in the flow. Oh also, it’s not Python it’s actually Jython, and it could mutate variables in the shared environment-an old teammate would (ab)use this functionality to set the date-time of some variable that another ETL-flow would use to make a decision (see issues above about mixing orchestration + transformation).
As mentioned elsewhere, AWS step functions are really the best in orchestration.
Our teams have many 1 or 2 step DAGs that are idempotent. They could have been lambdas and they're already pulling from SQS already. It could be just my misfortune, but in AWS, MWAA is kind of janky. It's difficult to track down problems in the logs (task failures look fine there) and the Airflow UI is randomly unavailable ("document returns empty", "connection reset" kind of things).
There's certainly some functionality overlap, but I don't see Lambda and Airflow as competitors. Each has capabilities that the other doesn't.
AWS Step Functions is a proprietary service provided exclusively by AWS, which reacts to events from AWS services and calls AWS Lambdas.
Unless you're already neck-deep in AWS, and are already comfortable paying through the nose for trivial things you can run yourself for free, it's hardly appropriate to even bring up AWS Step Functions as a valid alternative. For instance, Shopify's articles explicitly mention they are running their services in Google Cloud. Would it be appropriate to tell them to just migrate their whole services to AWS just because you like AWS Step Functions?
It’s like those stackoverflow answers that tell the user to stop using PHP and rewrite it in Python or something.
But fault tolerant workflow engine is not trivial thing, it may cost you many engineer hours to build, monitor and maintain it, so outsourcing it to someone else is totally viable solution.
The complexity and risk of migrating cloud providers eclipses whatever problem you assign to "fault tolerant workflow engines".
Any mention of AWS Step Functions makes absolutely no sense at all and reads at best like a non-sequitur.
If you get a good deal from one cloud provider, you can get started quickly.
It's useful even for individuals such as students who get free credits from these providers: create a cluster and you're up and running in no time.
Our rationale was that we didn't wanted to be tied to one cloud provider.
“Oh I have to learn how to use and setup this tool? I think I’ll just pay the equivalent salaries and be locked in…”
For a one person shop like me, AWS is a force multiplier. With it, I can do (say) 30% of what a dedicated engineer in a specific role could do. Without it, I'd be doing 0%.
I really like this tradeoff for my particular situation.
I don't know how management works through this math, maybe managing people gets exhausting and they just want to out-source it so leadership doesn't have to deal with it and then they can just focus on the core product.
And the above "I don't want to deal with it" reason isn't spoken of, the more more commonly touted benefit is cloud's "flexibility". Sure, but this is actually _really_ expensive. Every cloud migration effort I've experienced is only just worthwhile to begin to talk about because the costs are based on long-term contracts of cloud resources, not the per-hour fees. Nice flexibility.
With that said, the cloud may be a good place for prototyping where the infrastructure isn't the core value add and it's uncertain. A start-up is a prototype and so here we are. But, for an established company to migrate to the cloud and fire the staff that's maintaining the on premise resources.. I'm skeptical. More than likely, this leads to maintaining both cloud and on premise resources, not firing anyone, and thus, actually increasing costs for an uncomfortably long time.
And for the folks on the ground, who don't pay the bills, the increase of accidental complexity is rather painful.
You won’t get around any of the problems of airflow by moving to a cloud offering.
They don't always make sense, in certain scenarios it is worth taking an open source, cloud independent tool, in some scenarios you can roll your own, but there are circumstances where it's a good choice using a tool your cloud provider gives you.
I've been a one-man army at places because leveraging these cloud offerings allows me to crank out working software that scales to the moon without much thought.
I'd rather pay AWS/GCP to handle infra, so that I can get 2-3x as many project done.
Why? Where else is this mentioned?
At our company, AWS Step is a disaster. You're effectively writing code in JSON/YAML. Anything beyond very simple steps becomes 2 pages of YAML that's very hard to read or write. There is no way to debug, polling is mostly unusable. Changes need to be deployed with CF which can take forever, or worse hang.
It's the most one of the most annoying technologies I've used in my 20+ years of engineering.
Thoughtworks made a case for this distinction in https://martinfowler.com/articles/cant-buy-integration.html#...
The author also mentions that it’s used for machine learning models which will ultimately feed back into Shopify’s front end, for instance.
* run job for given reporting_period * backfill data for reporting_period between N and M * lookup failed jobs
Nothing in Databricks supports that.
Backfills are on our roadmap We are previewing looking up failed jobs soon.
Email me at bilal dot aslam at Databricks dot com if you want more info
I wish they had a "start small", self service, clear pricing option.
Dagster was a great solution for us. Their notion of software defined assets allowed us to track metadata of the Redshift and Snowflake tables we were working with. Working with re-runs and partitioned data was a breeze. It did take a while to onboard the whole team and get things working smoothly, which was a bit difficult because Dagster is still young and they were often making changes to how parts of the system worked (although nothing that was immediately backwards incompatible).
We also enjoyed some of the out of the box features like resources and unit testing jobs. Overall, I think it made our team focus more on our data and what we wanted to do with it rather than feeling like we had to wrangle with Airflow just to get things running.
Dagster, on the other hand, seems to let you scale from using it locally as a library all the way to running on ECS/K8s etc. Along with that there's unfortunately a ton of complexity in setting it up but that's not Dagster's fault and it seems like Dagster works once you get it set up. Agree about it being young and there being some rough spots but it's got lots of good ideas. We were nearly done setting it up but got pulled off onto more urgent things, so I haven't run it in production yet. I'm glad to hear it worked well for you!
I'd love to hear more on this. I've not evaluated Prefect, and am currently keeping an eye on Dagster. What trade-offs does Prefect win?
What Dagster has going for it in this space is pragmatism. It really nails all of the problem points of data ops (with resources & sensors specifically). If I was consulting for a shop that needed data pipelines and they had good eng, I'd recommend Dagster in a heartbeat.
Key takeaways roughly those:
- Dagster's integration with k8s really shines as compared to Airflow, it is also based on extendable Python code so it is easy to add custom features to the k8s launcher if needed.
- It is super easy to scale UI/server component horizontally, and since DAGs were running as pods in k8s, there was no problem scaling those as well. For scheduling component it is more complicated, e.g. builtin scheduling primitives like sensors are not easily integrated with state-of-art message queue systems. We ended up writing custom scheduling component that was reading messages from Kafka and creating DAG runs via networked API. It was like 500 lines of Python including tests, and worked rock-solid.
- networked API is GraphQL while Airflow is REST, both are really straightforward, however in Dagster it felt better designed, maybe due to tighter governance of Dagster's authors over the design.
- DAG definition Python API, e.g. solid/pipeline, or op/graph in a newer Dagster API, is somewhat complicated as compared to Airflow's operators, however it is easy to build custom DSL on top of that. One would need custom DSL for complicated logic in Airflow as well, and in case of Dagster it felt easier to generate its primitives, than doing never ending operators combinations in case of Airflow.
- Unit and integration testing are much easier in Dagster, the authors put testing as a first-class citizen, so mocks are supported everywhere, and the code tested with local runner is guaranteed to execute in the same way on k8s launcher. We never had any problems with test environment drift.
The biggest caveat was full change of internal APIs in 0.13, which forced the team to execute a fairly complicated refactor, due to deprecation of the features we were depending on e.g. execution modes. Had we spent more time on Elementl slack, it would be easier to put less dependencies on those features ^__^
After having a lot of pain points with an (admittedly older and probably not best-practices) Airflow setup, I am now at a different job running similar types of workflows on Temporal - we're pretty happy with it so far, but haven't done anything crazy with it.
* Both Apache Airflow and Temporal are open source
* Both create workflows from code, but the approach is different. With Airflow, you write some code and then generate a DAG that Airflow can execute. With Temporal, your code is your workflow, which means you can use your standard tools for testing, debugging, and managing your code.
* With Airflow, you must write Python code. Temporal has SDKs for several languages, including Go, Java, TypeScript, and PHP. The Python SDK is already in beta and there's work underway for a .NET SDK.
* Airflow is pretty focused on the data pipeline use case, while Temporal is a more general solution for making code run reliably in an unreliable world. You can certainly run data pipeline workloads on Temporal, but those are a small fraction of what developers are doing with Temporal (more here: https://temporal.io/use-cases).
While my experience with those specific tools is limited, I have used other DAG-based systems (notably Apache Oozie) quite a bit. I can't think of anything I did with those that I couldn't do with Temporal.
Temporal is distinct in that it's neither limited to DAGs nor to data pipelines. I suppose it is therefore a super-set, even though I don't really think of it that way (much as I don't think of Go or Rust as a super-set of bash).
It seems notable to me that the big Prefect rewrite mentioned elsewhere [0] leans into the same "workflow" terminology that Temporal uses. I have to wonder if Prefect saw Temporal as superceding the DAG tools in coming years and this is them trying to head that off.
That post's discussion of DAG vs workflow also sounds a lot like why PyTorch was created and has seen so much success. Tensorflow was static graphs, pytorch gave us dynamism.
One data center company that I know of uses airflow at scale with docker and k8s. They have a huge team of devops just to manage the orchestrator. They in turn have to fine tune the orchestrator to run smoothly and efficiently. Similar to what shopify has noted here, they have built on top of and extended airflow to take care of pain points like point 4. For companies like this it makes sense to run airflow.
Another issue I see companies/engineers who adopt airflow is that they use it as a substitute for a script than as an orchestrator. For example, say you want to download files from an API, upload to s3, load it to your warehouse (say snowflake) and do some transformations to get your final table - instead of writing separate scripts for each step of fetch/upload/ingest/transform and call each step from the dag, they end up writing everything as a task in a dag. A huge disadvantage is there is a lot of code duplication. If you had a script as a CLI, all your dag/task has to do is call the script with the respective args. I agree that airflow comes with a lot of convenience wrappers to create tasks for many things but I feel this results in losing flexibility.
This also results in them tying their workflow with airflow and any change they might need they have to modify their airflow code directly. If you want to modify how/what you upload to s3, you end up writing/modifying python functions in the respective dags' code. This removes the flexibility to modify/substitute any component of the workflow with something else or even change the orchestrator from airflow to something else. Additionally, different teams might write workflows in different ways - standardization of practice is really hard. This in turn results in pouring more investments to maintaining and hiring "airflow data engineers". Companies fall into steep tech debts.
Prefect/dagster are new orchestrators in town. I'm yet to try them out but I've heard mixed reviews about them.
EDIT: Forgot about upgrades. Lot of upgrades are breaking changes esp the recent change from 1->2. You end up spending a lot of time just trying to debug what went wrong. Just installing and running it is a pain.
One of my biggest annoyances in the orchestration space is that teams are mixing business logic with platform logic, while still touting "lack of vendor lock-in" because it's open source. At the point that you're importing Airflow specific operators into your script and changing the underlying code to make sure it works for the platform (XCom, task decorators, etc.), you are directly locking yourself in and making edits down the road even more difficult.
While some of the other players do a better job, their method of "code as workflow" still results in the same problems, where workflows get built as a "mega-script" instead of as modular components.
I'm a co-founder at Shipyard, a light-weight hosted orchestrator for data teams. One of our core principles is "Your code should run the same locally as it does on our platform". That means 0 changes to your code.
You can define the workflow in a drag and drop editor or with YAML. Each task is it's own independent script. At runtime, we automatically containerize each task and spin up ephemeral file storage for the workflow, letting you can run scripts one after the other, each in their own virtual environment, while still sharing generated files as if you were running them on your local machine. In practice, that means that individual tasks can be updated (in app or through GitHub sync) without having to touch the entire workflow.
I'm biased, but it seems crazy to me that so many engineers are willing to spend hours fighting the configuration of their orchestration platform rather than focusing on the solving the problems at hand with code.
With great control comes great responsibility.
We package all the code in 1-2 different Docker images and then create the DAG. We've faced many issues (logs out of order, missing, random race conditions, random task failures, etc.)
But what annoys me the most is that for that 1 big DAG, the UI is completely useless, tree view has insane dupplication, graph view is super slow and hard to navigate through and answering basic questions like, what exactly failed and what nodes are around it are not easy.
Can you point me to the documentation on this function? It's possible I'm not using it correctly.
I can do something like this, which works locally, but breaks when deployed:
res = annotated_task_function(...)
res.operator.task_id = 'manually assigned task id'def my_func():
...
In more recent versions of Airflow, TaskGroups (https://airflow.apache.org/docs/apache-airflow/stable/concep..., https://www.astronomer.io/guides/task-groups/ ) were made to help this a little bit. Hopefully that helps a bit.
At ~1k nodes in the graph introspection becomes hard anyway, as others have suggested, breaking it down if possible might be a good idea.
Databricks Orchestration pipeline, AWS Step Functions - good examples of DAGs as a configuration.
I am part of one such team. We were using Windows Task Scheduler on a windows VM to run jobs, we figured it would be a nice idea to (dramatically) modernise and move to airflow, but we grossly underestimated the complexity, learning curve, and surrounding tools it requires. In the end we (data science team) didn't get a single production task up and running. The data engineers had much more success with it thought, probably because they dedicated much more time to it.
Will look forward to trying AWS Step Functions.
We use it for a complex pricing process that invokes 30-40 micro services that can take up to minutes per step.
Our core is open source https://github.com/MagnivOrg/magniv-core
We can set you up with our hosted if you would like to poke around!
v1 of the system built before I joined was Step Functions and the like. It gets hairy just as you say.
v2 I built and designed with the lead data engineer, we called it Coriaria originally. We're hoping/planning to open source it eventually, although it's a little wrapped around our company's internal needs & systems.
It chooses neither "config" strictly speaking nor "code" for the DAG, instead the primary representation/state is all in the PostgreSQL database which tracks the dataset dependencies and how each dataset is built. It's a DAG in PostgreSQL as well.
To make dataset creation and management easier, I also wrote a custom Terraform provider for Coriaria. This made migrating datasets into the new system dramatically faster. The provider is really nice, supports `terraform import` and all that. Currently we have it setup so that there are separate roles/accounts that can modify an existing dataset, but reading state only requires authentication, not authorization. This enables one team to depend on another team's dataset as an upstream data source for their datasets without granting permission to modify it or create a potentially stale copy of the dataset. Terraform's internal DAG representation of the resource dependencies is leveraged because "parent_datasets" references the upstream datasets directly, including the ones we don't build.
We're able to depend on datasets we don't build ourselves because the system has support for Glue catalog backends to track and register partition availability.
Currently, it builds most of the datasets using AWS Athena & S3, however this is abstracted over a single step function. There's no DAG of step functions, it's just a convenient wrapper for the Athena query execution.
The system also explicitly understands dataset builds and validations as separate steps. The dashboard makes it easy to trace the DAG and see which datasets are blocking a dataset build.
We're adding more integrations to it soon so that other ways of kicking off dataset builds and validations are available.
If people are interested in this I can begin lobbying for open sourcing the system. My colleague wanted to open source it as well.
All else fails, I'll rebuild it from scratch because I don't like the existing solutions for managing datasets. We've been calling it a data-flow orchestration system or ETL orchestration system, not sure what would be most meaningful to people.
I think the main caveat to this system is that I'm not sure how much use it'd be for streaming data pipelines, but it could manage the discretization of streaming into validated partitions wherever streamed data is sunk into. Our operating assumptions are that you want validated datasets to drive business decisions, not raw event data streamed in from Kafka. Making sure the right data is located in each daily (or hourly) partition is part of that validation.
Quick google shows that others have done things like this already...
[1] https://noise.getoto.net/2021/10/14/using-jsonpath-effective...
[2] https://aws.amazon.com/about-aws/whats-new/2022/01/aws-step-...
[3] https://docs.aws.amazon.com/step-functions/latest/dg/sfn-loc...
Makes me wonder if Amazon internally was using Step Functions, ran into issues trying to scale to larger graphs, realized multiple teams were using Airflow, and created the Managed Airflow service.
(We extended Argo, works fantastically well by the way!)
Running the Airflow Helm is pretty straightforward, even with more "complex" use cases like heterogenous pods for different task sizes.
Sounds like basic SaaS needs to be provided as a capability, while the teams spin up their instances and shard to their needs.
One of the problems with enterprise workflows is putting everything together. Workflows are already cacophonous. A cacophony of cacophonies is madness.
Also, Cristophe Blefari has an excellent data newsletter. https://www.blef.fr/
And Modern Data Stack has a newsletter, tool information, Q&A www.moderndatastack.xyz
You write the modules as normal python/deno scripts, we infer the inputs by statically analyzing your script parameters and we take care of the rest. You can also reuse modules made by the community (building the script hub atm).
Personally, I think Airflow is currently being un-bundled and will continue to be with more task specific tools.
At the very least, if un-bundling doesnt occur, Prefect and Dagster are working hard to solve lots of these issues with Airflow.
Evolution of products and engineering practices is not linear and sometimes doesnt even make sense when looking at a-posteriori (as much as I would like it to follow some logical process). Will be interesting how this space will develop in the next year or so.
At my startup Cronitor we have an Airflow sdk[0] that makes it pretty easy to provision monitoring for each DAG, but essentially we are only monitoring that a DAG started on time and the total time taken. I keep thinking about how we could improve this and it would be great to hear about what’s working well today for monitoring.
We have something like orders. So people put orders into our system, some orders are imported from external system. We have something around 100-1000 orders per day, I think. Each order goes through several states. Like CREATED, SOME_INFO_ADDED, REVIEWED, CONCLUSION_CREATED, CONCLUSION_SENT_TO_EXTERNAL_SYSTEM and so on. Some states are simple to change, like few milliseconds to call some web services, some states are 5 minutes from operator, some states are few days. This logic is encoded into our program code. We have plenty of timers, every timer usually transfers orders from one state to another. This is further complicated by the fact that this processing is done via several services, so it's not a single monolith but some kind of service architecture.
Our management wants something to have clear monitoring, so you can find a given task by some property values, monitor its lifetime, check logs for every step, find out why it's failing, etc.
What I usually see is that Apache Airflow is used more like cron replacement. I've read some articles but it's still not clear whether it could be used as a business process engine. I had some experience with Java BPMN engines in the past, it was not very pleasant, but I guess time moved on.
- A few things I've noticed is Airflow generates a tonne load of logs that will fill up your disk quite fast. I started with 100GB and I'm now at 500GB, granted disk space isn't expensive, but still even with a few DAGs i'm surprised at how quickly. Apparently you need a DAG to run to clear those logs but I was too lazy so I just purge the logs using a cron job.
- The SQL Server Operator is buggy, I filed an issue with the Airflow team but I had to do some hacky stuff to get it to work.
- Even with a few DAGs, Airflow will spike the CPU utilization of the VM to 100% for X minutes (in my case about 15 minutes) which is quite interesting. My tasks basically query SQL Server -> dump to CSV (stored on GCS) -> import to BQ.
- My DAGs execute every hour, and if Airflow is down for X hours and I resolve the issue, it will try to run all the tasks for the hours it was down which isn't ideal because it will take hours to catch up. So I've had to delete tasks and only run the most recent ones.
Granted my set up is pretty simple and YMMV, but Airflow has done what it needs to do albeit with some pain.
Have you checked why that is? Airflow does Reimport every few seconds. We've had an issue where it didn't honor the airflowignore file making it execute our tests everx few seconds. The easy solution was to put them into the docker ignore.
You might also be having too much logic in your root levels. It's recommended to not even import at root level to make importing faster.
Not saying it's not an odd tool though.
- direktiv runs containers as part of workflows from any compliant container registry, passing JSON structured data between workflow states
- JSON structured data is passed to the containers using HTTP protocol on port 8080
- direktiv uses a primitive state declaration specification to describe the flow of the orchestration in YAML, or users can build the workflow using the workflow builder UI
- direktiv uses jq JSON processor to implement sophisticated control flow logic and data manipulation through states
- Workflows can be event-based triggers (Knative Eventing & CloudEvents), cron scheduling to handle periodic tasks, or can be scripted using the APIs
- Integrated into Prometheus (metrics), Fluent Bit (logging) & OpenTelemetry (instrumentation & tracing)
If you have time please try and jump on Slack to give us feedback!
I think it slowly ends up being sort of isomorphic to the set of problems that sharing database access across service and ownership boundaries has, and my view is increasingly of the "convince me this can't be an RPC call, please" camp, and when it really can't (for throughput reasons, for example), "ok, how about this big S3 bucket as the interface, with object notification on writes?"
I'm now at a tiny company without the infra to support Airflow, so I tested all the managed airflow providers (Composer, MWAA, and Astronomer) and ultimately settled on Astronomer. I also spent some time creating a framework that generates Dags from configuration and now 90% of my new etls just involve creating a config and a sql query. Pretty happy with my current setup and my only complaint is that Astronomer won't let you exec into a Task pod, which is actually kind of a big deal, but still better than having to manage a k8s cluster.
i actually like azkaban a lot better. of course, writing a plain text job config could also be painful. i think ideally you could write job def in python or other lang but it gets translates to plain text config and does not interfere with your job code in any way.
I don't think I ever worked anywhere that had automated workflows, though my I only worked for small startups so far.
But at the end of the day, what is the alternative? Dagster seems to be a good one, but after I tried it, it seems just not worth it to move.
That sore point makes running and using the software needlessly frustrating, and honestly I won't ever be using it again because of it.
And what kind of scheduler? Again, for "a single instance" it doesn't need to be distributed. For distributed operation, Nomad is as simple and generic as you can get. If you need to define a DAG, that's never going to be simple.
In a simplistic ETL/ELT pipeline you can model things as "Extract everything, then Load everything, then Transform everything", in which case you'll add a bunch of unnecessary complexity with Airflow.
If you're looking for a framework to make the plumbing of ELT itself easier, but don't need sub-graph dependency modeling, Meltano is a good option to consider.