Effective Airflow Development
curology.com
curology.com
That said, I make a point of using ETL-as-a-service whenever it's available, because there's no use solving a problem someone else has solved already.
Prefect - https://www.prefect.io/
Flyte - https://flyte.org/
" Licensee is not granted the right to, and Licensee shall not, exercise the License for an Excluded Purpose. For purposes of this Agreement, "Excluded Purpose" includes, but is not limited to, using the Software, or any derivative works thereof, to make available any software-as-a-service, platform-as-a-service, infrastructure-as-a-service or other similar service that competes with Prefect products or services."
While I'm interested in using Prefect as part of SaaS I'm working on, I'm having trouble defining whether it would compete with their offering or not. In my SaaS Prefect UI will not be exposed, I need an ETL engine "behind the scenes" for parts of the whole workflow (some action, sends an event, and on that event a job is triggered). In theory Prefect SaaS could be used to do the same, so I guess that would mean I'm competing with them?
On the other hand Flyte looks very young, that could mean it's not mature or hard to use for non-Lyft use cases.
Airflow 2.0 will have some pretty nice features for ML development as well.
Also they all seem to charge so very much for what amounts to a fairly straightforward service...
I've previously used Argo Worfklows, which I prefer because I already have a Kubernetes environment and because it runs containers, it's totally language-agnostic which I think is a huge benefit. It also has a huge number of features for defining and controlling the workflows. Downside is that it's configuration/definition YAML's can get large and a bit messy (as YAMLs are want to do) - however, templated workflows are coming soon, which should hopefully reduce the noise.
Personally I'm reticent to use anything that re-implements its own full scheduling system instead of hooking into a pre-existing (and probably more bulletproof) one (i.e. K8s scheduler), and anything that _requires_ me to write all of my ETL/schedulable code in Python.
We tried to make Polyaxon[0] work with Airflow for Machine Learning specific workflows, but it was very painful and it does not have a good state/artifacts management, which leaves the users tweaking around. We end up making a simple abstraction on top K8S, much easier, to provide features for parallel executions, dependencies, failure handling, retries, ... as well as handling ML specific graphs such as hyperparameter tuning and distributed scheduling.
By the way, Polyaxon looks awesome, I’ve been wanting to try it for a while, but just don’t have any machine learning projects in the pipeline at the moment alas.
The operators and scalability are somewhat useful. I was happy with the UI compared to cron. Testing is a mess. Also, Airflow isn't CI/CD-friendly (but it's possible to get it to work).
I'd recommend a managed option unless you have a skilled ops team. It reminds me of Hadoop in terms of how exciting it is to get set up, which isn't a good thing.
Now that I think about it though, most of the time I spent on testing wasn't caused by Airflow. Testing data pipelines just isn't easy with the current well-known tooling.
It does not help that the entirety of the documentation is written from the point of view of people who are definitely not of the devops variety doing things manually on their laptop. I.e. all the wrong things you should never do in a production setup. Configuring this thing for production usage is largely undocumented, non trivial, and you'll be piecing things together from stackoverflow and various third party github repositories for e.g. using docker, terraform, etc. rather than the official documentation which merely hints at these things being possibilities.
It also does not help that the internals are kind of buggy and wonky. We had a really hard time getting the basic plumbing for running workers, queues, etc. working properly. It would constantly grind to a halt and stop processing stuff. Also there's this minutes long uncertainty principle "is it actually running or still figuring out that it needs to catch up?!".
Also, the UI/UX is terrible IMHO. Think hitting cmd+r a lot because page refreshes are not a thing in Airflow and absolutely everything requires dealing with multiple clicks to navigate complex dialogs (modal, naturally). So, unless you just manually reloaded the page: you are looking at stale information. Jobs that have long finished. Green statuses that have turned red, etc. Even Jenkins/Hudson had auto reload 15 years ago. And given the significant overlap in functionality, you might actually be better off using that if all you need is the ability to run some simple job at specific intervals.
The only valid reason for using Airflow is the ecosystem of plugins. It's valid and it's basically the same reason that people tolerated the craptastic experience that was managing Nagios back in the day. Horribly complicated to setup, terrible/primitive UI, loads of performance issues, non trivial failure modes, etc. but world + dog used it and there were nagios plugins for just about everything. I've been that rabbit hole as well and I'd say the experience is similar enough.
So, definitely use it in hosted form if you can or avoid altogether unless you really need it.
We mainly use the K8sOperator, and the logic is mainly inside independent containers. Therefore our development and testing is not so tightly coupled to Python or Airflow.
Heh. ETL also stands for effective translational lift in helicopter aerodynamics.
As a side note how do you effectively google for a piece of software or product with a name as generic as "airflow"?
For unit/integration tests we ended up doing lot of Docker in Docker setup.
and then every task is just a standalone, vertically scalable service on k8s or a giant horizontally scalable compute job