How to Manage Apache Airflow with Systemd on Debian or Ubuntu
janakiev.com
janakiev.com
They seem to be coasting too, for quite some time. Their website is probably the most terrible site I've ever seen for an expensive piece of software. You can't even tell what it is, how to buy it, or even how to contact them. https://www.abinitio.com/en/
Other than that, there are a several other tools in the Enterprise Analytics space that fall into a similar pattern like Alteryx or Collibra. But from their perspective it makes perfect sense, I guess. When your sales is done by relationship building and there isn't much competition once you're in, there isn't really a need to boast a fancy website or make an effort.
If anyone has a good resource on how enterprise IT procurement is done or the dynamics around it, I'd love to read up on that.
"When your sales is done by relationship building and there isn't much competition once you're in, there isn't really a need to boast a fancy website or make an effort."
Sure, but try and find a phone number or email address on their site. They've taken coasting to a whole new level.
https://www.quora.com/Which-companies-in-the-USA-use-Ab-Init...
My suspicion is that their customers are mostly companies that use Teradata, because it has a fair amount of Teradata specific features. Probably not good news for their future, but lucrative for now.
What's actually mind boggling for me, and I wonder if it's a bit of over-engineering, is people going for complex setups (oh, it's just Airflow scripts with k8s and a little bit of SystemD services and configs plus some shell scripts) when there are COTS tools that do more for less engineering cost. Yes, these carry a price tag, but it's usually quite less than paying for engineers to babysit a tool with a ton of moving parts...
Airflow can do some things many commercial tools cannot, though, so for some it is the right option.
Why not kill ASAP and restart?
Do you have problems with load booming up sometimes?
Jupyter doesn't do scheduling and integrates pretty well with Airflow.
These activities are usually managed by cron and more often by advanced scheduler tools (depending on the vendor), so it's quite a core part of any architecture that needs to e.g. load/reload/refresh data periodically.
If the requirement is simply to connect notebooks to a data lake, then the only scheduling required is to load the data lake, and something like Airflow may be overkill for this, depending on what/how the data is processed and loaded.
I wonder what your issue is/was? Notebooks are supported by means of a Papermill operator (equivalent to how Netflix operationalizes notebooks) or PythonOperator/BashOperator which would just wrap around your notebook.
However to parralelize tasks Airflow needs to know a bit more hence you might have found it required to break up your notebook into individual tasks that combine into a DAG. Is that what you meant?
Prefect was started by an Airflow maintainer and friend of mine who also contributed the dask executor to Airflow. Hi @Jeremiah!
If you have idempotent tasks, which is a best practice, it is possible to use Airflow in HA even in active/active. It might occasionally schedule a task twice, which should be caught but in any case mitigated by idempotency.
If you are looking for more Enterprise support you can reach out to Astronomer (disclaimer: I'm a advisor to them) or use a Google cloud hosted version (Cloud Composer). Both are great products.
I'm guessing I misunderstand what is meant by a "Workflow"?
My assumption is that these are managing a state machine, where workflow is stand-in for "Business Process"? If it's doing that sort of job, I'd expect timers and loops?
However, it seems these are aimed at data conversion pipelines?
Data transformations are one thing. For us, it’s the most important thing. Our data warehouse runs as a massive DAG of nightly batched transformations over app-generated data.
We also use DAG-managing tools to call external APIs and get new data (eg for weather and geocoding) and batched ML training/inference pipelines too.
Why something like Airflow? Dependencies are easier to manage reliably. If you have hundreds or thousands of nodes in your DAG, then it is a lifesaver to be able to easily 1) run many threads of independent nodes; 2) re-run on failures; and 3) find nodes impacted by failure.
Pulling data from all the various teams' locally created data stores and external systems to push to analytics is definitely a large problem.
I was trying to figure out if these are aimed at data transformation pipelines, or state management systems - I've got state management problems, not data transformation problems.
Slightly different problems, but both fit with "Workflow".
The number of pipelines and executions is a function of the complexity of your application, and invariant of the number of records being processed by the batch jobs within those workflows.
I don't understand your question. Perhaps the answer is that workflows naturally require data processing tasks to spawn collections of child tasks when a parent task finishes, and conversely they are also require to spawn a child data processing task only after a collection of parent tasks finish executing. Therefore this requirement to fork and join tasks ends up being modelled as a directed acyclic graph of processing tasks.
Here is my current setup for Airflow:
- 1 container for the webserver
- 1 container for the scheduler
- 1 managed database (I use Postgres, it's a fairly small instance.)
- 1 S3 bucket to deploy DAGs. I mount it on my containers using s3fs-fuse.
- You can monitor the scheduler using a PID file, whereas the webserver can be monitored probing your admin URL.
- Most configuration can be done using environment variables, which is perfect for containers.
- I also configure DAG logs to be shipped to a S3 bucket.
A scripting language isn't well suited for process management, which is a very well specified task.
Now, if you mentioned https://ammonite.io/ you might have got my attention...