As mentioned elsewhere, AWS step functions are really the best in orchestration.
As mentioned elsewhere, AWS step functions are really the best in orchestration.
AWS Step Functions is a proprietary service provided exclusively by AWS, which reacts to events from AWS services and calls AWS Lambdas.
Unless you're already neck-deep in AWS, and are already comfortable paying through the nose for trivial things you can run yourself for free, it's hardly appropriate to even bring up AWS Step Functions as a valid alternative. For instance, Shopify's articles explicitly mention they are running their services in Google Cloud. Would it be appropriate to tell them to just migrate their whole services to AWS just because you like AWS Step Functions?
It’s like those stackoverflow answers that tell the user to stop using PHP and rewrite it in Python or something.
But fault tolerant workflow engine is not trivial thing, it may cost you many engineer hours to build, monitor and maintain it, so outsourcing it to someone else is totally viable solution.
The complexity and risk of migrating cloud providers eclipses whatever problem you assign to "fault tolerant workflow engines".
Any mention of AWS Step Functions makes absolutely no sense at all and reads at best like a non-sequitur.
If you get a good deal from one cloud provider, you can get started quickly.
It's useful even for individuals such as students who get free credits from these providers: create a cluster and you're up and running in no time.
Our rationale was that we didn't wanted to be tied to one cloud provider.
At our company, AWS Step is a disaster. You're effectively writing code in JSON/YAML. Anything beyond very simple steps becomes 2 pages of YAML that's very hard to read or write. There is no way to debug, polling is mostly unusable. Changes need to be deployed with CF which can take forever, or worse hang.
It's the most one of the most annoying technologies I've used in my 20+ years of engineering.
Thoughtworks made a case for this distinction in https://martinfowler.com/articles/cant-buy-integration.html#...
Our teams have many 1 or 2 step DAGs that are idempotent. They could have been lambdas and they're already pulling from SQS already. It could be just my misfortune, but in AWS, MWAA is kind of janky. It's difficult to track down problems in the logs (task failures look fine there) and the Airflow UI is randomly unavailable ("document returns empty", "connection reset" kind of things).
There's certainly some functionality overlap, but I don't see Lambda and Airflow as competitors. Each has capabilities that the other doesn't.
The author also mentions that it’s used for machine learning models which will ultimately feed back into Shopify’s front end, for instance.
“Oh I have to learn how to use and setup this tool? I think I’ll just pay the equivalent salaries and be locked in…”
For a one person shop like me, AWS is a force multiplier. With it, I can do (say) 30% of what a dedicated engineer in a specific role could do. Without it, I'd be doing 0%.
I really like this tradeoff for my particular situation.
I don't know how management works through this math, maybe managing people gets exhausting and they just want to out-source it so leadership doesn't have to deal with it and then they can just focus on the core product.
And the above "I don't want to deal with it" reason isn't spoken of, the more more commonly touted benefit is cloud's "flexibility". Sure, but this is actually _really_ expensive. Every cloud migration effort I've experienced is only just worthwhile to begin to talk about because the costs are based on long-term contracts of cloud resources, not the per-hour fees. Nice flexibility.
With that said, the cloud may be a good place for prototyping where the infrastructure isn't the core value add and it's uncertain. A start-up is a prototype and so here we are. But, for an established company to migrate to the cloud and fire the staff that's maintaining the on premise resources.. I'm skeptical. More than likely, this leads to maintaining both cloud and on premise resources, not firing anyone, and thus, actually increasing costs for an uncomfortably long time.
And for the folks on the ground, who don't pay the bills, the increase of accidental complexity is rather painful.
You won’t get around any of the problems of airflow by moving to a cloud offering.
They don't always make sense, in certain scenarios it is worth taking an open source, cloud independent tool, in some scenarios you can roll your own, but there are circumstances where it's a good choice using a tool your cloud provider gives you.
I've been a one-man army at places because leveraging these cloud offerings allows me to crank out working software that scales to the moon without much thought.
I'd rather pay AWS/GCP to handle infra, so that I can get 2-3x as many project done.
Why? Where else is this mentioned?
* run job for given reporting_period * backfill data for reporting_period between N and M * lookup failed jobs
Nothing in Databricks supports that.
Backfills are on our roadmap We are previewing looking up failed jobs soon.
Email me at bilal dot aslam at Databricks dot com if you want more info