AWS Batch – Fully Managed Batch Processing
aws.amazon.com
aws.amazon.com
http://tech.adroll.com/blog/data/2015/09/22/data-pipelines-d...
It is very convenient to be able to define any jobs as Docker containers, using the best language for the task, and define dependencies between their inputs and output explicitly, not being tied to any particular paradigm like MapReduce.
I am more than happy to let AWS handle plumbing for this instead of having to maintain and operate our custom solution.
Also, have you looked at Airflow?
Not comparing them, just curious.
We took a look at Airflow, which looks nice, but we haven't had a compelling reason not to use Luigi.
* time-based and dependency-based scheduling do not mix well in our experience, if you try to do both at the same time you end up with a very complicated execution model. Luigi is much simpler, it is basically a makefile. To handle time one can just add dummy dependencies where "exists()" function returns true if now()>X, and/or just run the whole thing by cron.
* We heavily use dynamic dependencies in Luigi, where one job decides how many childs to spawn and what parameters to use at run time. Makes it easy to do e.g. mapreduce style jobs purely in Luigi.
If you're interested in building this type of infrastructure for yourself, I'd recommend taking a look at http://github.com/pachyderm/pachyderm. I'm one of the founders so this is a somewhat shameless plug, but Pachyderm is designed to get people most of the way to what AdRoll has built without having to build as much of the "plumbing" from scratch. AdRoll was some of our earliest inspiration for the design of Pachyderm.
Optimal is a pretty strong claim, given that this is an NP-Hard problem. Or do we just toss the word "optimal" around now?
Maybe that is the secret sauce.
I haven't gotten far enough yet to see if it has any distributed compute support, or if each job runs on a single machine. EMR is designed for master/worker architectures with many workers. It looks like AWS Batch might be designed to run each job on a single machine.
Lambda is for shared-compute. You don't need a dedicated server to run a "function" that takes < 60 seconds and can be called in a stateless manner as an API endpoint.
This is dedicated host compute-heavy batch processing. It's a pain to do this at scale!
I've built systems for running large scale life science embarrassingly parallelizable problems on EC2 and wish I had something like this!
Imagine you have a set of input S3 files, each needs multiple-hours of compute to produce output S3 files. Doesn't seem that hard, until EC2 instances fail, programs crash, etc. etc.
Quite a lot of life science work is stand-alone programs that are domain-specific and read and write flat-files.