Petabyte-Scale Data Pipelines with Docker, Luigi and Elastic Spot Instances
tech.adroll.com
tech.adroll.com
Raw data is collected by our bidders and adservers that push data to S3 and Kinesis for real time consumers (http://tech.adroll.com/blog/data/2015/06/26/kinesis.html)
It is a very straightforward service, for example it doesn't do retries (done by Luigi), it is not highly available and not distributed. That makes things much easier to maintain, and we don't need HA since jobs inputs/outputs are in S3 anyway. Luigi makes it easy to restart parts of the pipeline if scheduler ever goes down.