Databot: High-performance Python data-driven programming framework
github.com
github.com
So here's an implementation of the example from the readme using make, and bash, and jq, and a silly python script to implement a timer by modifying a file every X seconds:
~/projects/makething$ cat Makefile
default: d
b: a
curl -o $@ "http://api.coindesk.com/v1/bpi/currentprice.json"
c: b
jq '.bpi.USD.rate_float' $< > $@
d: c
cat $<
cp $< $@
~/projects/makething$ cat timer.py
import sys
import time
def main():
delay = float(sys.argv[1])
fn = sys.argv[2]
while True:
with open(fn, 'w') as f:
f.write(str(time.time()))
time.sleep(delay)
if __name__ == '__main__':
main()
~/projects/makething$ cat go.sh
#! /usr/bin/env bash
python timer.py 2.0 a &
TIMER_PID=$!
function cleanup() {
kill $TIMER_PID
}
trap cleanup EXIT
while true
do
while make -j -q
do
sleep 0.1
done
make -j
done
make decides if things are up-to-date by comparing timestamps of files in the filesystem, so we can emulate a timer that triggers an event every 2 seconds by having a process modify a file every 2 seconds, and rig a rule in our makefile to use that file as an input.In your example, I would drop the shell and python scripts and simply run:
watch -n 2 -d "touch a && make"<opinion> for decades the *nix culture suffered from "suffer or you arent smart" syndrome, in my opinion. There is nothing wrong with easy, or nice to look at, its just that it is rare to find quality in both the implementation layer (like this) and the interface choices, and then to keep it simple-stupid going forward. I like the simplicity of this, btw</opinion>
However, one downside is the massive RAM consumption. 1GB of RAM even if it does pretty much nothing is quite a bill to start off with.
NiFi might not run on an AWS t2.micro instance. Whereas Apache Airflow does.
EDIT: plus, of course, things like your tech landscape. NiFi is Java and Airflow Python. If all of your tools are around one or the other, you can get both to work for the same use cases, but as mentioned above you will have more or less implementation work to do.
As someone else said Airflow and NiFi are decently different. However, Apache having multiple projects that overlap is pretty common.
Streaming you have Flink, Spark and Storm. Batch processing you have Flink, Spark, Hadoop MapReduce, Hive, and Tez. Serialization you have Avro and Thrift. Columnar storage you have ORC and Parquet. Batch ETL management you have Oozie and Airflow. NoSQL you have Cassandra and HBase. That's just off the top of my head.
I think this is the most popular one:
https://github.com/ReactiveX/RxPY
This one is a rewrite of RxPY, that makes use of async / await / asyncio:
https://github.com/dbrattli/aioreactive
Pretty interesting stuff!
I'm still looking for a perfect natural python ETL dsl, so I will follow that project.
So far I'm using https://github.com/petl-developers/petl and mostly happy with it.
How would that look like? I mean a pipe is something where one writes data into and another loads data from in the same order it was written into. I don't see how that could be nested, or why two pipes would need to be connected.