Really cool API, you should port this to concord! =)
i'd say major diff is dynamic topology. So during the pipeline execution you can add/remove workers for any stage.
Also each stage/operator can be written in any programming language.
Storm/Flink/SparkStreaming/etc... all have much higher level API's. We built the execution engine first, these great things (DSL, etc) should come soon. For example this API would be easy to support to execute on top (the pipe abstraction that is)
Here is an example of a DSL we prototyped in a couple hours.
Next up I think would be supporting custom sources/sinks such as Twitter, HDFS, RDMS, etc. What exactly would be involved in "porting" riko to concord and what would the advantages be for doing so?
Advantages are:
1. mesos integration, with that comes containerization support, multi tenancy, QoS, proper pipeline supervision, etc 2. Scheduling of pipelines. i.e.: Schedule them on 100 computers. 3. the outputs of your DAG could be consumed by other systems immediately and even written in different programming languages. So your python DSL could be the source to the Scala DSL at some point. so language interop 4. Available KV storage 5. Tracing (Zipkin) - ala Google Dapper. 6. Fast networking - C++ backed runtime is 1 order of magnitude faster than the python one.
What would be involved, is not much from what I can tell.
Each concord 'operator' is like a networked function.
so given a DAG, you could generate many operators internally, or literally write them to a file, i.e.: operator_one.py etc. The code generation or internal scheduling would be the glue that's needed.
if you ever become interested, ping me! would love to collab alex@concord.io
That's pretty impressive! So something like react/elm hot reloading?