- read from an input (source)
- perform some sort of processing
- write the data to some output (sink)
This may either be batch or continuous (stream). The inputs may change, the outputs may change.
I personally think that sql and duckdb are well positioned to do this. SQL is declarative, standardized and has decades worth of mature implementations.
The “source” can be modeled as a table.
The “sink” can also be modeled as a table.
What does a custom dsl provide over sql?
I have a side project called Sqlflow which is attempting to do something similar/
https://github.com/turbolytics/sql-flow
It’s not a DSL but the pipeline is standardized using the source, process, sink stages. Right now the process is pure sql but the source and sink are declarative. SQL has so much prior art, linters and a huge ecosystem with many practitioners.