Differential Dataflow but at what cost? (2017)
github.com
github.com
input = ...;
input += d1;
input += d2;
input += d3;
and you have some output that depends on the input. Whenever the input changes, the output needs to be recomputed, like this input = ...; output = f(input);
input += d1; output = f(input);
input += d2; output = f(input);
input += d3; output = f(input);
The idea here is to, instead of recomputing the output every time from scratch, find a function g that will compute the output-difference from the input-difference input = ...; output = f(input);
input += d1; output += g(d1);
input += d2; output += g(d2);
input += d3; output += g(d3);
And I think if you write f using the combinators from the differential dataflow framework it will be able to find g automatically for you.If you get into it, the library itself is well documented on docs.rs and there’s an active Gitter channel.
(Disclaimer: I work with Frank at his company, Materialize.)
[0]: https://timelydataflow.github.io/differential-dataflow/
[1]: http://sigops.org/s/conferences/sosp/2013/papers/p439-murray...
For example, consider a map combinator. The output diff is just the input diff with a function applied to the data. More formally, for every incoming (data, time, diff) triple, an output triple (f(data), time, diff) is produced.
For a filter combinator, the output is (data, time, diff) if f(data) is true, or (data, time, 0) if f(data) is false.
A computation that wants to filter and map some data then just requires piping together these operators.
Things get interesting when you want to aggregate and join data, but Frank’s blog posts and documentation explain how you build that in a dataflow system far better than I can here.
It appeared at the time analogous to multivariate calculus, though the better analogy appears to be to Moebius inversion (roughly: integration and differentiation on partial orders). Read more here: https://en.wikipedia.org/wiki/Möbius_inversion_formula
"Moebius dataflow" would have been a pretty bad-ass name, I think we can all agree. And it might have done a better job of preparing the reader for what was about to happen to their brain.