Strymonas – A library for OCaml and Scala for fast, bulk, in-memory processing
strymonas.github.io
strymonas.github.io
Julia's generated functions are also capable of similar designs.
I am not so sure, but I have also not looked into it too deeply. But in the end it doesn't matter as long as it is just a maturity issue. After all I guess the stream library in the submission is to be seen as a motivation for making MPS more mainstream.
I'd be interested in hearing some exameple applications of this, if anyone here uses MSP on streams in serious applications. Killing runtime overhead like this seems very powerful intuitively, but use cases where it would have a significant impact don't immediately jump to mind for me.
[1]: http://www.cs.rice.edu/~taha/publications/journal/dspg04a.pd...
Linear algebra is the other obvious-to-me application of this. Linear algebra libraries typically either create a lot of unnecessary intermediate arrays or provide ugly imperative APIs (hi BLAS!) and force that busy work onto the programmer. Deforestation is an excellent way to regain performance while still maintaining ease of use. Applications include deep learning, and this is broadly the approach taken by TensorFlow and other libraries.
It sounds like I need to read more about MSP and maybe start using it.
def filterFilterTest (xs : Rep[Array[Int]]) : Rep[Int] =
Stream[Int](xs)
.filter(d => d % 2 == 0)
.filter(d => d % 3 == 0)
.fold(unit(0), ((a : Rep[Int]) => (b : Rep[Int]) => a + b))
and xs.filter(d => d % 2 == 0).filter(d => d % 3 == 0).fold(0)(_+_)
(grabbed from unit tests here: https://github.com/strymonas/staged-streams.scala/blob/maste... )They both look pretty much the same and superficially do the same thing too, so what makes this new thing better/novel?
Check the structure of this library following the definitions of Run in line 41 [1] upwards. You will notice that the code is in continuation passing style. If you combine the signatures (which are there to handle different cases of execution) they could easily be e.g., ('t -> bool) -> bool. But the implementation is more elaborate than that and note that nothing is rewritten. The design relies on the smartness of CLR/JIT to exploit that structure and inline loops. I have just described a firstly push-based, library design. You can check the differences with pull-based in the Clash of the Lambdas [2] paper.
The strymonas library is pull-based first (combined with the push-based approach) to accommodate more cases in a fundamental pull-based way. On top of that we use multi-stage programming as you saw in the paper as we don't want to rely on a sufficiently smart JIT compiler. We use staging to generate tight-loops, fused, without space leaks.
[1] https://github.com/nessos/Streams/blob/master/src/Streams/St...
I would expect it to be very much faster than Akka Streams. It has quite a different design as far as I can tell, and doesn't support things like distribution across machines.
Does that help?