Show HN: Pidove, an Alternative to the Java Streams API
github.com
github.com
This does seem a bit better, would be curious to see in practice if it tends towards the same hard-to-readness or if it holds up.
Groovy on the other hand has very nice ways of doing the same thing that become second nature before too long.
Also the lambda syntax in Java is as concise as can be, how can you beat
x -> x*x
? Sometimes passing a lambda (or other function) as an argument is a simpler approach to specialization than defining a subclass. That I think is mainstream and accepted in Java today. More recently Java is also incorporating other ideas from the ML family such as switch expressions and algebraic types (records and sealed interfaces.)There is a lot more to "functional programming" than that, such as the use of persistent collections. In some cases (such as managing the symbol table in a compiler) those methods lead to good efficiency and great simplification, in other cases they are ways to make easy problems punishingly hard.
pidove builds on top of ordinary Java Collections and doesn't push more exotic approaches as does
One thing that always struck me as awkward about the Streams API was the need to create Streams and Collectors for every operation, say you want to do
var x = someList.stream().map(x->x*x).collect(Collectors.toList());
pidove simplifies that to var x = asList(map(x->x*x,someList));
Part of the simplification is that pidove operators take any Iterable as an input so there is no need to convert to a stream. This has some add-on effects in that operators like flatMap aren't creating so many heavyweight objects.Along the way I discovered that I really liked the Collectors in the Stream API in that
* Collectors have some quality engineering to improve numeric accuracy for sums
* Collectors can be used inside grouping and windowing operators
* another approach I tried took required fancy footwork w/ generics
so I decided to use them in pidove rather than write my own. I decided the main thing needed was some syntactic sugar and the Java maintainers agreed because they added this in JDK 16https://docs.oracle.com/en/java/javase/18/docs/api/java.base...()
Or, at the very least, being able to define functions taking a stream<A> and producing a stream<B> and compose it with another function taking a stream<B> as input and apply it all to a stream<A>.
Both are just not possible with the current Java stream API.
This is a common thing in many functional style programming languages.
For it to really shine I need to add more versions of the compose function with different lengths, but I am working on a code generator that will be able automatically generate those to arbitrary length.
static <A, B, X> Function<A, X> compose(Function<A, B> one, Function<B, X> two) {
return (a) -> {
return two.apply(one.apply(a));
}
}
Java sucks when there are a variable number of generic arguments, so it struggles with the composition of two functions that have a variable number of inputs. That doesn't really apply to stream processing though.Would be nice to be able to write, for example:
T result = streamOfA().apply(f).apply(g);
It's a bit unfair to compare stream API to similar APIs because it's so universally applicable. For example, Python (and I suspect groovy) impose some implementations and limitations that aren't imposed with streams.
Microsoft's LINQ for instance can compile a stream operation into a SQL statement and JooQ does the same. That system offers query optimization and efficient joins that depend on the query system having complete visibility into the queries. indexes built ahead of time, etc.
Another extreme is a system like
https://www.reactive-streams.org/
that are especially good for apply a filter and map and other operations to a stream of real time events, e.g. instead of having a pull operation such as a for-loop over an Iterable, items go into the system from a stream.
I've worked on systems that use the later kind of streaming to run batch jobs and you can get great performance (780% speedup with 8 cpus) on crazy heterogenous workloads. You do have to be careful though to shut the system down or flush it out or otherwise you get wrong answers. Frequently those frameworks don't shut themselves down properly unless you implement clean shutdown yourself.
The point is that operators like "filter" and "map" and the rest are so powerful because they are portable between the minimal pidove up to a Hadoop cluster.
Every implementation though has details, makes different decisions, has different limitations. For instance the way group by is done in Python and Java are completely different. A "Tuple" doesn't make as much sense in Java as it does in Python because it's no fun to "get(7)" and get just an Object.