Java LINQ – C# LINQ examples running on Android
github.com
github.com
https://gist.github.com/mythz/7e263820332a0fcea766
What's interesting was that Type Inference was actually considered and rejected because it was considered "un-Java Like" http://programmers.stackexchange.com/a/184183/14025:
> Humans benefit from the redundancy of the type declaration in two ways. First, the redundant type serves as valuable documentation - readers do not have to search for the declaration of getMap() to find out what type it returns. Second, the redundancy allows the programmer to declare the intended type, and thereby benefit from a cross check performed by the compiler.
With "Humans Benefit" cited as the rationale, which seems to contradict the behavior in IntelliJ IDE's which automatically collapses boiler-plate code into a "typeless lambda format" of what the source code would've looked like if Java had lambdas.
And Java tends to do type inference on the right hand side instead of the left hand side, so things look a lot cleaner if you chain them instead of assigning them to intermediate variables.
This example is chained, i.e. there are no intermediate variables.
I would expect in reality to do something like this in Java 8 (Disclaimer probably not exactly correct syntax, not near a Java compiler)
getCustomerList()
.map(c -> tuple(
c.companyName,
c.orders.stream()
.collect(groupingBy(order -> order.orderDate.year)).entrySet().stream()
.map(yearGroup -> tuple(
yearGroup.key,
yearGroup.items.stream().collect(
groupingBy(order -> order.orderDate.getMonth() + 1)
)
))
.forEach(Log::print)
Further it looks like we should probably have domain types for MonthGroup/OrderGroup etc instead of tuples, which would make it still cleanerThe `customers` and `customerOrderGroups` are variables in each of the examples. Which is the point of a close 1:1 translation between different languages. linq43 groups Customer Orders in a single expression by Year/Month into an anonymous structure. This just doesn't translate very well to Java 7 which lacks many of the language features to make it more readable.
Of course it can be rewritten using Java 8 streams which is much better suited for functional programming courtesy of lambda's (which is the primary cause for much of the boiler-plate). But until it's supported in Android we have to make do with what we have.
----
Should also give a shout-out to RetroLambda which I just discovered in the comments: https://github.com/orfjackal/retrolambda
Which looks like it's a post-build Maven and Gradle plugin to back port Java 8 features to be able to run on Java 7 Android and JVM runtime. Not sure if there are any disadvantages with this approach (debugging?) but it's an interesting alternative.
Another option would be to just use a different JVM language, Kotlin looks the most appealing to me which supports Android and has good Java interoperability: http://kotlinlang.org/docs/tutorials/kotlin-android.html
One could also go for implementing a more Linq-esq style in Java. Something like http://benjiweber.co.uk/blog/2013/12/28/typesafe-database-in... could be implemented on top of collections as well as databases.
void linq43() {
CustomerOrderGroups var = () -> getCustomerList()
.stream()
.map(c -> tuple(
c.companyName(),
group(c.orders()).by(o -> o.date().year())
.thenBy(o -> o.date().month()))
).collect(toList());
var.grouped().stream().forEach(
System.out::println
);
}
http://pastie.org/10324237I think.
If that's "bad-formed Java" what's a better functional form? Or do you mean for Java it's better to use an imperative-style with mutating collections?
var result = someEnumerable.Select(x => x.Bar < 5).ToList()
Still a lot cleaner than the type specific mess in the examples... I think sticking to for-loops is probably cleaner than that mess.I don't mean to be critical here, it's just I don't think it's really any better than without the "LINQ" for Java in this case.
I think it's an inner join, not a cross join. SelectMany() produces a cross join.
That is also totally not the point.
Running on Android was a constraint for this library, the natural approach to functional programming in Java 8 would've been to use Java 8's new stream API, but as that's not available on Android we fallback to using what currently does work on Android - which is the point of this library.
And I know, I know, bash is even less elegant and more boring. Haha.
C is also boring, but I assumed mocking it here would be asking for free kicks to the face. Apparently, I was wrong.
Ask any developer and they will have their favorite language. Java isn't bad (not my favorite). I like it better than C#, but then I hate Windows, so of course I would say that.
(Edit: Yes, I know C# is on more platforms than just Windows. If we're getting completely technical.)
FirstOrDefault being passed 2 rows = non-determinism. This one is always fun when your SQL server doesn't guarantee the order of rows returned so every time you hit .FirstOrDefault you get a different item. Never happens in dev due to the tiny datasets fitting in a single page.
There's also Single being passed 2 rows = crash. SingleOrDefault being passed >1 rows = crash.
And when it does crash it's entirely not helpful because the lambda block reports no state information in the stack. Inevitably this leads to adding precondition checks to avoid "Hey everyone, one of the 20 LINQ expressions in this method blew with more than one element in the sequence".
Then there's the fact that you don't know the size of the set of data nor really think about it. So if someone passes in a collection of 10 rows in dev and it works out that it's O(N^2) and has a random .ToList() in it then you're up shit creek in the memory and time departments when that 10000 items collection appears in production.
None of this unique to LINQ but it hides a lot of problems behind a wall of pain.
From extensive experience; there be dragons.
Which is the point of the different API's, you get to pick the behavior you want. Don't want it to crash? pick `FirstOrDefault`. A formal API is better than manually codifying the elected behavior each time, making it tedious to infer the intent each time whilst being susceptible to human errors manually copy+pasting imperative code.
It almost seems like your looking to blame Linq for what is clearly poor database and project design. If you have poor design you're going to have problems no matter what.
Single() is when there is absolute only 1 result. > 1 result is not expected ( i'll admit that i never use Single or SingleOrDefault() )
And using paging is the same thing in linq as it is in SQL. If you don't know how Linq works, you should learn it. What you should know is that .ToList() hits the database, so when you query all the objects, you get all the objects. Limit your query before the database query :)
Here are the results of page 4 with 10 items : Objects.Skip(10 x 3).Take(10).ToList(); //fetches 10 items = paging in the query
Here is how someone could write it without knowing Linq: Objects.ToList().Skip(10 x 3).Take(10);//fetches all the items and does the paging on the list.
As long as your DataType is an IQueryable<T>, it doesn't hits the database. When it's a List<T> (eg. when calling .ToList()) it's too late
The coalesce operator in C# 5 is welcome but then introduces another problem:
var x = myObject?.Property1?.Property2?.Property3;
If x is null, why and where?This is all in-memory manipulation of data. Mainly rules engine stuff for us.
A line like that is just lazy. I'm more keen to blame the developer and not the language. Striking nails with your fists as they say.
- Single and SingleOrDefault crashing with >1 row is a feature, not a bug. It means that your expectation that your filter is unique was incorrect. In that scenario, a crash is what you want. It's like when we use asserts. Similarly ..
- First and FirstOrDefault are only really appropriate after an OrderBy, i.e. to get the top item of a particular ranking. Using them anywhere else is almost always incorrect, and in the very few cases where it isn't, needs a comment to explain why, or I'll bounce your code back to you in review :).
LINQ is very expressive (you can do a lot with a little bit of code), especially with the embedded (non-fluent) syntax. That means that, yes, you can write some non-optimal code without realizing it, because it's only a few lines. On the flip side, you can also refactor that code quickly into more optimal patterns. By comparison, old-school iterating through collections is often much more difficult to refactor.
Perhaps that is the problem? :-)
More seriously, the thing has really leaky abstraction boundaries which is a main pain point. I have to write and manage a lot of boring enterprise code and creativity and expressiveness go a long way to introducing a lot of additional work.
This is actually one of the main reasons that the BCL team introduced the IReadOnlyXXX<> interfaces - over the years, many .NET developers trying to do the right thing and express intent in their code had taken to using IEnumerable<> in method signatures to signify that the method would not modify the list/collection... but an IEnumerable doesn't have to be a list/collection, it can be infinite! In trying to do the right thing and be expressive with intent, lots and lots of people were unwittingly losing the expression of another (very important) assumption.
Any time you are wrangling any kind of enumerable in code, you have to think about its guarantees and assumptions. LINQ's expressiveness might make it easier to forget this but it's still the rule - if I hand you a "malicious IEnumerable", it doesn't matter if you use LINQ or a couple of foreach loops, it's going to launch the missiles on every call to MoveNext() regardless.
If you are trying to express intent in your code, IEnumerable is almost never what you want to use.
There is your problem and it isn't LINQ. You'd have the same problem with C# without it. LINQ is a good declarative tool when declarative programming makes sense : data pipelines. And LINQ is 'lazy'...
You should be glad that an imperative language succeeded in getting declarative features. And that makes C# better than the competition in the same space, even though I never use MS techs.
And Single is supposed to crash if there's more than one row. Again, it's directly like selecting "TOP 2" then doing a check to make sure there isn't a second row. Your complaint seems like complaining that using a "while" conditional instead of an "if" can lead to trouble.
Seriously I'd love to know what code you are writing where you weren't sure if you wanted First or Single, decided Single (thus enforcing that data integrity rule), then decided it's the library's fault for doing exactly what you asked.
It's not a crash, it's an exception. You can catch an exception. But, as others have said, you could instead call the function that doesn't throw an exception if having more than one result isn't an error.