Java 8: No more loops
deadcoderising.com
deadcoderising.com
public IList<String> getDistinctTags(IEnumerable<Article> articles)
{
return articles.SelectMany(a => a.Tags).Distinct().ToList();
}
The entire LINQ "empire" (.NET 3.5) is built on top of IEnumerable<T> which was around since .NET 2.0. Streams seem to be very artificial; why not rely on Iterable<>?Oh, and no "yield" in Java.
One of the benefits of the Java version is it is easier to understand if you don't have a Java background but do have an FP background. With your C# example you'd need to find the documentation to find out what SelectMany does (which is probably just a helper method that abstracts a map and flatMap call)
Are there professional programmers with only FP and no imperative language experience?
Could the java equivalent be?
articles.stream().flatMap(article -> article.getTags().stream())
Or is the previous map required?It's not just a single method call; it's the entire approach. This entire method chain is something that cannot be done in Java as cleanly as it is in C#. Grouping is a prime example. Compare
articles.stream().collect(Collectors.groupingBy(Article::getAuthor))
with articles.GroupBy(a => a.Author)In other words, in order to support a gimmick you can actually use in production in a maybe a handful of use cases, they complicated the api for the use cases you hit 99% of the time. Awesome.
If you think that api is complicated then I don't think programming is for you, this is a very ordinary and usual construct in programming.
It's even fine in the case where you're pulling data from a file or other low-latency sequential data source, assuming that the cost of filling a spliterator buffer is less than your cost of processing.
But there's a list of gotchas all more dangerous than the "magic make it faster" button of .parallel() imply:
- For the sequential data source case, if the cost of filling the spliterator buffers is higher than the cost of processing, you're just wasting a ton of overhead trying to use parallel.
- You have to be aware that by default all uses of parallel() run on the same threadpool, which makes it a potential timebomb if someone uses it in the context of, say, a webserver where multiple requests might all individually process streams. This also means blocking operations during stream processing are very dangerous.
- Mutating an external variable goes from being fine for a sequential stream to a race condition for a parallel one.
- You can't hand out Streams that you intend to be executed sequentially, b/c your callers can just call parallel() whenever they want.
And, yes, all of these considerations make the api more complicated than one operating over plain old iterators.
articles.findAll{ it.tags.contains("Java") }
All I've done is adding groovy.jar
I have to read the underscore.js documentation every time I am using it :(.
I think it comes from FP's rooting in mathematics, and math is itself unnecessarily arcane.
But map is confusing: it's also a data structure. "Wait, does map() create a new key-value dictionary?" Filter and collect are okay. Reduce is pretty arcane but tolerable, but the actual operation feels intuitively closer to "categorize" or "group."
My point about math was cultural. Math seems to revel in having its own peculiar and often even domain specific (even within mathematics!) terms and symbols for things.
2. cadr is relatively easy, if you know your assembly (http://en.m.wikipedia.org/wiki/Car_and_cdr#Etymology) :-)
And math has locally defined terms because one cannot give every concept a unique short meaningful memorable name. That is no different in computer science, where 'integer' can include negative numbers or not, may or may not wrap around, can be any number of bits or potentially unlimited, may include minus zero, etc.
If you know IBM 704 assembly...
Maybe this will look more readable after getting used to it, but I'd much rather be looking at Clojure or Scala for now.
(def articles
[{:title "title1" :author "author1" :tags #{:Java :t2 :t3}}
{:title "title2" :author "author1" :tags #{:Jvxa :t2 :t3}}
{:title "title3" :author "author3" :tags #{:Java :t3}}])
;; find the first article in the collection that has the tag “Java”.
(first (filter #(contains? (:tags %) :Java) articles))
;; ==> {:tags #{:t2 :Java :t3}, :title "title1", :author "author1"}
;;get all the elements that match instead of just the first
(filter #(contains? (:tags %) :Java) articles)
;; ==> ({:tags #{:t2 :Java :t3}, :title "title1", :author "author1"}
{:tags #{:Java :t3}, :title "title3", :author "author3"})
;;group all the articles based on the author.
(group-by :author articles) ;; cheating?
;; ==> {"author1"
[{:tags #{:t2 :Java :t3}, :title "title1", :author "author1"}
{:tags #{:Jvxa :t2 :t3}, :title "title2", :author "author1"}],
"author3"
[{:tags #{:Java :t3}, :title "title3", :author "author3"}]}
;;find all the different tags used in the collections
(apply clojure.set/union (map :tags articles))
;; ==> #{:Jvxa :t2 :Java :t3} (group-by :author articles)
groupBy ((==) `on` author) articles
What's wrong with cheating? :P def topJavaArticle(articles:List[Article]) = articles.find(_.tags.contains("Java"))
def javaArticles(articles:List[Article]) = articles.filter(_.tags.contains("Java"))
def byAuthor(articles:List[Article) = articles.groupBy(_.author)
0: https://news.ycombinator.com/item?id=8874785On first look: "No loops?" "This reduce functions everyhere look ugly."
On second look: "Oh, I can import paralel reduce instead of the single-threaded one?" [1] "Somebody created a library to transparently switch between local reduce and one using hadoop?" [2]
[1] http://clojure.org/reducers [2] https://github.com/aphyr/tesser
Edit- there is a good flatmap example there that I glossed over when first reading. It looks like Java even has serviceable syntax for passing around lambadas like that now too, that is cool.
in scala, finding the first article looks like:
def topJavaArticle(articles:List[Article]) = articles.find(_.tags.contains("Java"))
which returns an Option[Article], so it handles the null check.
Since List in scala already handles functional constructs directly, then the filter example is just as easy:
def javaArticles(articles:List[Article]) = articles.filter(_.tags.contains("Java"))
group by author?
def byAuthor(articles:List[Article) = articles.groupBy(_.author)
And yes, that's the entire method definition, signature included. We could have added the return types for documentation if we felt like it.
You might still not like the style, but you can't say it's more verbose than the floor loop.
Backwards compatibility and design philosophy makes sure that Java 8 doesn't go hard enough on the sugar. It's why I am not optimistic of Java's future. There's awesome features out there in newer languages, like proper pattern matching, that Java just won't be able to borrow from functional languages.
"Elegance and familiarity are orthogonal."
I didn't fully understand this concept until I read your post and found myself disagreeing with you, thinking "what is obstinate talking about? OBVIOUSLY the streams way is way more readable and far less verbose".
The thing is, I can't believe I'm thinking that because I was exactly in your place a few months ago. Since that time, however, I've become very familiar with functional styles and now the Java 8 streams way seems "almost, but not quite right" and the imperative iteration style seems "gratuitously complicated and philosophically wrong... I mean... look at all that special syntax! That mutation! The horror!"
All of this is to say that I'm not quite sure either of us is more correct than the other, but familiarity does seem to cause a profound mental shift.
E.g. same Example in Dart:
class Article {
String title;
String author;
List<String> tags;
Article(this.title, this.author, this.tags);
}
Article getFirstJavaArticle() =>
articles.firstWhere((x) => x.tags.contains("Java"));
List<Article> getAllJavaArticles() =>
articles.where((x) => x.tags.contains("Java"));
List<String> getDistinctTags() =>
articles.expand((x) => x.tags).toSet().toList();
Can even be shorter without the Optional typing, but it's more readable to be explicit to have them. Dart benefits from having Collection and Stream mixins so you always get a rich API on Dart's collections.If anyone's interested to comparing FP collections in different languages, I've ported C# 101 LINQ examples in:
- Swift https://github.com/mythz/swift-linq-examples
- Clojure https://github.com/mythz/clojure-linq-examples
- Dart https://github.com/dartist/101LinqSamplesNice autologism.
I’d have to look up what .expand, .where etc means, while Java just uses the standard FP names that every CompSci student knows.
Here's the Clojure version:
;; given articles = [{:title "t1" :author "a1" :tags #{:t1 :t2}} .. etc. ]
;; These 4 Clojure one-liners replace all the J8 code examples in the article.
(first (filter #(contains? (:tags %) :Java) articles))
(filter #(contains? (:tags %) :Java) articles)
(group-by :author articles)
(apply clojure.set/union (map :tags articles))To wit: You can expect others to be familiar with basic features of the language and ecosystem, or prepared to learn them.
My gut says that loops, being a more primitive concept are likely to perform better in most situations. In addition I just find loops easier to reason about, but that is probably purely personal.
This is largely implementation specific. For instance, the .net LINQ to object implementations are largely syntactic sugar around loops (that is they compile to the same thing). Similarly, for loops are frequently just syntactic sugar around while loops.
[1] http://arxiv.org/pdf/1406.6631v2.pdfOn one hand, it does prove their point that in certain very specialized cases (looping through an array with no abstraction atop it), you will have significant performance penalties in the generic iterator case.
On the other, I'm not sure I would attribute this to LINQ. I'm reasonably certain (and in these cases the space "reasonably" represents could have a truck driven through it) if you were to write the same code as a foreach loop using the same iterator and generic collections you wouldn't see significant performance differences. I'm definitely confident in most "real world" uses, where you are already using generic collections and iterators, you should bias towards using the LINQ implementation (assuming you believe it is better code) until definitive performance testing proves otherwise. For instance, in the case of the sum of squares, that looks like classic loop unrolling optimizations not being applied which any indirection in the looping code can prevent.
Further, they show that there already exist optimization libraries that can eliminate much of the overhead.
I will say, I'm quite impressed by the java results on this benchmark.
Also, I didn't write any tests to prove any of this, so could be wildly off the mark. Further, I've spent more time than i ever wanted either hand translating or writing macros to, translate high level collections code into while loops. But that was in an extremely performance sensitive environment.
EDIT: There is more in-depth discussion of this in this thread. I missed it.
My gut feeling is that assuming streams are (in the end, but without looking) based on the loops, and Java JIT has great inlining capacity, there is really no measurable difference.
Again, one actually should measure it to conclude anything.
However, having that explicit "stream()" signifier is a very Java-y thing to do and appears to ask the programmer to decide how best to compile the given line. I would expect the compiler should be doing that work for us.
Who cares about streams? Who cares about Optional? We just want to filter a list in a clear, terse manner. (Some people do care about streams and Optional, and I wish them well, but that's orthogonal to the question at hand.)
Consider the examples given. Here they are implemented in Gosu:
getFirstJavaArticle() : Article {
return articles.firstWhere(\ article -> article.Tags.contains("Java"))
}
getAllJavaArticles() : List<Article> {
return articles.where(\ article -> article.Tags.contains("Java"))
}
groupByAuthor() : Map<String, List<Article>> {
return articles.partition( \ article -> article.Author )
}
public getDistinctTags() : Set<String> {
return articles.*Tags.toSet()
}
(I cheated a bit on the last one by just using a Set, but that's more appropriate and communicates the uniqueness of the elements in the collection to the API consumer.)Beyond the dot-star flatmap operator, there isn't anything very fancy going on: just closures being passed to methods, returning familiar classes that don't require additional transformation to pass on to the rest of the world.
It's too bad, because this is certainly good enough. As Jack Nicholson said: What if this... is as good as it gets?
public final class Article{
public final String title;
public final String author;
public final List<String> tags;
public Article(String title, String author, List<String> tags) {
this.title = title;
this.author = author;
this.tags = tags;
}
}
Getters don't seem very useful on an immutable object.Scala version of your code:
case class Article(title: String, author: String, tags: List[String])For instance, let's say that you don't want to store the author's name as a string anymore, and want to store a reference to an Author object. If you have a getAuthor() method, you can change it from a simple getter to instead call author.getName(), preserving your public-facing API.
And their justification fails. Sure, currently if you read x.foo you know that no code is being executed - but you never see x.foo because everything has getters and setters. So all it does in practice is make things (even) more verbose.
// version 1
var author: String
// version 2
var author: String {
get {"\(authorFirst) \(authorLast)"}
}It would be nice to have some syntax sugar for defining properties more tersely though, like in C#.
@Getter
private String author;
@Getter
private String title;
Or, on the class, here with a fluent (non-JavaBean) API: @Data
@Accessors(fluent = true) // experimental
public class Article {
private String author;
...
}I'm sure that library/framework people need to worry about that. Most normal developers do not. For many cases Java objects like Article are just structs. They are static maps. Sure, in the example on the site there was a getTags that added some logic, but still, pretty much a struct.
Switching to either public accessor for mutable or immutable objects actually will stream line code. You might say that public mutators are bad; encapsulation and all that. For the most part little is actually gained in the majority of getter/setter code to necessitate their weight.
Heck, as per the JavaBean spec (a spec for making components to create drag and drop UIs by the way), you can't even have fluent APIs where the setter returns "this". It has to be void.
API stability has value. If you're a library, you probably want to do this just to be safe. But really for most imperative Java code it's just a lot of fluff for little value.
Then it wouldn't be immutable, if I can change the implementation I can also create a mutable version.
Edit example
public class Article {
private final String title;
private final String author;
private final List<String> tags;
private Article(String title, String author, List<String> tags) {
this.title = title;
this.author = author;
this.tags = tags;
}
public String getTitle() {
return title;
}
public String getAuthor() {
return author;
}
public List<String> getTags() {
return tags;
}
}
public class MutableArticle extends Article {
private String title;
private String author;
private List<String> tags;
public MutableArticle() {
super(null, null, null);
}
public String getTitle() {
return title;
}
public void setTitle(String title) {
this.title = title;
}
public String getAuthor() {
return author;
}
public void setAuthor(String author) {
this.author = author;
}
public List<String> getTags() {
return tags;
}
public void setTags(List<String> tags) {
this.tags = tags;
}
} public class Article {
private final Author author;
...
public String getAuthor() {
return author.getName();
}
}
Also, your first class isn't fully immutable—getTags should be implemented as follows: /**
* @return Unmodifiable list of tags
*/
public List<String> getTags() {
return Collections.unmodifiableList(tags);
}True. It would be nice if Java had some immutable collection classes that don't have mutable methods. A method that gets a List<String> made with Collections.unmodifiableList(tags), doesn't know that it is actually immutable.
Say you had toString() in Article class that returned ("%s%" author, title).
Now if you create a MutableArticle on it, ma.toString() will return null unless you override toString on it as well (since Articles variables are private and not inherited).
I don't code much so I don't know is it normal to see code like that?
Also less relevant but Articles constructor should probably be public?
The Article class is unusable as provided due to the private constructor (in fact, the MutableArticle class wouldn't even compile), but private constructors can be useful. I often use private constructors and instead provide public static methods that call the constructors. This practice is often derided, and goes into Java's perception as a verbose, arcane language, but it allows for improving an API without breaking clients built to previous versions, as well as making an API more clear (since we can give the static methods more descriptive names).
I'm familiar with Java private constructors and builders/factory methods though, that was more of a minor nitpick, but again thanks for explaining on that.
Properly implemented getTags should create a copy of tags so that fiddling with the returned value doesn't affect the object.
https://github.com/jmoy/norvig-spell/blob/master/java/src/ma...
tagList = articles.tags.@flatten.@unique.@sort
Yes, some bits might be too architected for your MVP web app that you're going to re-write in a year (and even that depends mostly on the libraries you're using; there are plenty of lean libraries for the web-startup crowd).
data Article = Article { title :: String
, author :: String
, tags :: [String]
} deriving (Show)
articles = [ Article "Functional Java" "James Gosling" ["functional"]
, Article "Practical java" "James Gosling" ["enterprise","architechture"]
, Article "Imperative Haskell" "Simon P Jones" ["imperative", "purely imperative"] ]
firstJavaArticle = headMay . filter (isInfixOf "Java" . title)
allJavaArticles = filter (isInfixOf "Java" . title)
groupArticlesByAuthor = groupBy ((==) `on` author)
distinctArticleTags = nub . join . map tags
Full working code example with code imports and type signatures: import Data.List (isInfixOf, groupBy, nub)
import Safe (headMay)
import Data.Function (on)
import Control.Monad (join)
data Article = Article { title :: String
, author :: String
, tags :: [String]
} deriving (Show)
articles :: [Article]
articles = [ Article "Functional Java" "James Gosling" ["functional", "functional"]
, Article "Practical java" "James Gosling" ["enterprise","architechture"]
, Article "Imperative Haskell" "Simon P Jones" ["imperative", "purely imperative"] ]
firstJavaArticle :: [Article] -> Maybe Article
firstJavaArticle = headMay . filter (isInfixOf "Java" . title)
-- implemented in terms using allJavaArticles (NOTE: This IS performant in Haskell and IIUC due to stream fusion will only iterate once. Did not verify though.)
firstJavaArticle' :: [Article] -> Maybe Article
firstJavaArticle' = headMay . allJavaArticles
allJavaArticles :: [Article] -> [Article]
allJavaArticles = filter (isInfixOf "Java" . title)
groupArticlesByAuthor :: [Article] -> [[Article]]
groupArticlesByAuthor = groupBy ((==) `on` author)
distinctArticleTags :: [Article] -> [String]
distinctArticleTags = nub . join . map tags
main = undefined
0: http://en.wikipedia.org/wiki/Curryinghttp://tech.pro/tutorial/2011/functional-javascript-part-4-f....