(->> (string/split ... #"\s+")
frequencies
(sort-by val)
reverse
(take 10)) (->> (string/split ... #"\s+")
frequencies
(sort-by val)
reverse
(take 10))Much as I love Clojure, it's worth pointing out that Clojure's lazy sequences (include intermediate sequences) get cached, while LINQ evaluates more like Clojure's reducers. (You can even get it to do so in parallel.)
On a side note, I'd like to point out a key difference between this Clojure example and corresponding variants in C#, Python, &c: the Clojure variant has no variables. This isn't just a matter of concision: coming up with descriptive names is hard, and usually means duplicated information, either in the name or the type declaration. Often, those names are re-used over and over again with subtly different meanings, for each stage in a pipeline--forcing the reader to reason carefully about the declaring scope at each use.
Consider, for example:
Swimmer swimmer = new Swimmer("foo");
swimmer.setStyle("butterfly");
swimmer.swim();
return swimmer;
(doto (Swimmer. foo)
(.setStyle "butterfly")
.swim)
Same operation--but with (doto), four uses of a variable and one type declaration have been cleared away. Consider the original post: var words = s.Split(' ');
var wordCounts = words.GroupBy(x => x).Select(x => new { Name = x.Key, Count = x.Count() }).OrderByDescending(x => x.Count);
var countedWords = wordCounts.Select(x => x.Name).Take(10).ToList();
return ExtractTopTen(countedWords);
Six uses of three formal variables, including the bewilderingly confusable countedWords and wordCounts, plus nine uses of the delightfully generic "x". Even in the more compact C# example from this thread, consider: var top = (from w in text.Split(' ')
group w by w into g
orderby g.Count() descending
select g.Key).Take(10);
This variant refers to the temporary variables w (for "words") and g (for "groups"?) six times. Both authors felt the need to reduce the repetition of variables, but the best they could do was to choose single-character names.This is the real power of ->, .., ->>, doto, and friends: eliminating names for things. By thinking about the composition of transformations, instead of the intermediate results, you can make an algorithm easier to understand and change.
return s.Split(' ')
.GroupBy(word => word)
.OrderByDescending(wordGroup => wordGroup.Count())
.Select(wordGroup => wordGroup.Key)
.Take(10).ToList();I totally agree with you on this.
>Even in the more compact C# example from this thread, consider: var top = (from w in text.Split(' ') group w by w into g orderby g.Count() descending select g.Key).Take(10);
> This variant refers to the temporary variables w (for "words") and g (for "groups"?) six times. Both authors felt the need to reduce the repetition of variables, but the best they could do was to choose single-character names.
> This is the real power of ->, .., ->>, doto, and friends: eliminating names for things. By thinking about the composition of transformations, instead of the intermediate results, you can make an algorithm easier to understand and change.
You lost me somewhere along the way....are you saying the "from w in text"... snippet is bad, and something more along the lines of "This is the real power of ->, .., ->>" is more appropriate?
I ask because that code seems extremely readable to me. Personally, I don't give a shit if it's 30% more verbose or runs 50% slower, for 99% of code (written in the world), optimum performance doesn't matter. Maintainability does matter though. All of this software being written today has to be either maintained by someone, or replaced by something else. And the top ~2% of programmers like you sure as hell aren't going to be taking maintenance jobs any time soon.
I'm curious what the thoughts of a technically smart person such as yourself are on the subject of what companies will be left with 5 to 10 years down the road when consultants have come through and implemented using the currently most optimum platform/language/algorithms?
And I honestly don't mean for this question to be disrespectful. I'm just coming from a situation where I'm a former developer but on a project where I'm not coding, and I ask for features and the developers say they can't do it, or it will be a performance problem. And I know these guys are far more like you than me intelligence/education wise, but the things I ask for I've done tons of times in the past with 10 to 1000 times the data size, without a problem, on far older hardware.
I'm just curious what kind of a support problem you ultra smart people are leaving behind, or if you ever think about the idea that almost no on else is as smart as you?
1. As simple as humanly possible, so that the algorithm is easy to understand and change. Each component is isolated and can be understood and modified in isolation. Boundaries between component and environment are clearly thought-out.
2. As simply expressed as possible: broken up into distinct, well-organized functions, with descriptive, regular names for functions and variables, in context.
3. Well-documented; each function and each namespace come with contextual docs explaining their motivation, arguments, consequences, invariants, etc., with examples.
4. As short as possible, because humans have trouble holding large amounts of context in their head. Minimize the amount of scrolling or jumping between files necessary to understand the algorithm.
5. Well-tested, so that changes can be made freely. The test suite needs to be fast, so one can get feedback within seconds of making a change to the file; ideally a few milliseconds. A balance of typechecking, logical tests with mocks, integration tests, and full stress tests provides a continuum of safety.
I don't see these goals as particularly constrained to any language, but I will say that I feel best able to achieve them in a Lisp. Dunno whether that helps you project your problems at work onto me personally, though. ;-)
return new Swimmer("foo")
{ Style = "butterfly"}
.swim();
It is unidiomatic to have Swimmer.swim return 'this', though. var result = new Swimmer("foo")
{ Style = "butterfly"};
result.swim();
return result;I speak as someone who codes in C# for work and Clojure as a hobby. There's lots of things that just aren't quite possible in C#, and _everything_ is possible in Clojure.
Pointfree is beautiful to read and grok. It's a little trickier to debug, since pretty much every debugger on earth needs points to inspect values.
Edit: Actually, we can also get rid of the 'reverse' by replacing (sort-by val) with (sort-by (comp - val)). Not going to be shorter in terms of character count, though we win on line count.
(->> (string/split s #"\s+")
frequencies
(sort-by val >)
(take 10))
Just mentioning it here because I find it vastly more readable than (comp - val).