- groupBy is itertools.groupBy(lst, fn)
- sortBy is just lst.sort(key=fn)
- countBy is collections.Counter(map(fn, lst))
- Sibling comment mentioned flatten, which is just [item for sublist for sublist in lst]
More esoteric needs are usually met by itertools.
You mean sorted(). list.sort only works on list and in-place.
You can make lambdas in Python, but naming function is in my mind very much a "feature".
Same applies to then and else branches.
In a code review I don't what to read a thesis inside of a conditional branch.
Assuming you have five levels nested if, ignoring the fact that you might be doing something wrong, then you should name those, but that requires you to extract those code paths and put them into separate functions.
A function can do literally anything. Seeing a function call, it hasn't narrowed down the realm of what might be happening even a little bit.
"lambda f:" gives away very little information about what is about to happen. "if f:" gives away some information about what is about to happen.
Not a perfect rule of thumb, convoluted if statements do exist. But that is why functions are more in need of hints and comments than other control structures.
I think that's an unfair comparison:
If 'the thing that happens' for "if f:" is 'branch on the truthyness of f', then I agree that's pretty clear. However, if we're going to ignore the content of the branches then we should also ignore the body of the function, so "lambda f:" is also pretty clear: 'parameterise with f'.
If we're worried that "lambda f:" could do pretty much anything in the body, then we should also worry that "if f:" can do pretty much anything in the branches.
That is the part common to both. If you take that common part out, you will be left with more information in the if statement than the function call.
In the function, you know what arguments it (might) be interested in. You know they are evaluated. But in the if statement, you know what values it is interested in, usually get the same amoutn of information about whether they are evaluated, much clearer guarentees about how they are used and that the aspect of them that matters is their truthiness. Also the code might be about to branch.
Put it this way - if(...) is a specific - and common - named function. It is a very easy argument to make that it is carrying more information about what is going on than a general anonymous function.
I don’t think this argument generalises well. Once you start working with structured data and more complicated control logic, a generic loop with an arbitrary body often provides less immediate information to a developer reading the code than a function that explicitly names a specific pattern of computation like map or filter.
There are many recurring patterns of computation that are much more specific than a generic map or filter, too. While imperative languages tend to have only a small number of built-in control structures, if we’re using higher-order functions then we can name as many of these patterns as we like.
For example, suppose we want to generate a list of the first few triangle numbers, which is a sequence of the form:
[0, 0+1, 0+1+2, 0+1+2+3, ...]
Writing this in imperative Python code with a typical loop would give us something like this: triangle_numbers = []
total = 0
for n in range(6):
total = total + n
triangle_numbers.append(total)
Here’s an idiomatic functional version, which is real Haskell code using a standard library function that captures this pattern of sequence generation in just five characters: triangle_numbers = scanl (+) 0 [1..5]
Assuming you’re equally familiar with both programming styles, I think it’s fair to say that “scanl” tells you much more about what this code is doing than “for”.The inner function here is just addition in both cases. However, the code within the Python loop is allowed to do whatever it wants, so the reader has to recognise the pattern surrounding the addition. In the Haskell version, the function you pass to scanl must take an accumulated value so far and a next value as parameters and it must return the next accumulated value to add to the sequence being generated, and that is all it is allowed to do. So you just need to write (+), which is Haskell’s notation for what Python calls operator.add or an equivalent lambda, and that tells the reader exactly how the next accumulated value is being derived at each step to build up the sequence.
triangle_numbers = scanl(f, 0, [1,2,3,4,5])
gives me scant little information about what the value of triangle numbers is at the end of the operation. triangle_numbers = if(f, 0, [1,2,3,4,5])
or triangle_numbers = if(true, 0, f)
tells me that whatever f is something weird is going on. Commenting if statements to point out that they are branching code is foolish.My intended point was that comparing lambda: to if: is apples-to-oranges.
Both styles in my example have an outer structure that builds a sequence, wrapped around an inner calculation that says how you build the next value in the sequence. The inner calculation here was simple addition, a one-liner in the for loop and (+) in the Haskell version. However, it could equally have been something more complicated.
If it were, a reader would still have to figure out what the body of the for loop was really doing. In contrast, the (possibly anonymous) function passed into scanl would still have a more constrained purpose, which is immediately known because scanl represents one specific pattern of computation. The implicit algorithm it represents has a hole that is a certain shape, and you can only pass it a function with the right shape to fit into that hole.
Put another way, if you’re programming in this functional style, it’s not writing the word lambda that tells you what the next piece of code is for, it’s passing that lambda expression into a higher-order function like scanl. Passing an anonymous function is like writing the nested block for a branch of an if statement or the body of a for loop. It’s the call to scanl that is roughly analogous to writing the if or for statement itself.
Of course this assumes the lambda is pure, but in the kind of code we're talking about (pseudofunctional code with lots of higher-order functions) they should be pure for the most part.
An if statement is a lower level mechanism for directing the code path within a function. If it needs a name, wrap it in a function. I often do.
It also has a condition and a body to execute conditionally.
if (condition)
body
is just a goofy way of writing if(condition, body)Many Swing coders can also apply.
It is so hard to actually pack the functionality into a separate piece of code.
"Yeah, but then I have to look elsewhere"
Really? What about when a method is full of stuff like that?
Its readability is lost beyond hope.
If the idea of multi-line alone compromises readability, then why even have multi-line blocks at all? Function, loop and conditional blocks should all at most have one line each. I hope you can see the absurdity of this.
Why is having to name a function doing more harm than good?
Your second paragraph is a straw man.
There's nothing wrong with having to name a function. But there is something wrong with not having a choice and being forced to name a multi-line function. Just as you don't want to name every sub-expressions, you don't want to name every blocks of code. Unnecessarily naming things pollutes the namespace.
Where's the strawman? The premise was that functions that spans more than a line should be given a name because they are "unreadable". I have showed that, by this logic, anything that spans more than a line is "unreadable". You can't make special exceptions and say this only applies to lambdas. Lambdas (aka higher-order functions) are blocks of code just as much as anything else.
https://docs.python.org/3/tutorial/controlflow.html#lambda-e...
print((lambda x: [exec('x = x + 1; result = x'), eval('result')][1])(1))For the same reason that you wouldn’t necessarily define a function every time you wanted more than one line in the body of an if statement or for loop.
In a functional programming style, typically most of your control structures are represented as higher-order functions and you pass them other functions where you might use nested blocks of statements in an imperative style.
That is, where imperative pseudocode might look like this:
for n in [0..10]:
output_array[n] = n * n
some corresponding functional pseudocode might look more like this: output_array = for_each [0..10] (n -> n * n)
In some functional languages, it’s even idiomatic to write the supporting function so it reads more like this (borrowing Haskell’s $ notation, which avoids the awkward parentheses): output_array = for_each [0..10] $ \n ->
n * n
Now you might want a multiline body for the supporting function in much the same situations that you might want a multiline body for the imperative for loop: for n in [0..10]:
location = get_location(n)
route = fastest_route_to(location)
times[n] = average_journey_time(route)
times = for_each [0..10] $ \n ->
location = get_location(n)
route = fastest_route_to(location)
average_journey_time(route)
Similarly, if the logic you’re using in the inner part of the code is a self-contained concept, you might want to factor it out into a named function of its own in either case.This functional style can be tidy and very flexible in languages that are designed to support it. However, I wouldn’t necessarily encourage it in a language like Python, precisely because Python’s syntax and language features don’t make it natural and concise to write like this.
You can define a named function inside a named function and thus limiting its scope.
def outer_func():
def inner_func():
print("hi")
inner_func()For example, here's a multi-line lambda, containing only expressions:
>>> lambda f, pred, xs: [
... f(x) for ys in xs
... for x in ys
... if pred(x)
... ]
<function <lambda> at 0x1113e80d0>
Here's a one-line lambda, attempting to use a statement: >>> lambda x: (x += 1)
File "<stdin>", line 1
lambda x: (x += 1)
^
SyntaxError: invalid syntax
Note that the parentheses are to avoid parsing a '+=' expression with 'lambda x: x' on the left and '1' on the right: >>> lambda x: x += 1
File "<stdin>", line 1
SyntaxError: cannot assign to lambda
This restriction is problematic because of the following combination:- Python statements don't compose as well as expressions; we can write one statement followed by another, and we can nest statements (or blocks) inside control statements (if, else, with, etc.), but that's about it. We can't assign statements to variables, we can't call a function on a statement, or return a statement from a function, etc. i.e. we can't abstract over statements (unless we try wrapping them in a named function and using that as an expression, but that can break the underlying functionality of the statement, e.g. 'def ifThenElse(cond, true, false):' will evaluate both branches when called).
- Python uses statements for a whole bunch of stuff. In particular it has a 'with foo as bar: baz' statement, rather than e.g. a 'with(foo, lambda bar: baz)' expression; the recent pattern matching functionality is a statement rather than an expresssion; defining a named function is a statement; etc.
Taken together, we can often find ourselves needing to introduce a statement somewhere in our code; that requires us to turn a 'lambda' into a 'def'; but 'def' is a statement, so we have to turn any enclosing expression into a statement too; and this propagates up to disintegrate whatever expression we had; leaving us with a big pile of statements, full of boilerplate for managing scopes and names.
https://docs.python.org/3/library/itertools.html#itertools.g...
A bit more expressive overall.
At least it is better than lodash' useless groupBy which creates this weird key value mapping, loses order and converts keys to string and what not.
https://toolz.readthedocs.io/en/latest/streaming-analytics.h...
I never get this one right first time, but surely that's not it.
For Python coders new to functional programming, and how it can make working with data easier, I highly recommend reading the following sections of the pytoolz docs: Composability, Function Purity, Laziness, and Control Flow [0].
[0] https://toolz.readthedocs.io/en/latest/
The comment showing where groipBy, sortBy etc can be found just shows the problem - they are all in different libraries.That's just plain annoying! And don't get me started on the pain of trying to build an Ordered Dictionary with a default initial value!
>>> from collections import defaultdict
>>> from datetime import datetime
>>>
>>> d = defaultdict(datetime.now)
>>> d[1], d[2], d[3], d[4], d[0]
(datetime.datetime(2021, 7, 9, 15, 50, 52, 87605), datetime.datetime(2021, 7, 9, 15, 50, 52, 87613), datetime.datetime(2021, 7, 9, 15, 50, 52, 87614), datetime.datetime(2021, 7, 9, 15, 50, 52, 87615), datetime.datetime(2021, 7, 9, 15, 50, 52, 87616))
>>> for k,v in d.items():
... print(k,v)
...
1 2021-07-09 15:50:52.087605
2 2021-07-09 15:50:52.087613
3 2021-07-09 15:50:52.087614
4 2021-07-09 15:50:52.087615
0 2021-07-09 15:50:52.087616
>>>
Seems pretty good to me?What leaps out at me is that these 3 functions are all straight out of relational algebra style worldview.
Python the language doesn't support relational algebra as a first class concept. The reason it feels like IKEA self assembly is probably because you are implicitly implementing a data model that isn't how Python thinks about collections.
groupBy(fn): {fn(x): x for x in foo}
sortBy(fn): sorted(foo, key=fn)
countBy(fn): Counter(fn(x) for x in foo)
flatten: [x for y in foo for x in y]
filter(fn): [x for x in foo if fn(x)]
All the "for x in blah" can seem like a lot of boilerplate for a non-Pythonista, but actually it becomes subconscious once you're used to it and actually helps you feel the structure (a bit like indentation isn't necessary in C but it still helps you to see it).For compound operations (e.g. merge two lists, filter and flatten), I find the code to be a lot more easier to "feel" than if you'd combined several functional style functions, where you have to read the function names instead of just seeing the structure.
For example if
“def fn(i): return i//2”
You probably want to get an output like {0: [0,1]} but you’ll only get a scalar for the value (1) and it will be the last one evaluated.
It would be nice if there was a way to do this on one line but you’ll have to do
ret=collections.defaultdict(list) For i In foo: ret[fn(i)].append(i) Return ret
I suppose it isn’t that bad as it will be in the groupBy method.
Or you could add it to a method on a list. I think that’s easier to do with Rust.
g = {fn(x): [y for y in foo if f(y)==f(x)] for x in foo}
As you said, you'd have to use a for loop. I think sometimes people are unnecessarily afraid of for loops in Python (not everything has to fit on one line!) but this is a case it's a pity to take up so many lines. Or, back to square one, probably this is a case it's best just to use itertools groupby (after a sort) after all.It's not the boilerplate that's the problem, it's that using the same construct for everything doesn't convey its intention.
If I see a chain of filter, map, group etc., I instantly know what each part does. I know that filter does a filtering, and then I can look at the function passed to it do know how exactly it's filtering. If I see a list comprehension, I have to first understand if it's filtering or mapping or whatever (or even all at once), before I can grok it.
filter(fn): [x for x in foo if fn(x)]
Python has filter(predicate, iterable) -> iterable built in.> About 12 years ago, Python aquired lambda, reduce(), filter() and map(), courtesy of (I believe) a Lisp hacker who missed them and submitted working patches. But, despite of the PR value, I think these features should be cut from Python 3000.
("Python 3000" was the code name for Python 3.) He eventually stepped back from this hard line view but it remains a fact that it's a bit of a historical accident that it's present at all.
[1] https://www.artima.com/weblogs/viewpost.jsp?thread=98196
Most popular languages these days have aligned patterns for common collections operations. Python stands out IMO.
I wrote a bit more details in a parent comment (https://news.ycombinator.com/item?id=27782030)
I'm working on some python these days and I find quite unpleasant the 1-line lambdas, and general messy options to implement common collections operations.
from itertools import chain
flatten = chain.from_iterable
def lflatten(x): return list(flatten(x))
def sflatten(x): return set(flatten(x))I could tolerate a slightly different name, but `itertools.chain.from_iterable` is absurd to me.
If you reach for itertools imports often in an interactive REPL, you might be interested in Pyflyby: https://labs.quansight.org/blog/2021/07/pyflyby-improving-ef...