Functools – The Power of Higher-Order Functions in Python
martinheinz.dev
martinheinz.dev
Years ago I found out the Guido wouldn't let tail recursion included, and even tried to remove map function from built in functions. Therefore I got the impression that python wouldn't have further support in FP. I really wish that is not the case. With the coming pattern matching in 3.10, My hope is high again.
I have very high respect for Guido van Rossum, I'm not here to discredit him. He's one of the authors in PEP 634.
I wish python would have simpler syntax for lambda which is currently using the lambda expression, even JS is doing better on this. A built-in syntax for partial would also be great. It could be even better if we can have function composition.
Some problems are better solved the FP way. It could even make the program more readable which is one of the strengths of Python.
The irrational hatred of FP is still there, to the point of the absurdity of implementing a hobbled procedural version of pattern matching!
It is bordering on insane.
Pattern matching took so long because it was extremely hard to find that compromise, and I actually think they did an okay job. No, it won’t cover every case, but it will also cover the important ones.
x = match y:
case 0: "zero"
case 1: "one"
case x:
y = f(x)
y+y
print(x)
it can require semicolons or parentheses or even braces somewhere if that's necessary to make the grammar work, idc. just don't make me type out `x = ...` a hundred times...(yes, i know you could use the walrus operator in the third case)
(also, if anyone wants to reply with "use a named function because READABILITY", please save it)
In general, breaking Python's syntax rules to introduce block expressions or what you want to call it seems like a steep price to pay for what amounts to syntactic sugar in the end.
However, your proposal actually misses the mark in the same way as the actual match/case does as well. What would be /really/ handy in Python is something like
{'some_key': bind_value_to_this_name} = a_mapping
because it is so common, especially if you're consuming a JSON API. You of course see the myriad problems with the above.I think you should read the PEPs on match/case, these things have actually been considered. One idea was to introduce a keyword to say which variables are inputs and which are outputs, but that also violates the general grammar in a way that isn't very nice.
The accepted solution has warts, I agree, but at some point you just have to accept that you can't reach some perfect solution.
i'm at peace with that :) [1]
> In general, breaking Python's syntax rules to introduce block expressions or what you want to call it seems like a steep price to pay for what amounts to syntactic sugar in the end.
that's your opinion, i disagree! i think Python would be a better, more pleasant to use language if they figured this out. a girl can dream, ok?
> What would be /really/ handy in Python is something like
{'some_key': bind_value_to_this_name} = a_mapping
yes, extending the destructuring syntax would be nice, i agree! but the original question was "Please suggest a pattern matching expression syntax [...]", and the post you were responding to was specifically talking about `match`'s syntax. [2]> I think you should read the PEPs on match/case
i read them when they came out. i'm assuming you're referring to this section from PEP 622:
> "In most other languages pattern matching is represented by an expression, not statement. But making it an expression would be inconsistent with other syntactic choices in Python."
https://www.python.org/dev/peps/pep-0622/#id77
i believe Python's statement-orientatation kinda sucks, so to me this is just "things aren't great, let's stay consistent with that". yeah, yeah, "use a different language if you don't like it" etc.
---
[1] well, maybe not quite, after all i'm here arguing about it.
[2] `match` uses patterns like the one you described in `case` arms, but is distinct from them. i don't see why dict-destructuring syntax couldn't be added in a separate PEP.
looks like i messed up my PEPs -- i linked to PEP622, which was superseded by PEP635 [1]. the gist of it hasn't changed though.
---
[1] https://www.python.org/dev/peps/pep-0635/#the-match-statemen..., Ctrl+F "Statement vs. Expression"
I'm less interested in the syntax (will take anything accepted by a modified ast module), more in the semantics.
Here's a syntax I played with in the past:
def area(s: Shape):
pmatch(s):
Circle(c):
pi * c.r * c.r
Rectangle(w, h):
w * h
;;
def times10(i):
x = pmatch(i):
1:
10
2:
20
;;
return x
Two notes: * Had to be pmatch, since match is more likely to run into namespace collision with existing code.
* The delimiter `;;` at the end was necessary to avoid ambiguity in the grammar for the old parser. Perhaps things are different with the new PEG parser.
Code is on my github.https://github.com/santinic/pampy
Clearly it's possible. It's also more ergonomic than PEP 622.
Mutations everywhere. But mainly because mapping and flatting stuff is so burdensome. And it lacking many fp functions and a proper way to call them ergonomically makes a crazy list comprehension the goto tool, which often is much less explicit about what's going on than calling a function with a named and widely understood concept.
I've said roughly this before somewhere on HN but cannot find it right now: list comprehensions are like regexes. Below a certain point in complexity, they're a much cleaner and immediately-graspable way of expressing what's going on in a given piece of code. Past that point, they rapidly become far worse than the alternative.
For example, "{func(y): y for x, y in somedict.items() if filterfunc(x)}" is clearer than the equivalent loop-with-tempvars, and significantly clearer than "dict(map(lambda k: (func(somedict[k]), somedict[k]), filter(filterfunc, somedict.keys()))))". Even with some better mapping utilities (something like "mapkeys" or "filtervalues", or a mapping utility for dictionaries that inferred keys-vs-values based on the arity of map functions), I think the "more functional" version remains the least easily intelligible.
However, once you cross into nested or simultaneous comprehensions, you really need to stop, drop, and bust out some loops, named functions, and maybe a comment or two. So too with regex! Think about the regex "^prefix[.](\S+) [1-7]suffix$"; writing the equivalent stack of splits and conditionals would be more confusing and easier to screw up than using the regex; below a point of complexity it's a much clearer and better tool to slice up strings. Past a point regexes, too, break down (for me that point is roughly "the regex doesn't fit on a line" or "it has lookaround expressions", but others may have different standards here).
But even the lambda keyword isn't so bad, you can create a dictionary of expressions to call by name, a lot more compact them declaring them the usual way imo: https://github.com/jazzyjackson/py-validate/blob/master/pyva...
To your point, I only recently learned there's a Map function in Python, while in JS I'm .map(x=>y).filter(x=>y).reduce(x=>y)ing left and right.
> But even the lambda keyword isn't so bad, you can create a dictionary of expressions to call by name, a lot more compact them declaring them the usual way imo: https://github.com/jazzyjackson/py-validate/blob/master/pyva...
lambda keyword is better than nothing, it definitely can be improved. Just imaging using javascript syntax in your example.
> To your point, I only recently learned there's a Map function in Python, while in JS I'm .map(x=>y).filter(x=>y).reduce(x=>y)ing left and right.
I think with the introduction of list comprehension Guido saw map function was no longer needed, that was why he wanted it removed. I don't deny it, but using map and filter sometimes are just easier to read. Say [foo(v) for v in a] vs map(foo, a).
[1]: https://stackoverflow.com/questions/1247486/list-comprehensi...
I think composition and piping are such basic programming tools that make a lot of code much cleaner. It's a shame they're not built-in in Python.
So, shameless plug, in the spirit of functools and itertools I made the pipetools library [0]. On top of forward-composition and piping it also enables more concise lambda expressions (e.g. X + 1) and more powerful partial application.
One gripe that I have with functions like map is that it returns a generator, so you have to be careful when reusing results. I fell into this trap a few times.
I'd also like a simpler syntax for closures, it would make writing embedded DSLs less cumbersome.
I hope that is never changed; I often write code in which the map function's results are very large, and keeping those around by default would cause serious memory consumption even when I am mapping over a sequence (rather than another generator).
Instead, I'd advocate the opposite extreme: Python makes it a little too easy to make sequences/vectors (with comprehensions); I wish generators were the default in more cases, and that it was harder to accidentally reify things into a list/tuple/whatever.
I think that if the only comprehension syntax available was the one that created a generator--"(_ for _ in _)"--and you always had to explicitly reify it by calling list(genexp) or tuple(genexp) if you wanted a sequence, then the conventions in the Python ecosystem would be much more laziness-oriented and more predictable memory-consumption wise.
Ah well, water under the bridge, I know.
First you claim functional languages are scary to look at, then say that you want Python to become more like functional languages. But maybe the reason Python is elegant and easier to read is exactly because Guido had the self-constraint to not go full functional. You also miss the part of the story where he actually did remove `reduce` from builtin, exactly because of how unreadable and confusing it is to most.
It's exactly that kind of decision making I expected from a BDFL, and I think Guido did a great job while he was one keeping Python from going down such paths.
Functional programming is kind of cool unless it's someone else's code and you need to debug it.
Any decision he made is infinitely more than I could. Because I am just a python user, and an outsider in any decision making process. So for me, he's right all the time. That's a perfect definition of a dictator :)
But I do have wishes. It's like I love my parents but I do want to stay up late sometimes.
Yeah, I totally missed the part he removed reduce from builtin. Sorry about my memory. map, filter, or reduce, it does not matter. As I stated, some problems are better solved functional way. Because Python is such a friendly language, if it includes functional paradigm properly, it would make the functional part more readable than other functional languages.
FP is scary not because it has evil syntax to keep people at distance, it's just an alien paradigm to many. Lot's of non functional languages has functional support, which doesn't make them less readable. E.g. C#, JS. I suspect these languages have helped many understanding FP more. Python could make the jump by including more FP, but not turning into a full-fledged FP.
BTW. I'm still glad reduce is kept in functools.
I guess my argument is that Python is friendly exactly because it uses FP paradigms sparingly and conscientiously. Some paradigms such as map and filter can definitely make your code cleaner, while others (such as reduce) only lead to headaches. That being said as you mention we are getting pattern matching, I'm curious to see how that ends up.
i would argue that C# and JS do also use it sparingly and don't go too hard, but also that Python is generally a cleaner language than those. It's hard to know how much of it can be attributed to what, but generally, it's not a trivial thing to predict what impacts a given language feature will have on the readability of the code. A given FP paradigm might be cool in isolation, but when stacked with 4 other ones, it can quickly lead to gibberish code.
You are right, C# and JS use FP sparingly. I think I expect more from Python because I enjoy writing in Python and my functional adventures were rooted in Python.
Personally, lambda expression is painful to look at. Many times I have deliberately avoided it. It's possible to make improvements, but it seems the FP part just stagnated. Looking forward to pattern matching though.
from itertools import chain
flatten = chain.from_iterable
def lflatten(x): return list(flatten(x))
def sflatten(x): return set(flatten(x))I could tolerate a slightly different name, but `itertools.chain.from_iterable` is absurd to me.
- groupBy is itertools.groupBy(lst, fn)
- sortBy is just lst.sort(key=fn)
- countBy is collections.Counter(map(fn, lst))
- Sibling comment mentioned flatten, which is just [item for sublist for sublist in lst]
More esoteric needs are usually met by itertools.
You mean sorted(). list.sort only works on list and in-place.
You can make lambdas in Python, but naming function is in my mind very much a "feature".
Same applies to then and else branches.
In a code review I don't what to read a thesis inside of a conditional branch.
Assuming you have five levels nested if, ignoring the fact that you might be doing something wrong, then you should name those, but that requires you to extract those code paths and put them into separate functions.
A function can do literally anything. Seeing a function call, it hasn't narrowed down the realm of what might be happening even a little bit.
"lambda f:" gives away very little information about what is about to happen. "if f:" gives away some information about what is about to happen.
Not a perfect rule of thumb, convoluted if statements do exist. But that is why functions are more in need of hints and comments than other control structures.
I think that's an unfair comparison:
If 'the thing that happens' for "if f:" is 'branch on the truthyness of f', then I agree that's pretty clear. However, if we're going to ignore the content of the branches then we should also ignore the body of the function, so "lambda f:" is also pretty clear: 'parameterise with f'.
If we're worried that "lambda f:" could do pretty much anything in the body, then we should also worry that "if f:" can do pretty much anything in the branches.
That is the part common to both. If you take that common part out, you will be left with more information in the if statement than the function call.
In the function, you know what arguments it (might) be interested in. You know they are evaluated. But in the if statement, you know what values it is interested in, usually get the same amoutn of information about whether they are evaluated, much clearer guarentees about how they are used and that the aspect of them that matters is their truthiness. Also the code might be about to branch.
Put it this way - if(...) is a specific - and common - named function. It is a very easy argument to make that it is carrying more information about what is going on than a general anonymous function.
I don’t think this argument generalises well. Once you start working with structured data and more complicated control logic, a generic loop with an arbitrary body often provides less immediate information to a developer reading the code than a function that explicitly names a specific pattern of computation like map or filter.
There are many recurring patterns of computation that are much more specific than a generic map or filter, too. While imperative languages tend to have only a small number of built-in control structures, if we’re using higher-order functions then we can name as many of these patterns as we like.
For example, suppose we want to generate a list of the first few triangle numbers, which is a sequence of the form:
[0, 0+1, 0+1+2, 0+1+2+3, ...]
Writing this in imperative Python code with a typical loop would give us something like this: triangle_numbers = []
total = 0
for n in range(6):
total = total + n
triangle_numbers.append(total)
Here’s an idiomatic functional version, which is real Haskell code using a standard library function that captures this pattern of sequence generation in just five characters: triangle_numbers = scanl (+) 0 [1..5]
Assuming you’re equally familiar with both programming styles, I think it’s fair to say that “scanl” tells you much more about what this code is doing than “for”.The inner function here is just addition in both cases. However, the code within the Python loop is allowed to do whatever it wants, so the reader has to recognise the pattern surrounding the addition. In the Haskell version, the function you pass to scanl must take an accumulated value so far and a next value as parameters and it must return the next accumulated value to add to the sequence being generated, and that is all it is allowed to do. So you just need to write (+), which is Haskell’s notation for what Python calls operator.add or an equivalent lambda, and that tells the reader exactly how the next accumulated value is being derived at each step to build up the sequence.
triangle_numbers = scanl(f, 0, [1,2,3,4,5])
gives me scant little information about what the value of triangle numbers is at the end of the operation. triangle_numbers = if(f, 0, [1,2,3,4,5])
or triangle_numbers = if(true, 0, f)
tells me that whatever f is something weird is going on. Commenting if statements to point out that they are branching code is foolish.My intended point was that comparing lambda: to if: is apples-to-oranges.
Both styles in my example have an outer structure that builds a sequence, wrapped around an inner calculation that says how you build the next value in the sequence. The inner calculation here was simple addition, a one-liner in the for loop and (+) in the Haskell version. However, it could equally have been something more complicated.
If it were, a reader would still have to figure out what the body of the for loop was really doing. In contrast, the (possibly anonymous) function passed into scanl would still have a more constrained purpose, which is immediately known because scanl represents one specific pattern of computation. The implicit algorithm it represents has a hole that is a certain shape, and you can only pass it a function with the right shape to fit into that hole.
Put another way, if you’re programming in this functional style, it’s not writing the word lambda that tells you what the next piece of code is for, it’s passing that lambda expression into a higher-order function like scanl. Passing an anonymous function is like writing the nested block for a branch of an if statement or the body of a for loop. It’s the call to scanl that is roughly analogous to writing the if or for statement itself.
Of course this assumes the lambda is pure, but in the kind of code we're talking about (pseudofunctional code with lots of higher-order functions) they should be pure for the most part.
An if statement is a lower level mechanism for directing the code path within a function. If it needs a name, wrap it in a function. I often do.
It also has a condition and a body to execute conditionally.
if (condition)
body
is just a goofy way of writing if(condition, body)Many Swing coders can also apply.
It is so hard to actually pack the functionality into a separate piece of code.
"Yeah, but then I have to look elsewhere"
Really? What about when a method is full of stuff like that?
Its readability is lost beyond hope.
If the idea of multi-line alone compromises readability, then why even have multi-line blocks at all? Function, loop and conditional blocks should all at most have one line each. I hope you can see the absurdity of this.
Why is having to name a function doing more harm than good?
Your second paragraph is a straw man.
There's nothing wrong with having to name a function. But there is something wrong with not having a choice and being forced to name a multi-line function. Just as you don't want to name every sub-expressions, you don't want to name every blocks of code. Unnecessarily naming things pollutes the namespace.
Where's the strawman? The premise was that functions that spans more than a line should be given a name because they are "unreadable". I have showed that, by this logic, anything that spans more than a line is "unreadable". You can't make special exceptions and say this only applies to lambdas. Lambdas (aka higher-order functions) are blocks of code just as much as anything else.
https://docs.python.org/3/tutorial/controlflow.html#lambda-e...
print((lambda x: [exec('x = x + 1; result = x'), eval('result')][1])(1))For the same reason that you wouldn’t necessarily define a function every time you wanted more than one line in the body of an if statement or for loop.
In a functional programming style, typically most of your control structures are represented as higher-order functions and you pass them other functions where you might use nested blocks of statements in an imperative style.
That is, where imperative pseudocode might look like this:
for n in [0..10]:
output_array[n] = n * n
some corresponding functional pseudocode might look more like this: output_array = for_each [0..10] (n -> n * n)
In some functional languages, it’s even idiomatic to write the supporting function so it reads more like this (borrowing Haskell’s $ notation, which avoids the awkward parentheses): output_array = for_each [0..10] $ \n ->
n * n
Now you might want a multiline body for the supporting function in much the same situations that you might want a multiline body for the imperative for loop: for n in [0..10]:
location = get_location(n)
route = fastest_route_to(location)
times[n] = average_journey_time(route)
times = for_each [0..10] $ \n ->
location = get_location(n)
route = fastest_route_to(location)
average_journey_time(route)
Similarly, if the logic you’re using in the inner part of the code is a self-contained concept, you might want to factor it out into a named function of its own in either case.This functional style can be tidy and very flexible in languages that are designed to support it. However, I wouldn’t necessarily encourage it in a language like Python, precisely because Python’s syntax and language features don’t make it natural and concise to write like this.
You can define a named function inside a named function and thus limiting its scope.
def outer_func():
def inner_func():
print("hi")
inner_func()For example, here's a multi-line lambda, containing only expressions:
>>> lambda f, pred, xs: [
... f(x) for ys in xs
... for x in ys
... if pred(x)
... ]
<function <lambda> at 0x1113e80d0>
Here's a one-line lambda, attempting to use a statement: >>> lambda x: (x += 1)
File "<stdin>", line 1
lambda x: (x += 1)
^
SyntaxError: invalid syntax
Note that the parentheses are to avoid parsing a '+=' expression with 'lambda x: x' on the left and '1' on the right: >>> lambda x: x += 1
File "<stdin>", line 1
SyntaxError: cannot assign to lambda
This restriction is problematic because of the following combination:- Python statements don't compose as well as expressions; we can write one statement followed by another, and we can nest statements (or blocks) inside control statements (if, else, with, etc.), but that's about it. We can't assign statements to variables, we can't call a function on a statement, or return a statement from a function, etc. i.e. we can't abstract over statements (unless we try wrapping them in a named function and using that as an expression, but that can break the underlying functionality of the statement, e.g. 'def ifThenElse(cond, true, false):' will evaluate both branches when called).
- Python uses statements for a whole bunch of stuff. In particular it has a 'with foo as bar: baz' statement, rather than e.g. a 'with(foo, lambda bar: baz)' expression; the recent pattern matching functionality is a statement rather than an expresssion; defining a named function is a statement; etc.
Taken together, we can often find ourselves needing to introduce a statement somewhere in our code; that requires us to turn a 'lambda' into a 'def'; but 'def' is a statement, so we have to turn any enclosing expression into a statement too; and this propagates up to disintegrate whatever expression we had; leaving us with a big pile of statements, full of boilerplate for managing scopes and names.
https://docs.python.org/3/library/itertools.html#itertools.g...
A bit more expressive overall.
At least it is better than lodash' useless groupBy which creates this weird key value mapping, loses order and converts keys to string and what not.
https://toolz.readthedocs.io/en/latest/streaming-analytics.h...
I never get this one right first time, but surely that's not it.
If you reach for itertools imports often in an interactive REPL, you might be interested in Pyflyby: https://labs.quansight.org/blog/2021/07/pyflyby-improving-ef...
The comment showing where groipBy, sortBy etc can be found just shows the problem - they are all in different libraries.That's just plain annoying! And don't get me started on the pain of trying to build an Ordered Dictionary with a default initial value!
>>> from collections import defaultdict
>>> from datetime import datetime
>>>
>>> d = defaultdict(datetime.now)
>>> d[1], d[2], d[3], d[4], d[0]
(datetime.datetime(2021, 7, 9, 15, 50, 52, 87605), datetime.datetime(2021, 7, 9, 15, 50, 52, 87613), datetime.datetime(2021, 7, 9, 15, 50, 52, 87614), datetime.datetime(2021, 7, 9, 15, 50, 52, 87615), datetime.datetime(2021, 7, 9, 15, 50, 52, 87616))
>>> for k,v in d.items():
... print(k,v)
...
1 2021-07-09 15:50:52.087605
2 2021-07-09 15:50:52.087613
3 2021-07-09 15:50:52.087614
4 2021-07-09 15:50:52.087615
0 2021-07-09 15:50:52.087616
>>>
Seems pretty good to me?What leaps out at me is that these 3 functions are all straight out of relational algebra style worldview.
Python the language doesn't support relational algebra as a first class concept. The reason it feels like IKEA self assembly is probably because you are implicitly implementing a data model that isn't how Python thinks about collections.
groupBy(fn): {fn(x): x for x in foo}
sortBy(fn): sorted(foo, key=fn)
countBy(fn): Counter(fn(x) for x in foo)
flatten: [x for y in foo for x in y]
filter(fn): [x for x in foo if fn(x)]
All the "for x in blah" can seem like a lot of boilerplate for a non-Pythonista, but actually it becomes subconscious once you're used to it and actually helps you feel the structure (a bit like indentation isn't necessary in C but it still helps you to see it).For compound operations (e.g. merge two lists, filter and flatten), I find the code to be a lot more easier to "feel" than if you'd combined several functional style functions, where you have to read the function names instead of just seeing the structure.
For example if
“def fn(i): return i//2”
You probably want to get an output like {0: [0,1]} but you’ll only get a scalar for the value (1) and it will be the last one evaluated.
It would be nice if there was a way to do this on one line but you’ll have to do
ret=collections.defaultdict(list) For i In foo: ret[fn(i)].append(i) Return ret
I suppose it isn’t that bad as it will be in the groupBy method.
Or you could add it to a method on a list. I think that’s easier to do with Rust.
g = {fn(x): [y for y in foo if f(y)==f(x)] for x in foo}
As you said, you'd have to use a for loop. I think sometimes people are unnecessarily afraid of for loops in Python (not everything has to fit on one line!) but this is a case it's a pity to take up so many lines. Or, back to square one, probably this is a case it's best just to use itertools groupby (after a sort) after all.It's not the boilerplate that's the problem, it's that using the same construct for everything doesn't convey its intention.
If I see a chain of filter, map, group etc., I instantly know what each part does. I know that filter does a filtering, and then I can look at the function passed to it do know how exactly it's filtering. If I see a list comprehension, I have to first understand if it's filtering or mapping or whatever (or even all at once), before I can grok it.
filter(fn): [x for x in foo if fn(x)]
Python has filter(predicate, iterable) -> iterable built in.> About 12 years ago, Python aquired lambda, reduce(), filter() and map(), courtesy of (I believe) a Lisp hacker who missed them and submitted working patches. But, despite of the PR value, I think these features should be cut from Python 3000.
("Python 3000" was the code name for Python 3.) He eventually stepped back from this hard line view but it remains a fact that it's a bit of a historical accident that it's present at all.
[1] https://www.artima.com/weblogs/viewpost.jsp?thread=98196
Most popular languages these days have aligned patterns for common collections operations. Python stands out IMO.
I wrote a bit more details in a parent comment (https://news.ycombinator.com/item?id=27782030)
I'm working on some python these days and I find quite unpleasant the 1-line lambdas, and general messy options to implement common collections operations.
For Python coders new to functional programming, and how it can make working with data easier, I highly recommend reading the following sections of the pytoolz docs: Composability, Function Purity, Laziness, and Control Flow [0].
[0] https://toolz.readthedocs.io/en/latest/
https://towardsdatascience.com/functools-the-power-of-higher...
from functions import *
…on the front and @cached_property
def troll_face(self):
return ‘:)’
…on the back.functools and itertools are amazing and I love them both. They are especially useful for teaching high school CS without having to stray from Python, which the kids at all levels know well and are comfortable with.
However.
Using @cached_property feels like a bad code smell and it’s a controversial design decision. The example given is more like an instance of a factory? Perhaps if the result of the render function was an Important Object™ (not just a mere string) then the function’s callers might not be so cavalier about discarding the generated instance.
I would like to see the calling code that calls `render` more than once in two different places (hence the need for caching with no option for cache invalidation?). When every stack frame is a descendant of main() then there’s no such thing as “two different places”!
def main():
page = HN().render()
while ‘too true’:
do_work(page)
tea_break(page)
Ironically it doesn’t feel very functional.What web server is this? I’ve never heard about this behavior before.
This is different than for instance java, which there is one process and then multiple threads. But because of the GIL in python, only one thread can ever run at the same time (in the same process), therefore one has to launch multiple instances/processes of the application.
What if we have multiple containers running the same service and load balanced using an ingress controller?
Local Caches are a headache in such a scenario.
Best pattern I have read is use a cache as a sidecar in a pod. I feel that scales very well.
Why?
It also doesn’t really matter. The only time things like this have actually been important was in giant Java codebases, where attempts were made to keep packages cleanly separated (and so therefore passing the simplest possible public types reduces the amount of public classes).
FP gains are real. Anyone who tells you otherwise doesn't know what FP is.
To anyone out there who isn't on the FP train yet, get on. You'll become a better programmer for it.
The last time I was surprised was the itertools library.
Also check https://pymotw.com/3/. It's a great tour of the Python standard library
>>> pow(5, -1, 13)
8
>>> 5 * 8 % 13
1- Lack of pipe operator
- Multiple arguments everywhere instead of currying
- No do-notation or equivalent
- Reference equality rather than structural equality for objects etc.
If you want to program in this style, consider using F#, R or even JavaScript with a few Babel plugins.
Combine it with Lodash FP: https://github.com/lodash/lodash/wiki/FP-Guide
Chaining functions is somewhat annoying in python if you want to do any sort of cross library work (which is super common in the data side of the language). Pandas has nice function chaining, but if you want chain data through cleaning functions and maybe do some nlp work you don't end up with a very readable or nice syntax IMO.
> Coconut is a functional programming language that compiles to Python. Since all valid Python is valid Coconut, using Coconut will only extend and enhance what you're already capable of in Python.
> Why use Coconut? Coconut is built to be useful. Coconut enhances the repertoire of Python programmers to include the tools of modern functional programming, in such a way that those tools are easy to use and immensely powerful; that is, Coconut does to functional programming what Python did to imperative programming. And Coconut code runs the same on any Python version, making the Python 2/3 split a thing of the past.
If you want a real dynamic language functional programming experience, you would be far better with Elixir.
That’s not proper caching.
Was the issue memory and storage?
Much better off using the same assignment semantics used throughout Python and let people choose to deepcopy() all the objects they read from the cache if they really really want to modify them later; it would even work as a simple decorator to stack onto the existing one for such cases.
Aside: If we’re missing “deep coping” anything by default in Python, it’s definitely default parameters :) deep copying caches makes much less sense than that.
Is that really what's going on here? I'm in way too deep to be sure what's best for beginner programmers, but I feel like Python must surely optimise...
sheep = 1
goats = sheep
sheep = sheep + 10
... by simply copying that 1 into goats, rather than tracking that goats is for now an alias to the same value as sheep and then updating that information when sheep changes on the next line.Now, if we imagine those numbers are way bigger (say 200 digits), Python still just works, whereas the low level languages I spend most time with will object because 200 digit integers don't fit in a machine word. You could imagine that copying isn't cheaper in that case, but I don't think I buy it. The 200 digit integer is a relatively small data structure, still probably cheaper to copy it than mess about with small objects which need garbage collecting.
I don't think it does that. You can try it yourself:
>>> x = 123456
>>> y = x
>>> id(x)
139871830996560
>>> id(y)
139871830996560For these particular numbers, CPython has one optimization I know of: Small integers (from -5 to 256) are pre-initialized and shared.
See https://docs.python.org/3/c-api/long.html#c.PyLong_FromLong
This is generally invisible to the python programmer, but leads to https://wsvincent.com/python-wat-integer-cache/
On my system these "cached" integers seem to each be 32 byte objects, so, over 8kB of RAM is used by CPython to "cache" the integers -5 through 256 in this way.
There's similar craziness over in String town. If I mint the exact same string a dozen times from a constant, those all have the same id, presumably Python has a similar "cache" of such constant strings. But if I assemble the same result string with concatenation, each of the identical strings has a different id (this is in 3.9)
So, my model of what's going on in a Python program was completely wrong. But the simple pedagogic model in the article was also wrong, just not in a way that's going to trip up new Python programmers.
Not really. You're hitting constant folding: https://arpitbhayani.me/blogs/constant-folding-python
This isn't a pre-made list of certain strings that should be cached, this is the compiler noticing that you mentioned the same constant a bunch of times.
Also in general you would see a lot of things with the same id because python uses references all over the place. E.g. assignment never copies.
The behavior that results from your example follows from this rule. First, evaluate `1` and call it `sheep`. Then evaluate whatever `sheep` is, once, to get `1` (the same object in memory as every other literal `1` in Python) and call it `goats`.
The last line is where the rule matters: The statement `sheep = sheep + 10` can be read as, “Evaluate `sheep + 10` and call the result `sheep`.” The statement reassigns the name `sheep` in the local namespace to point to a different object, one created by evaluating `sheep + 10`. The actual memory location that `sheep` referred to previously (containing the `int` object `1`) is not changed at all — assignment to a local will never change the value of any other local.
This is easy to remember if you recall that a local namespace is effectively just a `dict`. Your example is equivalent to:
d = {}
d["sheep"] = 1
d["goats"] = d["sheep"]
d["sheep"] = d["sheep"] + 10
It should be clear even to beginners that `d["goats"]` has a final value of `1`, not `11`, because the right-hand side of `d["goats"] = d["sheep"]` is only evaluated once, and at that time it evaluates to `1`. Assignment using locals behaves in exactly the same way. @lru_cache(post=deepcopy, ...)
to have a new (copied) instance in cases where that's required. Or do whatever else you needed. Maybe you happen to know some details about the returnee that would let you get away with copying less and could call my_copy instead.Something something power of function composition.
`lru_cache` does not wrap your cached values in container types on its own, so if you’re getting bit by mutable values it’s probably because what you cached was a mutable container (like `list`, `dict`, or `set`). If you use primitive values in the cache, or immutable containers (like `tuple`, `namedtuple`, or `frozenset`), you won’t hit this problem.
If you were to manually cache values yourself using a `dict`, you’d see similar issues with mutability when storing `list` objects as values. It’s not a problem specific to `lru_cache`.
Edit: Can you show me a caching backend other than local memory (Redis, for example) that implements this behavior?
Imagine you `lru_cache` a zero-argument function that sleeps for a long time, then creates a new temp file with random name and returns a writable handle for it. It should be no surprise that a cached version would have to behave differently: Only one random file would ever be made, and repeated calls to the cached function would return the same handle that was opened originally.
If you didn’t expect this, you might complain that writes and seeks suddenly and mysteriously persist across what should be distinct handles, and maybe this `lru_cache` thing isn’t as harmless as advertised. But it’s not the cache’s fault that you cached an object that can mutate state outside the cache, like on a filesystem.
Logically, the behavior remains the same if instead of actual file handles we simulate files in memory, or if we use a mutable list of strings representing lines of data. Generalizing, you could cache any mutable Python collection and see that in-place changes you make to it, much like data written to a file, will still be there when you read the cache the next time.
The reason you don’t see “frameworks” for this is because tracking references to instantiated Python objects outside of the Python process is pointless — objects are garbage-collected and are not guaranteed to stay at the same memory location from one moment to the next. Further, if the lists themselves are small enough to fit in memory, surely there’s no need for out-of-memory scale to cache simple references to those objects.
The only way you can have this confusion is if the actual cached object could be somehow returned. That only happens when everything in the cache is actually just a Python object in your same process.
I disagree- it works wonders for expensive functions that return strings and ints. Maybe "don't use lru_cache with mutable values" would be more accurate.
Thanks for noting that edge case though, I've never thought about that.
If you want to reduce resource consumption, you could zip the results, or deduplicate the values using a separate dict.
There’s nothing unPythonic about a function returning mutable values.
A common use for caching is wrapping a call to a remote servicer. If the response is JSON, it’s likely to be mutable.
Nothing weird or edgy about that.
The dev creates a PR.
The cache is a decorator. It always returns a CachedValue object. The CachedValue object has an API for reading and writing the returned value.
If you write to the CachedValue object it will update Redis with the new, altered value.
It updated Redis not only for the value derived from that key, but all identical values.
This behavior was not included in the ticket. It is not mentioned in the documentation.
It only occurs when the value is mutable.
What do you say to the developer? Job well done? Or what the f*ck is wrong with you?
I’m a nice person so I’d say something nice. I’d ask what they were thinking.
But if their reasons was “this is how caches works” I’d take a serious look at their resume and ask whether they might need some mentoring.