A few things to remember while coding in Python
satyajit.ranjeev.in
satyajit.ranjeev.in
The problem with mutable defaults is that they are evaluated once only when the function is defined. Each time the function is called you'll be using the same mutable variable that was created during function definition.
def f(x=0, y="foo", z=3.14159):
This, however, is a perfectly Pythonic idiom: def f(L=None):
if L is None:
L = []
[1]: http://docs.python.org/reference/compound_stmts.html#functio... def f(L=()):
... >>> def f(x):
... return x + 5
>>> type(f)
<type 'function'>
>>> dis.dis(f.func_code)
2 0 LOAD_FAST 0 (x)
3 LOAD_CONST 1 (5)
6 BINARY_ADD
7 RETURN_VALUE
>>> g = f
>>> g(10)
15
>>> g is f
True
They're just variables in the current scope. If you're quite clever, your brain is already figuring out that has some interesting implications that some libraries use: >>> import socket
>>> socket.gethostbyname('www.google.com')
'74.125.71.103'
>>> socket.gethostbyname = lambda i: '10.0.0.1'
>>> socket.gethostbyname('www.google.com')
'10.0.0.1'
I will not pass judgement on monkey patching like this, just pointing out it's doable. I know for a fact Ruby can as well.Functions just being variables has useful properties when you're doing something like fancy switch/case type things (the readability of this is questionable, but it's cool to look at it, like a Duff's device):
>>> i = 1
>>> { str: func1, unicode: func2, int: func3 }.get(type(i), func1)(i)
in func3
Also consider something like this, which is how decorators work (and they're incredibly useful), sort of like a closure: >>> def maker(i):
... def ret(x):
... return x + i
... return ret
>>> f, g = maker(10), maker(100)
>>> f(5), g(10)
(15, 110)
I'd be surprised if Ruby couldn't do everything I just did.By useful he means mucking up your program in totally unexpected ways.
He's being nice about it being a silly decision to have it behave that way. The entire post reads more like a list of unexpected things that will bite you in the ass.
You can use a mutable default argument as an ersatz static variable, e.g. for memoization.
Maybe I misunderstood you, but they are perfectly mutable:
class A
attr_accessor :a
def initialize
@a = []
end
def b x = a
x << 1
end
end
obj = A.new
# => #<A:0x007fd92b2e9850 @a=[]>
obj.b
# => [1]
obj.b
# => [1, 1]
obj.a
# => [1, 1]
If you mean "inline default arguments are not mutable", that's not true either. What is true is that the default argument is evaluated when the function is called, not when it is defined: a = 0
# => 0
x = lambda {|y = (a + 1)| y }
# => #<Proc:0x007fd92b1dae50@(irb):33 (lambda)>
x[]
# => 1
a = 5
# => 5
x[]
# => 6 ruby> class A; end
=> nil
ruby> def foo(bar = A.new); return bar; end
=> nil
ruby> foo
=> #<A:0x00000101985060>
ruby> foo
=> #<A:0x0000010197a6b0>
and get back the same object each time `foo` is called in Ruby.I've even seen a major Python library with this bug (I'm sorry, I don't recall which off-hand). It's really surprising behavior for new Python devs.
Basically, sometimes you do want to reuse the mutable between function calls, and in those cases it can save a fair bit of code passing it in repeatedly.
def calculate(a, b, c, memo={}):
try:
value = memo[a, b, c] # return already calculated value
except KeyError:
value = heavy_calculation(a, b, c)
memo[a, b, c] = value # update the memo dictionary
return valuedef f(L=None): L = L or []
These are the little assumptions that keep blowing off my feet. Thanks.
def f(x=None): x if x is not None else [] def f(x=None): if x is None: x = []That said, people still seem to favor your form as the more Pythonic way. Personally, I think that's just because the ternary expression is relatively new.
But if this expression is a line in a larger function, and is intended to reset the value of argument x if no other value is passed in for it, does this really act as an assigment to x? Because I sort of read this expression as evaluating to some value -- the passed value for x, or a [] -- but does this assign that value to the argument x? Or must it be x = [expression]
def f(x=None):
x = x if x is not None else []
return x
Now it will assign that value back to x. Otherwise it would just evaluate the expression. L = [] if L is None else LNot to say I think python should be changed on this point. It shouldn't, there are code checkers that warn you on the gotcha, let's use them.
If you say "f(x=[])" (assuming that worked without the actual side effects it has), someone could still say "x(None)" instead of "x()", causing the function to die. Since a robust program isn't able to avoid checking for None, it might as well set defaults there too.
There is another case where this is important; you might want the equivalent of "f(x=expensive_function_to_calculate_useful_default())", and you don't want that function called unless it needs to be. Only the x=None approach allows this to be deferred.
In the expensive case, I'd just calculate it once and store it somewhere (possibly as a lookup dictionary if there are multiple inputs) and access that from within the function.
But None is a result that can happen in situations that would otherwise return exactly the expected type. If "nothingness" can be meaningful (especially in a function that accepts an empty list as a parameter, say), it's nicer if the code just deals with None itself instead of requiring checks for None in all the callers.
> (defvar *fn*
(let ((x 3))
(lambda (&optional (y (list nil x)))
(push 7 (car y)) ; modifies the list
y)))
*FN*
> (funcall *fn*)
((7) 3)
> (funcall *fn*)
((7) 3)
From this example you can see two things. First, the binding of 'x' is closed over when the lambda expression is evaluated. And second, the expression that provides the default value of 'y' is evaluated every time the function is called.There's no fundamental reason it couldn't have worked that way in Python. (I understand that changing the language so it worked that way now would likely break some code.)
EDIT: fixed formatting.
http://stackoverflow.com/questions/1651154/why-are-default-a...
The code that evaluates the default expression doesn't need to be in a separate function, either, so the argument that calling that function is too expensive also doesn't hold water.
I just tried a test in SBCL:
(defun foo1 (x) x)
(defun test1 (n) (dotimes (i n) (foo1 (cons nil nil))))
(time (test1 100000000))
=> 4.4 sec, or 44ns / iteration
(defun foo2 (&optional (x (cons nil nil))) x)
(defun test2 (n) (dotimes (i n) (foo2)))
(time (test2 100000000))
=> 4.1 sec, or 41ns / iteration
The version with the optional parameter is actually slightly faster, which completely blows a hole in the performance argument.Look, no language is perfect -- not even Common Lisp :-) I think users are better served when design flaws in a language are acknowledged without defensiveness than when bogus justifications are offered.
I don't agree that this is a design flaw. As I recall it bit me once as a beginner, and never again in over a decade of using python, and as a lisp hacker you know you don't design a language for beginners. :-)
The equivalent Python would be something like this:
def function():
x = 3
def internal(x, foo=[]):
foo.append([7])
foo.append(x)
return foo
return internal(x)
print function()
print function()
Which does what you would expect: [[7], 3]
[[7], 3]You've still got that &optional argument though. I don't see a huge amount of difference from a semantic point of view between that and the Python version though (ie. if x == None: ...).
fn = lambda y=[y]: y.push(7); return y
if you accept the ; to separate statements, as the lambda in python is syntactically only allowed to contain one statement.(The introduction of the variable x into the example is not important for the behavior of default arguments, however, it is important for a separate issue. I've stripped it out here.)
To the curious - SO has an explanation of why Python was designed like this, which I found interesting: http://stackoverflow.com/questions/1132941/least-astonishmen...
Actually, this is not a design flaw, and it is not because of internals, or performance. It comes simply from the fact that functions in Python are first-class objects, and not only a piece of code.
Why in Common Lisp defaults behave the way one would expect, then? Functions are also first class, but defaults are evaluated at every call.
def foo():
def bar():
...
return bar
foo() == foo() # false
Two different function objects are created. If, in the above examle, bar took a pram thelist=[], each call to foo would produce a bar function with a different list instance for thelist. The default values can be read as expressions passed to the function object constructor, rather than a bit of code to be evaluated each function run.I don't know how CL works in this regard, nor do I know which is better or worse. I think the explanation linked did a terrible job conflating first class functions with execution and runtime models. Some of the answers below it explain better tho. :)
When the function represented by the lambda expression is applied to arguments, the arguments and parameters are processed in order from left to right. (...) If optional parameters are specified, then each one is processed as follows. If any unprocessed arguments remain, then the parameter variable var is bound to the next remaining arguments, just as for required parameter. If no arguments remain, however, then the initform part of the parameter specifier is evaluated, and the parameter variable is bound to the resulting value (...).
The CLTL2 specifies that the form representing the default value of optional parameter shall be evaluated every time the parameter is not provided.
def f(seq=[]):
for x in seq:
# do somethingPass all code under pylint scrutiny, comply to its complains or adjust its rules, do it early. That is the recommendation I wish all devs could read.
They can also be good. Here's an example from the Reddit discussion, showing how a mutable default can be used to very neatly and cleanly add memorization to a function:
def fib(n, m={}):
if n not in m:
m[n] = 1 if n < 2 else fib(n-1) + fib(n-2)
return m[n]I won't say it would be clear anyone that reads it, because we live in a world where people who claim to be programmers can't do fizz buzz.
Add a @memoize decorator and do it there, you need to always be as obvious as possible. Compare:
@memoize
def fibonacci(n): pass
def fibonacci(n, memory=[]): pass
You don't even need documentation for the first example. def fib(n, m:"donotusethisparameter"={}): freqs = {}
for c in "abracadabra":
try:
freqs[c] += 1
except:
freqs[c] = 1
If this is really the common idiom, I'd say this is a sign that professional programming has yet to fully mature as a field.Some may say a better solution would be:
freqs = {}
for c in "abracadabra":
freqs[c] = freqs.get(c, 0) + 1
Okay, so I understood immediately what was going on with the 2nd bit of code.Rather go for the collection type defaultdict
from collections import defaultdict
freqs = defaultdict(int)
for c in "abracadabra":
freqs[c] += 1
As a non-pythonista, the 3rd bit of code, I had to Google "defaultdict" to figure out. It's only a couple of seconds to Google, and a professional should know this tidbit, but it seems like premature optimization to me. This brings to mind this post:http://news.ycombinator.com/item?id=3995185
As a programmer, one's most valuable resource is brainpower. Supposedly, a programmer's most important goal is writing clear code. Look around at what goes on in our industry. There's a lot of our most valuable resource spent on showing off our cleverness, not directed towards the clearest code. To me this is like spending money to show one can spend money or playing an instrument to show off dexterity instead of producing gorgeous sounds.
(I think this starts in school and other environments where one is motivated to show off one's coding chops.)
Most of the complexity in our field accrues like litter: a bit here and a bit there. I think it says something about the culture of the folks who live there.
So you spent a couple seconds Googling defaultdict. Great, you now know what a defaultdict is and can use it in your own code. It's useful in a lot of places besides this toy example.
You should avoid gratuitous complexity, where you force the reader to learn something that will never, ever be useful to them again. A great example might be writing your own encryption algorithm, which will be complicated, wrong, under-performant, and totally useless on any other project. If you just use bcrypt (or whatever the recommended best practice is now), then your code works well, and all readers of your code now know about bcrypt and can use it themselves.
I have a different set of policies than most programmers, which arises from my observation that our field's priorities are out of whack with the actual cost-benefit.
Our greatest costs involve understanding systems, so our first priority should typically be to produce readable and understandable code.
You shouldn't avoid language features just because some people don't know about them.
One should pick language features to optimize for readability, which is entirely contextual. If your shop has a culture of using ?: to the point where it's like a coding standard then you should keep on doing that.
So long as code can be read and understood, programmers will learn. Better yet, if the culture of a shop is that use of language features and other tools are motivated by contextual cost-benefit, then programmers will learn from this example. As it is, programmers generally are more interested in showing off, having fun, and writing things as easily as possible. It's less common to have a culture of prioritizing reading.
return [i/2 for i in nums if not i % 2]
"not i % 2"? That's beautiful? We do a division and ask whether the remainder is not true? What does it mean for a remainder of a division to be not true? Is sqrt(3) untrue? Is 17 not yellow? Is this really the clearest way to say even number?Yes, after years of working in C, I'm well aware of how C does bools, but that's because C's values are small and fast, not beautiful and clear. In C, there's barely any abstraction to leak--you just manipulate bits and don't worry about mixing your metaphors. But Python has different priorities, which is why I prefer it to C when my users (and I) won't be hurt by the performance difference.
If we're showing off clarity instead of cleverness, wouldn't this be a better way to demonstrate it:
return [i/2 for i in nums if i%2 == 0]
And I prefer the concept of "clarity" to "beauty" when it comes to code. Beauty is, well, whatever in the eyes of the beholder. Clarity, for me, is the question of how fast code can be read and understood correctly by a given programmer who is familiar (not more) with the language, not necessarily familiar with other languages, and unfamiliar with what the code does.The faster such a person can skim the code and understand it correctly, the easier it will be to modify and keep free of bugs. If we're so smart, why don't we show it by using our brains to write code that is quicker to read and understand correctly than code written by lesser lights?
And I tested with the `timeit` module, there's not really much of a performance difference either (although the version with `not` is slightly faster, it's just by less than a percent or so).
What metrics do you use to determine if code is readable or not? What metrics do you use to determine "actual cost-benefit"?
Well, if one were to take into account hard metrics for every 4 line snippet of code, then I'm not sure enough would get done fast enough. In the context of everyday programming and of the 3 examples I quoted, it's enough to ask yourself questions like: what would a newbie understand? What would an average programmer recognize immediately? If there's a quick obvious answer to either of those questions, and no onerous externalities involved, then that's what you write. (The code being 10X longer or too slow or too likely to contain bugs would be an onerous externality.)
This isn't to say that metrics aren't useful here. The question is how to apply them at a low enough cost. In a large company, perhaps one could A/B test variations on coding standards. In a diverse group of small programming shops with internally consistent coding standards or styles, one might gather metrics on bugs per line of code versus features of coding styles.
Something to think about. It's not as if metrics are commonly used to make these decisions now. Either edicts come from on high, or the local alpha-coder declarers what's best in her/his experience.
from collections import Counter
freqs = Counter("abracadabra")
I was surprised to see that missing, given that Counter was mentioned in the next section.Here's my (now mostly obsolete) version:
def count(string):
counts = {}
for item in set(string):
counts[item] = string.count(item)
return counts
print count("abracadabra")
I had a look into the collections library, and it just uses iterable.iteritems(). I suspect that this might be faster for larger strings with multiple repeating characters, since set() and count() will pass the string directly to C.In fact, I almost came here to write a parallel comment: I'm really not sure that reversing the list 'a' with 'a[::-1]' is better than 'reversed(a)', which usually effectively does the same thing, but whose meaning is much more obvious.
But, while I agree with your general point, in the specific case of 'defaultdict', I differ.
I use 'defaultdict' all the time and I'm glad its there. It feels cleaner than 'freqs.get(c,0)'. I define the default value in one place, and then the interface to my datastructure is simpler; hence as I continue writing, I can spend more of my brainpower in the problem domain.
Its a small detail, but its one less thing to think about when writing a complex algorithm.
Actually, this speaks to point: optimization should be for reading, not for writing.
>>> reversed([1, 2, 3])
<listreverseiterator object at 0x107265650>
>>> [1, 2, 3][::-1]
[3, 2, 1]
To clarify your point, list(reversed(a)) and a[::-1] are equivalent. It's a slightly subtle point, but extremely important if you're keeping the result of reversed() around for any length of time. If you're just iterating at the moment that you use it, yes, they're effectively equivalent.1). For the most part, "c = collections.Counter()" is almost always better than "c = defaultdict(int)"
* Counter only supplies missing values rather than automatically inserting them upon lookup.
* The Counter version is much clearer about what it is trying to do. The defaultdict version is cryptic to the uninitiated (understanding it entails knowing that it has a __missing__ method to insert values computed by a factory function and that int() with no arguments returns zero).
* The Counter version provides helpful methods such as "most_common(n)".
2). An ellipsis in Python is normally used in a much different way than shown in the article (it's used for an extended slice notation in NumPy).
halve_evens_only = lambda nums: map(lambda i: i/2, filter(lambda i: not i%2, nums))
I still find it rather silly that python doesn't supper a nice list map/filter; it could be so much nicer nums.filter(lambda i: i%2 == 0).map(lambda i: i/2)
If they did, even including the annoyingly long-to-type "lambda". List comprehensions are cool and all, but do not really scale visually (i.e. get rather messy) when you have more than one map and filter step.These arbitrary break-away from OO method style into module+data style (len(L) is another!) are one of the things I hate most about Python. There are some reasons for doing so, but a pure-OO (like Scala) or pure-method+data (like F#) would have saved me many a runtime error.
nums.filter!(i => i%2 == 0).map!(i => i/2);
stuff.filter(_ % 2 == 0).map(_ / 2)
and C#, demonstrating that this isn't some obscure feature that only language geeks care about: Stuff.Where(x => x % 2 == 0).Select(x => x / 2)
My point isn't that this sort of syntax is new and novel, it's just that in Python it's annoyingly inconsistent. There are reasons where you would want to use type-class style modules to structure your code in a certain way, but I do not think python's map() filter() reduce() and len() qualify as these cases halve_evens_only = (i / 2) for i in nums if (i%2 == 0)
The parens aren't necessary, but they help readability for people who aren't used to the generator order of operations. (Again, Lisp geek, more parens means more readable in my fractured mind.)You can omit them in the generator expression if it's being passed directly as the only parameter to a function:
halve_evens_only = list(i/2 for i in nums if i % 2 != 0) halve_evens_only = map (/2) . filter even(lambda i: i%2 == 0).filter((lambda i: i/2).map(nums))
makes (marginally) more sense. But I like filter and lambda alright as they are. I agree Python's inconsistency in this is a bit unfortunate, but I haven't had much of a problem with the runtime errors you mention.
impl methods<T> for [T] {
fn filter(f : fn(T)->bool) -> [T] { ... }
fn map<U>(f : fn(T)->U) -> [U] { ... }
}
And then you can call it with syntax like: println(#fmt("%?", [ 1, 2, 3 ].map { |x| x + 3 }));
// prints "[ 4, 5, 6 ]"
The methods are properly scoped, so code that isn't in your module needs to import your methods to use them. That way, you avoid introducing strange action-at-a-distance in your code. varname, = [x for x in l if predicate_with_single_truth_value(x)]
The comma after varname is an implicit assert that the list comprehension only contains one element.This sort of code would be very confusing when I'm just quickly reading through a procedure trying to find the potential bug.
(varname,) = [x for x in l if predicate_with_single_truth_value(x)] [varname] = [x for x in l if predicate_with_single_truth_value(x)]varname ,= [...]
Superficially it looks like an operator, but I suspect that's merely because of whitespace freedom; i.e., a, = [0] is equivalent to a,=[0] and a ,= [0].
http://stackoverflow.com/questions/1642028/what-is-the-name-...
Sure it is! And in Python 3 there's the ,_*= operator, similar to lisp's car:
varname ,_*= [1, 2, 3] # varname == 1Even Python removed one of its most prominent cases of this, the print command. In Python 2.x you could have a trailing comma after a print to omit the new-line but Python 3's print() function requires print('something', end=' ') to be more explicit about it.
[1] http://stackoverflow.com/questions/118370/how-do-you-use-the...
def foo():
...
It is really highly unusual and I wouldn't recommend the practice shown in the blog post at all as this is not a common pattern.The recommendation on `iteritems` had better be generalized to include `iterkeys`, `itervalues`, and other opportunities for using iterators rather than building lists. A note that the 'iter...' versions are removed in Python 3 (because iterator behaviour becomes the default) would be appropriate here.
In relation to collections, itertools is a great module to get familiarized with. I import * from this module. I consider functions there as if they were builtins.
"Conditional assignment" is a weak and misleading name. "Conditional expressions" is more descriptive. There is no assignment in
print "yes" if some_condition else "no"
as the article acknowledges later.Using Ellipsis for getting all items is a violation of the Only One Way To Do It principle. The standard notation is [:].
I commend the good intentions of the writer, but I'm surprised that this article got 144 upvotes in HN.
Reversed produces an iterator and not a copy with all the dangers of mutable semantics. The [::-1] syntax is not equivalent as it returns a copy of the list as if it were reversed, changing up somewhat what it's doing.
You'd better remember that you should know what you're talking about before you start dictating to other people how they should write their code.
Generally, most use this:
freqs = {}
for c in "abracadabra":
try:
freqs[c] += 1
except:
freqs[c] = 1
Who does this?! if not os.path.exists("foo"):
os.mkdir("foo")
That introduces a race condition. If foo does not exist on the first line but is created by something else on the second line then this will raise an exception. The proper code is: import errno
try:
os.mkdir("foo")
except OSError as exc:
if exc.errno != errno.EEXIST:
raise
That doesn't have the race condition.(Yes, that's wordy. Eventually they plan to add to Python 3 a fancier exception hierarchy described at http://www.python.org/dev/peps/pep-3151/ , which lets you filter at a more fine-grained level, as below.)
try:
os.mkdir("foo")
except FileExistsError:
passTo solve this particular problem in future code more compactly however, note that Python 3.2 (finally) adds an "exist_ok" Boolean keyword parameter to the multi-directory variant, os.makedirs(). In other words, calling os.makedirs("mydir", exist_ok=True) will silently ignore existing directories and only raise if other errors occur.
freqs = {}
for c in "abracadabra":
if c in freqs:
freqs[c] += 1
else:
freqs[c] = 1Since this is posted by the you as the author, I'll comment here: some JavaScript is running on page load that blanks the entire page in Safari and iCab on iPad, making the page turn white except for the bullet symbols, and making the article unreadable.
I was able to read it only by disabling JavaScript or parsing it with Readability. (Both disable my ability to comment about this bug there.)
Python: ?
Ruby: ?
Ruby: Twitter.