- Bare except: statements (that catches everything, even Ctrl-C)
- Mutables as default function/method arguments
- Wildcard imports!
- Bare except: statements (that catches everything, even Ctrl-C)
- Mutables as default function/method arguments
- Wildcard imports!
Given a function like:
def append_one(l=[]):
l.append(1)
return l
What does this return each time? >>> append_one()
>>> append_one()
>>> append_one()Still, I'd change this to something like:
def append_five(l=[]):
l.append(5)
return l
It tests the same thing (knowledge of how default parameters work), but without the confounding problem of similar-looking characters. Of course, syntax highlighting would help the applicant out.All of that being said, I still don't doubt that many developers don't know what they should about default parameters.
No. The default value gets "created" (the expression is evaluated and stored) when the def statement is executed. Take the following example:
In [1]: def foo():
...: def append_five(l=[]):
...: l.append(5)
...: return l
...: return append_five
...:
In [2]: a = foo()
In [3]: b = foo()
In [4]: a()
Out[4]: [5]
In [5]: b()
Out[5]: [5]
In [6]: _4 is _5
Out[6]: False
We only wrote one function definition, but multiple lists are created. (They are created when the "def append_five" definition executes, during the execution of foo.)If the candidate correctly deduces what will happen, I'll ask them to write a bug-free version, which looks like one of the below:
def append_one(var=None):
var = var or []
var.append(1)
return var
def append_one(var=None): if var is None:
var = []
var.append(1)
return var
Mutability is a very subtle but very important concept to understand in python. Everyone who uses python for non-trivial code should know it well: https://docs.python.org/2/reference/datamodel.htmlI don't think the question has much to do with mutability, it isn't surprising to me nor would I imagine most programmers that a list is mutable, that's very common.
The surprising part of this question is that the default value of 'l' continues to exist outside the lexical scope of the function, the expected behavior is that the value of 'l' is initialized at function call time and is garbage collected after each call. As it sits, using default values in python is sort of like defining a global that only has a named reference inside the function block, which is very strange.
Don't even get into unexpected behavior in classes:
In [1]: class A(object):
...: l = []
...:
In [2]: a, b = A(), A()
In [3]: a.l.append("Something")
In [4]: a.l
Out[4]: ['Something']
In [5]: b.l
Out[5]: ['Something']
In [6]: class B(object0:
...:
KeyboardInterrupt
In [6]: class B(object):
...: l = None
...: def __init__(self):
...: self.l = []
...:
In [7]: c, d = B(), B()
In [8]: c.l.append("Something")
In [9]: c.l, d.l
Out[9]: (['Something'], []) >>> for item in [1]:
... print item
1
>>> item
1
>>> for i in []:
... print i
>>> i
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
NameError: name 'i' is not defined
I would expect i == None. That oddity makes it dangerous to use the feature unless you're really careful (e.g. using a for - else construct). In [1]: for i in []:
...: pass
...: else:
...: print 'Else!'
...:
Else!
In [2]: for i in []:
...: break
...: else:
...: print 'Else!'
...:
Else!
In [3]: for i in range(2):
...: break
...: else:
...: print 'Else!'
...:
In [4]: for i in range(2):
...: pass
...: else:
...: print 'Else!'
...:
Else!
The syntax could be interpreted as: if len(l) == 0:
print "Else!"
else:
for i in l:
pass
The "catch cases where a `break` is triggered" case isn't common enough for this syntax feature to be encountered very often, leading to confusion when people come across it (though at least it's not a bug where a common use-case has weird behavior to new-comers).if you conceptualize how a for-loop has to work as a while-loop using Python's iterator protocol (which is the only way the iterator protocol itself makes sense), it seems pretty intuitive.
That is, this:
for item in items:
...1
else:
...2
becomes, approximately: try:
while True:
__hidden_iter = items.iter()
try:
item = __hidden_iter.next()
except StopIteration:
raise __NormalLoopExit
...1
except __NormalLoopExit:
...2
If you have an empty loop, the first assignment doesn't complete (instead raising StopIteration in evaluating the right side, which raises the notional exception __NormalLoopExit, which invokes the else: clause, if any) so the variable never gets around to being created.In most of the languages I'm familiar with, there are very clear syntax differences when working with class attributes. For example, in many languages class attributes have to be accessed via the class name instead of from an instance of the class making it clear to the programmer they are working with a class attribute, e.g. MyClass.myClassVariable not myInstance.myClassVariable. Additionally, the way you define class attributes in python is the way you define instance attributes in many languages, which just adds to the confusion. e.g. in Java or C# you can define class variables directly in the class body, but an explicit 'static' keyword is needed, undecorated definitions are assumed to be instance variables.
Finally, I think the definition of class B above is a little more nuanced, class B has both a class attribute named l AND an instance attribute named l.
B.l == None and B().l == []
It's been a while since I've done major OOP coding in any language other than Python, so I'm a little rusty. The issues you raise are perfectly legitimate and would be understandably confusing to newcomers to the language. :)
If the object was immutable then append wouldn't work. That's hardly matching expectations.
I guess the clarification to what I was saying is that, in the simple case (integers, strings, None) the objects are immutable. It's only getting into cases where the value of the object itself is mutable, that you run into issues. If all objects (or all objects 'allowed' as default values) were immutable, then this behavior would not trigger.
So saying that mutability has nothing to do with it isn't entirely true. It's the immutability of the types of values used in most simple cases that hides this issue from developers until they run into a more complex case.
if val is None: val = []
or the more idiomatic python way:
val = val or []You want to explicitly check against `None` so that you're not overwriting all falsey values of `val` - even though you should generally try to enforce argument types, your second example would cause unexpected behavior in some cases, particularly those that have non-falsey 'default' assignments
>>> [1,2,3] + [4,5]
[1, 2, 3, 4, 5]
Thus appending should do something different than addition. >>> x = [1,2,3]
>>> x.append([4,5])
>>> x
[1, 2, 3, [4, 5]]I wonder if people who weren't exposed to languages which work differently ala C++ would be as surprised?
I understand mutability and immutability in other languages (and I gave your link a quick read to make sure there weren't any weird Python-specific rules), so I understand how the list can change and still be the same object, but a tuple or string would not. But why does that mean that the default parameter object remains in existence throughout all calls, instead of being recreated each time it is called?
Is there a reason for this being the default behavior? It seems like the majority of the time you would want to use a default parameter, you'd want it to behave like your bug-free examples.
While that explains how it works, I actually completely agree with you. This is surprising behavior and, in a language that prides itself on not being surprising, seems, well, surprising.
I have to wonder if performance isn't the big reason for it. If your default is [], it isn't a big deal to re-evaluate, but if your default is get_default_cities_from_slow_web_service(), having that re-evaluated on every function call would be catastrophic. Given the choice between two negatives, the choice they made is probably reasonable.
Before I ever ask this question (I do a lot of tech interviews sadly) I always ask the candidate about object mutability vs immutability. Almost everyone knows the textbook answer, and only a few know the actual implications of it. This tests which they know :)
Default kwargs of a function are defined at function definition. However, they are only in scope, for the scope of said function. It is a weird but important subtle difference.
def append_one(var=None):
return (var or []) + [1]
Would this take longer and/or use more storage for long lists as vars?And besides all that, there is nothing wrong with doing an append on one line, and returning the variable on the next. It's clear and readable.
my_list = []
append_one(my_list)
# my_list didn't get anything appended to it
This shows up another subtle trap related to the "truthiness" (or falsiness in this case) of things like the empty list.Because that's how it works in a lot of other languages, such as Ruby and Javascript.
because the behavior is the same, whether or not they misread an 'l' as a 1.
In [1]: a = []
In [2]: a.append(a)
In [3]: a
Out[3]: [[...]]
In [4]: a[0]
Out[4]: [[...]]
In [5]: a[0][0]
Out[5]: [[...]]
In [6]: a[0][0][0]
Out[6]: [[...]]
In [7]: a[0][0][0][0]
Out[7]: [[...]]
In [8]: a.append(a)
In [9]: a
Out[9]: [[...], [...]]
In [10]: a[0][1][0] is a
Out[10]: True
In [11]: id(a)
Out[11]: 4547140064
In [12]: id(a[0][1][0])
Out[12]: 4547140064I know python is not unique in having warts like this, but it's pretty b.s. in general that unexpected behavior is just thought to be okay, especially in a language meant to be very accessible, and most especially since it's being used as a perfectly valid metric for disqualifying new python programmers from employment.
If the industry as a whole cared about evidence-based, non-superstitious, non-monoculture-reinforcing hiring practices, we'd realize that tripping people up and judging programming capability based on minutia is as unfair as it is self-defeating.
It is simply a easy way to gauge a candidate's proficiency with the language. It also helps if they know that this is a problem. You'd be shocked to know a lot of people on the market for jobs writing python don't get this question correct, but the smart ones often do when talking through it even if they didn't originally.
In my experence your much better off with people that look at odd syntax and say, "I don't know what that does" vs those who do.
Using the default value in some capacity isn't that uncommon... Though maybe you were speaking to a more general case? for example, decoding a an obfuscated C file.
a_list_of_words = "my list of words".split(" ")
I never enquired why, since there were bigger issues in the code e.g. "unit testing" by running the code, taking the result and putting it as the check value. By running repr(value), copying out the string then comparing self.assertEqual(repr(value), '[<Object1: unicode_value>, ...]')
df_subset = df['date buyer nwidgets'.split()]
That is far easier to type than the explicit list, with all its punctuation. Now, it's definitely weird that they did a `split(" ")` rather than just using the default, but the idea is the same.I do try to strip stuff like that out before I put it into a script, replacing it with the explicit list, but I'm never sure if that actually improves anything. It's not as if the explicit list is any easier to read.
In [1]: "a string r".split(" ")
Out[1]: ['a', '', '', '', 'string', '', '', '', '', 'r']
In [2]: "a string r".split()
Out[2]: ['a', 'string', 'r'] >>> filter(None, " quick hack for split".split(" "))
['quick', 'hack', 'for', 'split']From the comment a few levels up I understood that the code which used the str.split with " " argument didn't signify that someone who written it knew about its semantics. If he did and it was really what was intended then ofc it's completely ok, but if not, it can easily lead to bugs.
For example, if the user is required to input several ints separated with whitespace, this:
map(int, input_str.split())
will rise only in expected cases, while this: map(int, input_str.split(" "))
can lead to rejecting correct input just because someone pressed space twice. It's very frustrating for the user, too, because whitespace are hard to spot visually.So, I don't know if this qualifies as antipattern, but I think if I saw .split(" ") instead of .split() in the code I'd at the very least expect the comment explaining why it's used.
(That sentence I wrote about using hashable types need not apply, sorry!)
If you ran that code, you would get this error:
TypeError: list indices must be integers, not listInefficent, or bizarre way to do it maybe.
Anti-pattern is supposed to mean something more, though.
In this case there are no adverse effects and no ambiguity -- so, I guess the programmer was just lazy to construct the list.
a_tuple_of_words = ("my", "tuple", "of", "words")
or
a_tuple_of_words = "my", "tuple", "of", "words"
But yes, Python misses entirely the point of tuples, treating them as read-only lists.
http://dozzie.jogger.pl/2014/04/11/python-tuples-the-useless...
Lighter-weight, immutable collections have a use case. The code in OP appears to be one where it makes sense. I follow the rule where variables are mutable IFF they need to be mutable.
For the rest of the world, tuples are not immutable lists. They are tuples, i.e. collections of "objects" that could share nothing about their type. Tuples often are not even iterable! (Erlang, Haskell)
The fact that tuples in Python can have as much structure as one wants is derived from dynamic typing, not from the tuples' nature. The same you could say about Python's lists.
This is a really subtle issue. It takes to know more languages to see it clearly.
Would you feel better if they named it "ImmutableList" instead?
(Although I agree with you that statically-typed-language-tuples don't seem to make sense in Python.)
But hey... Python's weird choice of how to name the ImmutableList could be worse, right?
For example, someone could be malicious enough to call their general-purpose associative array a "hash", just because a hashmap (note: not a hash) is a good implementation for large associative arrays. Wow, that'd be hilariously misleading, wouldn't it? Good times!
Or imagine someone was silly enough to name their auto-resizing arrays "vectors", even though in all previously existing contexts a "vector" is a sort of thing which absolutely cannot be meaningfully resized/extended. Ha. Think of the tiny cognitive burden placed on generations of future programmers-who-study-math, trying to juggle these two very-similar-but-distinct concepts, multiplied by the number of such future programmers. Amazing practical joke, right?
/rant
Yes, I would feel better if it was named "ImmutableList" or any other way that is not misleading about the purpose.
1. Iterate over a tuple
2. Convert a list to a tuple
3. Construct a tuple of a length not known at compile-time
Python allows these because "why not?" but it does break their "one and only one way to do it" rule and confuses beginners a hell of a lot.
There are definitely borderline cases. For instance, should a Vector be a list or a tuple? A Vec3 type is obviously a tuple, but a large Vector destined for BLAS is obviously a list.
No, it allows them because the distinction that those restrictions are founded on is only useful in a statically-typed languages, and Python isn't statically typed.
> For instance, should a Vector be a list or a tuple?
A real vector/array should be its own data type (probably implemented in a C, or similar low-level, extension) that happens to implement the interface expected of an indexable, iterable collection, neither a list nor a tuple.
...like Erlang.
There is a deep difference that goes beyond use of tuples in language approach between Python and Erlang here where it comes to types in which Erlang, while dynamically typed, has a deep concern for types in its pattern matching system to make path decisions while Python is very much centered on using dynamic OO techniques -- how objects respond to messages -- to do that.
So I'd still say its the same kind of deep language approach difference at work.
A typical rule of thumb in Python land is that heterogeneous data probably belongs in a tuple, so practice goes a little further than immutable lists.
I think you could improve your demonstration of the usage in the standard library by examining a random selection of usages to try to find out what is typical. But maybe you already looked at more than you talk about in the article (and I understand that this might not be an interesting use of your time).
> A typical rule of thumb in Python land is that heterogeneous data probably belongs in a tuple, so practice goes a little further than immutable lists.
The problem with Python tuples is it's two things mixed: immutable lists and a container for heterogenous data. It's the same situation as JavaScript's objects.
Even in Haskell, though, people often write all kinds of type-class magic to allow "iterating" over a tuple. For example, a Binary instance over a tuple wants to call "put" on each element.
Haskell's (Oleg's) HList is basically a tuple with iteration/list-like operations.
No, a structure is, you know, a structure -- what C calls a struct. Python calls it a namedtuple. If some people call it just a tuple, well, that's a difference in terminology, but it doesn't mean Python is confused about the concepts, it's just using terminology you're not used to.
Also, if we're going to be pedantic about the meaning of data types, your blog post is wrong about lists. You say "position in the list doesn't matter", but that means ordering doesn't matter, and an unordered collection of similar objects is a set, not a list. Python makes this distinction clear: a list is ordered, a set is not.
No, it menas exactly this. The term "tuple" and its use predates Python. Sorry, no banana.
> [...] your blog post is wrong about lists. You say "position in the list doesn't matter", but that means ordering doesn't matter
Oh, so what's the difference in meaning of element True on position 1 and element True on position 20? Position in list doesn't matter if we're talking about meaning of the elements.
References, please? And not mathematical references; programming references. C was using the keyword "struct" long before Python to refer to what you are calling a tuple.
> what's the difference in meaning of element True on position 1 and element True on position 20?
The fact that the index is 1 instead of 20. Both elements have the same type, and might well refer to the same property of some sequence of things; but the index being 1 instead of 20 means the element True is describing that property relative to the first item in some sequence, instead of the 20th item. That's why position in the list makes a difference: the ordering of the items, as well as the type of the items, carries information.
(Of course, in Python the list items don't even have to be of the same type; but most uses of Python lists in practice that I've seen do assume that all the elements are "the same kind of thing".)
ML has had tuples several decades before Python existed.
@a_list_of_words = qw/my list of words/;
there
a_list_of_words = %w{my list of words} >>> qw = str.split
>>> qw('my list of words')
['my', 'list', 'of', 'words']It would really make sense to change the semantics of Python to fix this issue.
def foo(default_arg = []):
Why can't that just be shorthand for: def foo(default_arg = ParamNone):
if default_arg == ParamNone:
default_arg = []
How would that break first class functions?What breaks is something like:
def foo(default_arg = slow_f()):
pass
Under the shorthand gets turned into: ParamNone = object()
def foo(default_arg = ParamNone):
if default_arg is ParamNone:
default_arg = slow_f()
pass
This is fine, since everyone would know that the shorthand means to not put slow code there. Instead, people will start writing it as: _foo_arg = slow_f()
def foo(default_arg = _foo_arg):
pass
Of course, then what happens with: _foo_arg = slow_f()
def foo(default_arg = _foo_arg):
_foo_arg = 5
? Under expansion it becomes: _foo_arg = slow_f()
def foo(default_arg = ParamNone):
if default_arg is ParamNone:
default_arg = _foo_arg
_foo_arg.add(5)
This violates Python's scoping rules, because _foo_arg is now being used in local scope instead of global scope. Eg: >>> def f(x=None):
... if x is None:
... x = spam
... spam = 3
...
>>> spam = 9
>>>
>>> f()
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
File "<stdin>", line 3, in f
UnboundLocalError: local variable 'spam' referenced before assignment
Which means you now need a new scoping rule, just to handle default parameters without making things more confusing.It also turns what was a simple O(1) offset into a precomputed list into a globals() lookup for many cases.
The default arguments thing is worse than a lot of the stuff Python 3 corrected.
More specifically, given a better 'default arguments thing', how would you interpret:
x = [2]
def f(x=x*5):
x.append(4)
With the earlier conversion it's: x = [2]
def f(x=DefaultArg):
if x is DefaultArg:
x = x*5
x.append(4)
This isn't going to work because the x inside of f() is different than the outside x, and you'll get the error message I mentioned.If you add a nonlocal, as in:
x = [2]
def f(x=DefaultArg):
nonlocal x
if x is DefaultArg:
x = x*5
x.append(4)
then you'll get "SyntaxError: name 'x' is parameter and nonlocal".What other solution are you thinking of?
def foo(default_arg = [0]*(256*256)):
...
and still get the same namespace issues.Memoization is not always going to be an available solution. For example, it may be that slow_f() returns a stateless object, so can be reused, while slow_f(x) returns something stateful. You can think of my examples as either using default arguments as a single element memo, or using a module variable for the same. Both premised on the idea that the developer knows enough to make the right decision.
There are other dynamic languages with functions as first class objects which don't share the "mutable default arguments" gotcha.
But having said that, any change regarding this would break backward compatibility.
Example: nose.tools
For example, this is a real method in one of my projects:
def listen(self, address, ssl=False, ssl_args={}):
pass
I like the way this turns up in the docs because it's immediately clear that ssl_args needs to be a dict. Otherwise I have to describe it in words.Why not just add @param annotations in your docstrings instead?
If they need to touch this argument in an overridden method and they don't know what they are doing, then yes.
> Why not just add @param annotations in your docstrings instead?
I'm using Sphinx and it renders them separately. I want the empty dict to show up in the function signature.
There are other ways to emphasize it ought to be a dict/mappable. Change its name to be suffixed as "_dict", for example?