Python idiom for taking the single item from a list
blog.garlicsim.org
blog.garlicsim.org
def get_single(l):
assert l and len(l) == 1
return l[0]
Then you get the best of both worlds: readability and a concise one-liner.Also, the performance here is probably much worse. (Although in many cases it would not matter.)
But that's just my opinion, your suggestion is legitimate.
This is an argument against using functions at all. If it works against get_single, it works against all functions.
That is to say, it doesn't work at all.
A function that squares a number is simple, one that computes the standard deviation is definitely more complex.
I'm talking about the complexity difference between this:
def stddev(pop):
total = 0
count = 0
for x in pop:
total += x
count += 1
mean = total / float(count)
variance = 0
for x in pop:
variance += (x - mean)**2
return math.sqrt(variance)
and this: def stddev(pop):
return math.sqrt(variance(pop))
def variance(pop):
m = mean(pop)
return sum(square(x - m) for x in pop)
def mean(pop):
return sum(pop) / float(len(pop))
def square(x):
return x**2
The first is a (mildly) complex function. The latter are all simple functions, and the complex result is constructed by composing simple operations.Good programmers write functions in the latter style, not the former.
Well, that stddev function could be much less verbose:
def stddev(pop):
mean = sum(pop) / float(len(pop))
variance = sum( (x-mean)**2 for x in pop)
return math.sqrt(variance)
To me that's easier to read than jumping back and forth between multiple function definitions. Of course, if you need the mean or variance independently then your way is better.Sure, it could, but I was demonstrating what it looked like without the use of functions. Your example proves my point just as mine does: sum(), like mean() or variance() in my example, is just a simple function, the kind that I'm arguing for. The fact that it's built into Python (rather recently, I note) doesn't change that fact or reduce the impact of the argument. Your example simply goes one step down the path, and mine goes further.
> To me that's easier to read than jumping back and forth between multiple function definitions.
You don't have to jump back and forth between function definitions. Let's say you don't know what the standard deviation is, but you know what the mean is. You can look at the definition of stddev() and see, "Ah, it's clearly the sqrt() of the variance. What's the variance? Ah, it's the sum of the squares of difference between each element and the mean." You know what the mean() does (its name is pretty clear) and you know what sum() and square() do, so you never have to look at those functions. Someone else who knows what the variance is would never have to look that deep. When someone is reading the stddev() in my example, he doesn't have concern himself with implementation details of functions he already understands. When someone is reading yours, he has to at least read how the mean is calculated. He can't avoid it--it's right there.
> Of course, if you need the mean or variance independently then your way is better.
You almost certainly will in any case where you're using the standard deviation, but that's just an artifact of the example. Other advantages of using small, simple functions like in my example:
* More reusable (as you noted) * More easily testable. * More easily comprehensible (as I showed above) * More easily documented (especially in a language like Python with its docstring support) * More conceptual abstraction
def get_single(l):
i = iter(l)
val = i.next()
try:
i.next() # expected to throw exception for one-element iterable
except StopIteration:
return val
raise AssertionError('More than one object')
Eww. assert s and len(s)==1
return s.pop()
Or if you want to stay in the immutable land: assert s and len(s)==1
return tuple(s)[0] >>> x = (lambda: (yield 1))() # generator with one step
>>> y = tuple(x)[0]
>>> y
1
>>> list(x) # exhausted
[]
>>> x = (lambda: (yield 1))()
>>> y, = x
>>> y
1
>>> list(x) # exhausted, too!
[]
Because that's just what you inevitably need to do to fetch a value from a generator. There is no peeking action or some such.This is redundant: if a list's length is 1, then it's true in a boolean context.
Also, please stop naming your lists 'l'. On a vast array of fonts, it differs only in a few pixels from '1'. Use "L" instead :)
assert l is not None and len(l) == 1And as was mentioned the "assert l" part is defending against l being None. I suppose I could be more explicit by saying "assert l is not None and len(l) == 1".
I'm not sure why you'd defend against None anyway. Why defend against None, but not against 3.1459 or 4j or ''?
He's defending against None because calling __len__ on None results in an exception.
3.14159 also results in an exception but it's far more likely that the object passed was None than that it was a completely different type than the one expected.
FYI, None is a completely different type than the one expected.
In a boolean context, the list is true if the length is non-zero. This example and the one the article is about is for the case where you know the list to have exactly one element. Not zero and not more than one.
The `if L and` part safeguards against that.
Use xs. If you have multiple lists use ys etc.
This has multiple benefits over L in terms of readability anad understandability as a single-item variable names can be made to match the list naming scheme:
for x in xs:
for y in ys:
do_some_fancy_calculation(x,y) If one element
do this one element thing
else
do this more than one element thingFirst: You are often working with someone-else's library which for various reasons you cannot change.
Second: It is not uncommon to use a standard method that may well be able to return multiple items, but in your use case it should only return one. Case in point: a database call that returns the result of a query.
Another example is if you know there is a single item in a data structure, and use list comprehension to extract it. You'll end up with a list of one item.
These situations happen, and for good reasons.
As an example: Say you have a GUI widget that can contain many entries, and you call `widget.get_entries()` which returns a list with all the entries. But if you know there must be only one entry, you can do `(entry,) = widget.get_entries()`.
In reality, you shouldn't be mucking with the entries of a widget at all; you should tell the widget what to do and it should adjust its entries as necessary). Demeter's law and all.
class SpartanList(list):
def foot(self):
return self[0]
assert l.foot()
return l.foot()Did you know that "tuple unpacking" actually works on arbitrary iterables (in fact, the relevant function in Python's C code is called ``unpack_iterable``)?
>>> def f(): yield 3
...
>>> b, = f()
>>> b
3
you should be careful with such declarations.> Yes. Why is that surprising?
because you called it "tuple unpacking". Furthermore, because in functional languages where this feature comes from you actually unpack tuples, and pattern-matching of other structures is not called "tuple unpacking".
> I've always thought that the tuple referred to by 'tuple unpacking' was the target containing the variables being assigned
That's what the "tuple" is unpacked into, it makes no sense that the name of the pattern would come from there. At best and stretching it, you'd have "into-a-tuple unpacking". Furthermore, you're not actually unpacking into a tuple (even less so in Python 3, `(1, * a)` isn't a valid expression... anywhere that I know of), you're unpacking into free variables. That's the point of unpacking.
> because you frequently unpack things other than tuples.
The only other thing frequently unpacked is a list, and Python's tuples are immutable lists, it does not stretch the imagination that a read-only feature of tuples would work on their mutable cousin as well. As for other structures, in 6 years of Python I'd say I've seen unpacking of arbitrary iterators thrice at best. You do not frequently unpack dicts or iterators.
edit: fucking hell, yc's comment format sucks donkey balls.
I imagine a great many people will be "interested" in your response, then.
And really, you choose to be uncivil over a name? Arguing about names is one thing; insulting your opponent because he uses a different name than you do is just childish.
Anyway, since I can't seem to resist trollbait:
> because you called it "tuple unpacking".
It surprises you that historical terminology persists even a decade after limitations have been removed?
> Furthermore, because in functional languages where this feature comes from
Algol 60 had this feature under the name of "multiple assignment" I'll bet long before LISP had "destructuring-bind". This feature did not originate in functional languages.
> That's what the "tuple" is unpacked into, it makes no sense that the name of the pattern would come from there.
Prior to Python 1.5, the expression being unpacked had to be a tuple; now it does not. The name is derived from the original functionality. If you'd spent your effort referring to the Python Language Reference instead of telling the world how you can't fathom someone would use a different name than you would for this sort of assignment, you'd know this :)
> The only other thing frequently unpacked is a list, and Python's tuples are immutable lists
Python's tuples are not immutable lists. See http://mail.python.org/pipermail/python-dev/2003-March/03396... for Guido's own words saying it.
> edit: fucking hell, yc's comment format sucks donkey balls.
You could have saved yourself a lot of trouble and others (like myself) a lot of annoyance by simply not commenting.
The only other thing frequently unpacked is a list
Which isn't a tuple. You're the one trying to be pedantic, but you want to ding me for making this distinction?
it does not stretch the imagination that a read-only feature of tuples would work on their mutable cousin as well.
I agree, it doesn't. I don't see how it stretches the imagination that it works on general iterables, either.
A couple of things:
1) The key word is obvious. There are oftentimes less obvious ways to do things that may be better for whatever reason.
2) It's not really reasonable to expect that there can only be one way to do everything.
What it really means is that (for instance) Python only allows one way to denote where a code block begins and ends (via indentation) while Ruby allows you to use curly brackets and begin/end. Nor does it have an unless statement that is equivalent to "if not"
[thing] = stuff
works too. I think I like that even better.
There doesn't appear to be any difference in speed between the two. Assigning the result to one variable (totally different to this) is about 6% faster, so this syntax is what I'll use, thank you!
On your exact machine with your exact version of Python, today.
Please don't let microbenchmarks dictate what code you write. If you need a 6% performance gain in a microbenchmark, you've chosen the wrong language to use. Python is about readability and maintainability, not syntax hacks to make some benchmark slightly faster.
thing, = [stuff]
though some people think that it's less legible since it has a completely different meaning without the comma.
EDIT: you can even play with the whitespace and pretend that ",=" is a list->scalar conversion operator:
x ,= [1]
also works and is more readable IMHO.
so if I was wrong in my original assumption that stuff has exactly one element, Python will shout at me before this will manifest itself as a hard-to-find bug someplace else in the program.
and then later on,
This method works even when stuff is a set or any other kind of collection. stuff[0] wouldn’t work on a set because set doesn’t support access by index number.
The second argument is basically in favor of duck-typing which is the pythonic way of writing code, ie, not caring about the actual object but only if it responds to the given message.
But what about the first argument? Couldn't you make the case that the pythonic way of handling it is to only care about if the object responds to __getitem__(0)?
Which idiom you use would depend entirely on context, wouldn't it?
`.__iter__()` is a more general and common interface than `.__getitem__(0)`, so it's preferable to assume that the object implements the former rather than the latter.
But sometimes you do want to assume stuff about your object. In this case we want to assume that the list has exactly one item, and we want our code to break immediately if it isn't true.
Mmm no? Here, what he cares about is that it's a single-element collection, he doesn't just want the collection's first item. If he did, foo[0] would do a better job.
So the object responding to __getitem__(0) is not a sufficient condition.
In fact, it's entirely wrong as sets do _not_ implement __getitem__. Worse, __getitem__(0) does a very different thing than unpacking on a dict. Unpacking is more ducky than __getitem__ for this case, because it expresses the following: the right-hand is a single-element iterable. Any iterable will work as long as it only has a single item, it doesn't have to be a sequence, a mapping, or even a collection (generators or callable_iterators will work just as well)
They are both decent ideas, not hard fast rules. Let your context be your guide.
if 1==len(lst):
return lst[0]
else:
return reduce(fn, lst)
I don't really want to have to do: if 1==len(lst):
(single,)=lst
return single
else:
return reduce(fn, lst)Furthermore, OP's assertion-unpacking will work not just on lists, but also on dicts (will return the only key, not the only value, and the key doesn't have to be ``0``), on sets (which don't implement __getitem__ at all), on arbitrary collections and even on arbitrary iterables (including callable_iterator and generators)
Also, please don't put the constant on the left-hand of a comparison in Python, it's useless and ugly.
Also, I don't worry about order on simple equality expressions unless someone asks me to do it a certain way. That's a religious argument and a complete waste of time for me.
return reduce(fn, lst)
does the same thing.I've never found a case where I knew there would only be one element in a list, but I do use tuple unpacking all of the time.
Pretty cool I guess; didn't know it wasn't a common idiom. Good to know. Thanks for sharing. :)
"""The result is a tuple even if it contains exactly one item."""
I run into this constantly using struct.unpack(), and I find the author's idiom to be ideal way to handle this. I will use this from now on.
thing = some_dict[my_object.get_foobar_handler()].get_only_element()
self-documents the purpose of the operation and its assumption much more legibly than the typo-esque (thing,) = ...
in my opinion.;)
1) In python, arrays and lists are two very different things.
2) Neither the array module nor the array class have a single function.
> a = [0]
> b, = a
> b == 0
=> true Python 2.6.5...
Type "help", "copyright", "credits" or "license" for more information.
>>> (a, (b, c, (d, e))) = (1, (2, 3, (4, 5)))
>>> a, b, c, d, e
(1, 2, 3, 4, 5)
>>> (a, (b, c, (d, e))) = (1, (2, 3, (4, 5, 6)))
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
ValueError: too many values to unpack
It isn't quite as flexible as functional languages and it's not as idiomatic as it is in functional languages, but it's not a hack or quirky edge-case either.When you want to get the (n+1)th item from a list, do:
item = stuff[n]
To get the first item, do:
item = stuff[0]
Unless the list has one element, then do:
(item, ) = stuff
At least you'll never be accused of consistency.
It's not really inconsistent, because you aren't performing the same operation in the two cases.
I see this fallacious reasoning all the time in design critiques: this is inconsistent with that. Well, yes, but so what? Why is consistency desirable in this instance?