The Pythonic Emptiness
blog.codingconfessions.com
blog.codingconfessions.com
> Similarly, when you use len() to check a sequence for emptiness, you are reaching out for a more powerful tool than necessary. As a result it may make the reader question the intent behind it, e.g. if there is a possibility of handling objects other than a built-in sequence type here?
Given that checking for truthiness is less strict than a length test, by the same token, whenever you use it, you're reaching for an even more powerful tool than necessary. And, if anything, seeing `not items` is what makes me question the intent - did the author mean to also check for None etc here, or are they just assuming that it's never going to be that? And sure, well-written code will have other checks and asserts about it - but when I'm reading your code, I don't know if it's missing an assert like that because you intended it, or because you couldn't be bothered to write it.
OTOH len() is very explicit about what is actually checked, and will fail fast and loudly if the argument is not actually a sequence.
Also note that it's not, strictly speaking, an either-or - you can use `not len(x)` instead of `len(x) == 0` if you want a distinctive pattern specifically for empty collection checks.
An empty array or string are not false. Zero is not false.
In Lisp dialects with nil, we don't have endless discussions about how to test for false, or for empty, as the case may be.
Okay? That still doesn't justify using it as replacement for #f (which is also a unique value). Or as #t for that matter, there is also only one #t in the entire system.
> An empty array or string are not false. Zero is not false.
And an empty list is also not false, only #f is false. So simple and uniform! Wait, there is no #f, but an empty list is false? Why? Why is empty list so special? Or, alternatively, why are you using the false value as an empty list? The idea to represent a list as
[1 | [2 | [3 | False]]]
is quite strange. Just spare an atom, it's only what, eight bytes nowadays?> how to test for false
In sane programming languages, the booleans are self-testing, so to speak, and they are the only values with such property, so:
if not val then ...
(if (not val) ...
tests for false.> or for empty
if val_specific_is_empty(val) then ...
(if (null? val) ...
But of course, it only matters for languages that have other options for building composite data types than CONS-ing two items together.(Scheme isn't that language, because everything that is not #f is true.)
List processing is cumbersome in Scheme because the empty list is't false, and because car and cdr cannot be applied to (). (Which, stupidly, isn't even self-evaluating; you have to use '()).
The empty list is special because once upon a time lisp emphasized list processing more. Complex list processing is still important in metaprogramming, and representations of flexible dynamic data sets.
You're not working with lists, it utterly doesn't matter that nil also has a role as the empty list. Just like you don't care that your physician also plays saxophone in a jazz band; he's not doing it during your appointment.
Also in Pandas if you try to check the truth value of a dataframe to see if it’s empty, it will fail. It will say “the truth value of a dataframe is ambiguous”.
df.empty is less ambiguous but you have to remember it specifically for dataframes.
But len(df) > 0 almost always works for any type of collection.
https://github.com/python/cpython/blob/v2.7.1/Lib/cgi.py#L60...
Python's std lib `cgi.FieldStorage` object was falsy if it did not define any headers, even if it contained file data.
Thus my conditional trying to check whether a file was being uploaded "if request.field_storage" was going through the False branch when files were being uploaded but only in certain header-free scenarios not covered by automated and manual testing. This resulted in us dropping user data on the floor and losing uploaded files for a very large number of users before we realized and shut it off
The other sibling post contains another example where people may be confused, and Google pulls up others. But this is the concrete case that caused us to send out a hundred thousand apology emails to affected customers after losing their files
If we're talking about readability, they're far clearer than either of the options in the article and require no pre-knowledge of truthiness rules.
if not mylist: # 1.061
if len(mylist) == 0: # 1.924However, this is not an example of that. "if list:" has been the idiomatic Python since inception, and "if len(list):" has been an unnecessary complication for the same period of time. Python's "preferably one way to do it" has never been about "there is literally only one syntactically valid way to do it", for fairly obvious reasons if you think about it.
Perhaps it's just because I'm not Dutch.
I would argue the reverse. It was a bad idea to begin with and the start of something worse.
> Plenty of languages have changed their guiding principles over time.
What would you say is the guiding principal of Python now then?
It has unambiguously been idiomatic Python since the beginning. My opinion is that it is bad, but that's much more an opinion than the fact it has been idiomatic. And to Python's credit, part of the reason why I am so sure it's bad is precisely the experience I gained in Python using it. At the time Python was implementing the principle, I don't think the general experience of the programming language community was strong enough to know that it was a bad idea.
Python's guiding principle right now seems to be the same guiding principle as almost every other language, "let's solve as many problems as possible by adding features to the language". If a year goes by without at least one major new feature, the language must be "dead" or "failing". As the years wear on and so many languages have piled up so many features, I wonder when people will finally look around and realize that all these features, for all their superficial appeal in the small scale, are not generally helping them write better programs, or write programs they couldn't have written before, and often harm their programs on larger scales. There are exceptions. I have a hard time imagining any modern language without some concept of closures, for instance. I could name a few more; some sort of easy polymorphism (there's a few ways to get there but you need to take at least one of them... but preferably not all of them...), some sort of concurrency solution in this era, solutions for memory safety (again, multiple solutions, but you need at least one of them). But so many of these features are, in my opinion, not a net positive, their benefits far smaller than meets the eye and the costs so much greater.
Just use a normal check, like everyone is expecting to see.
“Oh but what if it’s not a sequence”, well then you have bigger problems. Why are you emptiness testing something that may-or-may-not-be-a-sequence? Maybe solve that problem first.
Indeed. Relying on truthiness has always felt very un-Pythonic to me, not least because it contradicts several principles in the Zen of Python:
• Explicit is better than implicit.
• Special cases aren’t special enough to break the rules.
• There should be one — and preferably only one — obvious way to do it.
The original Hungarian notation was a less readable implementation of the same idea. The Hungarian notation most people hated replaced functional types like count and index with data types like unsigned long. And used them everywhere.
It’s the duck typing nature of, classic Python, that leads the community to recommend the broader `if items:`, which allows for numbers and such.
You just have to not write any bug in your code. Also, use type checking everywhere. And rewrite your mind, too.
The benefits are worth it .. ! Oh well;
I can't think of anything more Pythonic than that!
The programming language that has the least measured bug in practice is Clojure because it is duck typed and because it doesn't use OOP. Both static typing and OOP have a significant measurable negative effect on code correctness.
My invented numbers say there's no meaningful difference in line counts between statically typed and dynamically typed languages.
So how will we prove any of us wrong?