Python wats
github.com
github.com
I can tell you what I consider WTF worthy - things that I and others have undoubtedly been bitten by because they are ubiquitous. Those are mutable defaults, lack of loop scoping, and "global"/"nonlocal" scoping. Oh, and add the crazy "ascii" I/O encoding default on 2.7 (and the inability to cleanly reset it) to that list.
if not x:
do_something()
Where do_something will occur if x is None, zero, 0.0, a blank string or even midnight in some versions of python.This has bitten me multiple times. It's a violation of python's explicit over implicit. It's an example of weak typing in an otherwise strongly typed language. There's really no good reason for it.
IMO, if x isn't a boolean, that line should just throw an exception. If you want to check for x being 0.0 or "", you should do x == "" or x == 0.0.
To me, this is the absolutely most reasonable way to check for truthiness or falsiness. If you want to check for equality or identity, use == or is.
E.g. one can mean "value not provided" while the other means "there aren't any of that thing" and the behavior you want will often be different depending upon which one it is.
It's this type of subtle difference that often causes obscure bugs.
The idiomatic way to check for an empty list is "if not mylist", but I still enjoy the fact that I can check for "if this is anything other than a thing I care about, do this" with one check.
Not to mention that doing anything else would violate duck typing. You don't care if the thing passed is an empty list, you care if it's falsy. If you had a "reverse()" function that checked for an empty list and raised an exception, you'd get in trouble when someone passed an empty tuple. What you care for is whether the object is iterable or not, not whether it's a list.
They will both be true if x is None. What you probably want is actually:
if x is None:
and if len(x) == 0:This is one of the warts which wasn't changed in Python 3.
Because it's no more explicit than "not x" and because it's 5 characters longer.
There's always a trade off to be made between clarity and verbosity.
Judging by the number of bugs I've squashed caused by the developer either not realizing or simply forgetting that ("", 0, 0.0, midnight, False, [], {}, None) all resolve to false, I've come to believe this is one of those times where the clarity should probably take precedence.
You can go to the other extreme by making everything explicitly statically typed which will make your code much more verbose (e.g. see Java). I believe that trade off isn't worth it either, despite the small increase in type safety.
Java is not an example of a good static typing implementation and is not exemplar of verbosity in modern statically typed languages.
[1] https://docs.python.org/2/library/stdtypes.html#truth-value-...
[2] https://docs.python.org/2/library/datetime.html#time-objects
def myfunc(mylist):
mylist += [1,2,3]
is NOT the same as: def myfunc(mylist):
mylist = mylist + [1,2,3]
Try this on each of those definitions: a = [4,5,6]
myfunc(a)
print a
The first one mutates 'a', the second only reassigns (and discards) a temporary `mylist`.Put another way, += is not just syntactic sugar. With lists, it actually mutates the original list.
I'm not sure but I think that list.__iadd__ is even more tricky: When the new object still fits into the allocated array, it will modify it inplace, otherwise it returns a newly allocated list and doesn't modify it inplace. So you cannot rely on __iadd__ modifying a list (or anything else) inplace.
False; __iadd__ is always in-place for lists.
To me the real wtf is the second, where python creates a variable with the same name and discards it.
void increment(int x) {
x += 1;
}
void printit() {
int y = 0;
increment(y);
printf("%d\n", y);
}
C doesn't have lists so I had to use an int here, but the Python example does the equivalent of that code printing 1, not 0. (But only for lists. That same Python code with ints does what you expect.) >>> k1 = "foo"
>>> k2 = hash(k1)
>>> hash(k1) == hash(k2)
True
>>> d = {k1: "v1", k2: "v2"}
>>> d[k1]
'v1'
>>> d[k2]
'v2'
Now that I've been pedantic, I hope I get this right: in Python, it is equality (==) that defines same-ness for dicts and sets. But if x == y, then you have to ensure that hash(x) == hash(y). Python uses this characteristic to make the initial check for dict and set membership integer comparisons on the hash, but when two items have the same hash, Python goes on to distinguish them on the basis of equality. class A:
def __hash__(self):
return -1
print(hash(A()))
Prints -2.For the same reason, if you ever need an easy hash collision, use the hashes of -1 and -2.
This was my first thought too, so I checked but
>>> hash(0)
0
>>> hash(0.0)
0
>>> hash('')
0
>>> {0: 4}['']
KeyError: ''
Which makes sense, because there is checking for hash collisions, and when they are resolved 0 == 0.0 while 0 != ''
def make_key(obj): return (obj.__class__.__module__, obj.__class__.__name__, hash(obj))
make_key(x) == make_key(y)
This has nothing to do with tuples, I think, as one can constructs many other examples just like this one.
https://docs.python.org/3.4/reference/simple_stmts.html#gram...
>>> x = 0*1e400
>>> set({x, x, float(x), float(x), 0*1e400, 0*1e400})
set([nan, nan, nan])
>>> set({x, float(x), 0*1e400, 0*1e400})
set([nan, nan, nan])
>>> set({x, float(x), 0*1e400})
set([nan, nan])
EDIT: it gets worse: >>> x = 0*1e400
>>> y = 0*1e400
>>> z = 0*1e400
>>> set({x, x, x})
set([nan])
>>> set({x, y, z})
set([nan, nan, nan])The only actual surprise on this list is, I agree, that += operator. Especially because tup[0] = tup[0] + [1] fails as expected.
[0] https://docs.python.org/3/reference/datamodel.html?highlight...
So while it looks weird there is a rational backwards-compatibility reason for it.
I think the more important reason is that Python doesn't really have an emphasis on the boolean type in general. The `if` statement works for every type, in contrast to other languages allowing only the boolean type in `if`. It is very common to use `if string_or_number_or_anything: blah blah` in Python. While I won't make a stance whether this is a good thing or not, it is still awkward that Python allows implicit conversion between `bool` and `int`.
That was a while ago, however. This behavior ought to be phased out in favor of treating it as entirely its own type.
http://stackoverflow.com/a/6865824/1763356
Alex Martelli's answer on the same page goes into more depth.
I get the worries, but this is Python - you should know what domain you're working on anyway, because you have to. It's a different philosophy to statically typed languages, and since having True and False as aliases is convenient I'm personally glad for it:
# Truth as samples
sum(x > 10 for x in xs)
(numpy.random.randint(0, 10, 1000) == 7).mean()
# Several properties are made obvious
assert True > False
my_flag ^= TrueThe only real benefit to this appears to be some shortcuts that shave a few characters off and do so, IMO, at the expense of readability.
E.g. this reads easier for me:
len([x for x in xs if x > 10])
Than this does: sum(x > 10 for x in xs)
Martelli may call this stuff contortions but to me it feels more natural.At the expense of allocating a new list, maybe. But the `sum` convention is faster, space-efficient and cleaner. It's also pretty obvious once you've seen it.
I'm sure some people will look at this and think that it's summing all of the numbers over 10 in xs:
sum(x > 10 for x in xs)
Whereas with this, which is explicit, it's substantially less likely they'll think that: sum(1 for x in xs if x > 10)
Again, I don't see the point of maintaining the weakened type system for a few minor 'clever' barely-short cuts like this. Look at what that kind of thinking did to perl.If python started throwing exceptions on all the code that treated True and False as integers, I'm pretty sure all the fixes done to accommodate that would probably make the code cleaner and easier to understand.
Perhaps it's just a bias of familiarity, but that argument seems contrived to me.
I'm worried the rest of my arguments will amount to "I like what I know", so perhaps we should lay this to rest as a difference in tastes. It does remind me a little of concatenation with `+`, which is oft-hated... and yet I've not seen a single error caused by it.
sum(int(x > 10) for x in xs)
Just require the conversion to be explicit.In fact, I think the ones he picks reflect tremendously on his skill as a programmer.
Take for instance his apparent insistence that int(2 * '3') should produce 6. Cool. Magic "special" behaviour for the multiply operator if the string happens to contain a number.
The same goes for his desire for the string "False" to be falsey. Wonderful magic special behaviour dependent on the (possibly unknown) content of a string means I've got to go round checking my strings for magic values.
In python 2.x you could redefine True and False, and this is fun:
>>> object() < object()
True
>>> object() < object()
False
>>> object() < object()
True
But python3 dropped a lot of wat-ness.This one is fun too: https://bugs.python.org/issue13936
The only exception I think is the bool coercion. Even though I can predict the behavior, I find it fairly disturbing. I believe that should have been removed when transitioning to Python 3. It is very unfortunate, considering Python is generally reluctant to implicit type conversion.
all([[[]]]) == True
makes perfect sense, because the all all does is check the truthiness of the argument's elements. An empty list is false, non-empty is true, no matter the contents. >>> all([[[]]]) == True
True
>>> all([[]]) == True
False
>>> all([]) == True
True
is pretty counter-intuitive at first glance.If we denote
x = []
then bool(x) == False # empty list is falsy
bool([x]) == True # nonempty list is truthy
bool([[x]]) == True # nonempty again
and finally, [bool(elem) for elem in [[x]]] == [True] # all True!
which is the thing `all' is interested in. It is more like a newbie mistake or careless document reading if the user thinks `all' runs through the lists and all of the nesting too.Cast false to a string, then check the Boolean value of that string. It's not an empty string so of course it would be true.
What would a C or C++ programmer expect? Because my experience in those languages says "casting" is a meaningless way to understand what's going on.
Perl doesn't have bareword true/falue values. Ruby doesn't use "casting", I think. That is, I think the idiomatic way to express this in Ruby is:
>> !!true.to_s
=> true
>> !!false.to_s
=> true >>> import numpy as np
>>> 'x'*np.float64(3.5)
'xxx'
(on recent NumPy releases this raises a deprecation warning)Clearly the correct answer is 'xxx>' ;)
>>> type(np.uint64(1) + 1) numpy.float64
>>> False == False in [False]
True False == False and False in [False]
https://twitter.com/marcusaureliusf/status/55794887300903731...True
x (operator1) y (operator2) z
is defined by the language as equivalent to (x (operator1) y) and (y (operator2) z)
So 1 in [1] in [[1]]
is equivalent to (1 in [1]) and ([1] in [[1]])
For any comparison operator >>> False == (False in [False, ])
FalseThis is much faster than checking if the string has some semantic meaning first, which would have to be localized.
[0]: http://www.javapuzzlers.com/
[1]: https://www.youtube.com/watch?v=wbp-3BJWsU8
* Reassigning interned integers via reflection
* private fields work on a class level, not an instance level, i.e. Instances of a class can read private fields of other instances of the same class.
* Arguably package-private fields as the docs like to avoid mentioning they exist.
* ==, particularly in regards to boxed types.
* List<String> x = new ArrayList<String>() {{ add("hello"); add("world"); }};
* The behaviour of .equals with the object you create above
* Type Erasure
System.out.println(0.0 == -0.0); // true
System.out.println(java.util.Arrays.equals(new double[] {0.0}, new double[] {-0.0})); // false
(It's documented in the contract for Arrays.equals, but still kind of ridiculous.) # let x = ref 0;;
val x : int ref = {contents = 0}
# let y = ref 0;;
val y : int ref = {contents = 0}
# if false then
x := 1;
y := 1;
;;
- : unit = ()
# (!x, !y);;
- : int * int = (0, 1)
# y := 0;;
- : unit = ()
# if false then
let z = 1 in
x := z;
y := z;
;;
- : unit = ()
# (!x, !y);;
- : int * int = (0, 0)Oh, and `null`, in any language ;-)
The one wat was adding a float to a huge integer. Granted, I assumed something fishy might happen with such large numbers, so I steer away from doing those operations.
Since the 1 is shifted 53 bits to the left the number is right at the start of the number range where only even numbers can be represented.
You get 900719925470993 as the decimal number which you cast to float : 900719925470992.0
Then you add 1.0 and get 900719925470992.0 because of floating point imprecision and rounding. The next floating point number would be 900719925470994.0.
92 is less than 93 and this gets you this seemingly weird x+1.0 < x = true.
That said, I'm not aware of a language that does better, and I'm aware of many that do much worse.
9999999999999999999002 + 49.0
rounds to 9.9999999999999999990e+21, whereas the lesser value 9999999999999999999000 + 50.0
rounds to 9.9999999999999999991e+21.Plus, I'm not a fan of computing with decimal arithmetic, since it's less stable. For instance,
min(a, b) <= (a + b) / 2 <= max(a, b)
doesn't always hold for decimal floats, whereas it does for (non-overflowing) binary floats. Decimals are generally more prone to this kind of inaccuracy, since they lose more bits when the exponent changes.(Consider a = 1.00000000000000000001, b = 1.00000000000000000003.)
Interval arithmetic support is cool, but not useful without effort - bounds like to grow. Plus, Python has bindings for them anyway ;).
If all would work any other way, it would be seriously broken. (Given the rules for how to convert list to bool.)
A more direct solution might be to have a means of storing off every assertion failure into a (clearable) list. The test harness could then check that list at the end of each test. If a particular test should succeed after triggering an assertion failure, it can check the list to make sure it triggered exactly the expected failures and then clear it.
Right, I guess I'm not suggesting changing the language today to do this. Rather, I'm suggesting that it would be better if it were already done that way from the beginning. If it were designed that way from the start then there would be no question about the intention because it would just be part of the definition of Exception
Heh, doesn't that argument apply currently? I think I was trying to ask, "when someone fails to think deeply enough about it, what is likely to give the 'correct' results"?
Edited to add - thanks for responding, by the way. I'm confused by all the down-votes.
I still don't quite get what you meant.
If a test case triggers an assertion violation down in some method, there is a bug. That should break the test, so that I'm told about the bug, and can investigate and fix it. If there happens to be a `try...except Exception` anywhere in the stack above that method, the test never learns that an assertion fired and might even pass. This makes every test less useful than it could be.
(Then you investigate and hopefully remove the catch-all.)
This isn't terribly relevant to narrow unit tests, where I should be able to know what I expect to have happened and presumably if an assertion pops it won't have happened, but it makes larger scale fuzzing substantially less useful.
def f():
x = 0
def g():
x += 1
g()
f()Any assignment operator (compound or not) follows the same namespace rules. Scope lookups are consistent within a scope, regardless of order. Any "namespaced" lookup (`x[y]` and `foo.bar`) will act inside that namespace.
Non-namespaced identifiers will always assign into the nearest enclosing scope (a scope is always a named function) by default, since clobbering outer scopes - particularly globals - is dangerous. You can opt in to clobbering with `global` and `nonlocal`. If you don't ever assign to a name in the current scope, lookups will be lexical (since a local lookup cannot succeed).
---
Hopefully most of these value judgements should be relatively obvious.
>>> 3.1 is 3.1
True
>>> a = 3.1
>>> a is 3.1
False >>> a = 'hn'
>>> a is 'hn'
True
>>> a = ''.join(['h', 'n'])
>>> a is 'hn'
False
>>> a
'hn'
>>> a = 'h' + 'n'
>>> a is 'hn'
True
Edit: found another interesting case >>> a = 1
>>> b = 1
>>> a is b
True
>>> a = 500
>>> b = 500
>>> a is b
False a = 500
b = 500
a is b
will give `True`. But the REPL forces each line to compile separately and CPython won't cache across them. However, there's a global cache for integers less than 256 so in that case they're `is`-equivalent even in the REPL.For inequivalent values or `True`, `False` and `None`, the result is defined. Identity is also preserved across assignment. For everything else, the answer is "whatever the runtime feels like". PyPy, for instance, will always say True for integers and floats.
It's not just hard to predict - it's inherently unpredictable. The runtime gets free reign as long as you can't prove that it's lying.
awesome idea - thank you
foo = 1.0 + 2 # Foo is now 3.0
foo = 1,0 + 2 # Foo is now a tuple: (1,2)
foo = 3 # Foo is now 3
foo = 3, # Foo is now a tuple: (3,)
source - https://wiki.theory.org/YourLanguageSucks#Python_sucks_becau...