WTFPython: Exploring and understanding Python through surprising snippets
github.com
github.com
class A:
a = (1,) # tuples are immutable and don't have "+="
b = [1,] # lists have "+="
obj = A()
obj.a += (2,) # this creates obj.a
obj.b += (2,) # this modifies A.b
print(A.a, A.b)
The snippet prints "(1,) [1,2]".I think this behavior is actually "default args" but reversed. In case of default args it is the mutability that causes the nasty surprise (ooops, all calls share the same object, which is mutable). In case of the "+=" it is the _immutability_ that causes the surprise - I would not expect "obj.a += ..." to create an instance attribute that shadows the class one.
I think the practical conclusion here is "don't call += on objects that don't support it - because it may work and this is not what you want!)".
I'm curious if you see it the same way?
Side note: the object can be immutable, but then the (variable or attribute) left side should be replaced with a new immutable object.
But what does that have to do with assignment?
>>> a = 3
>>> a += 4
>>> print(a)
7
3 and 7 are immutable, yet a takes on those objects.Obviously, this is inspired by the C += operator, and behaves accordingly, when it's not being weird.
The difference is in how the Class attributes are affected i.e. the values of "A.a" and "A.b".
# obj.a += (2,) has no impact on A.a since tuples
# are immutable and copies are made.
# This is unexpected from a C/C++ intuition.
>>> print(A.a)
(1,)
# obj.b += [2,] has an indirect impact on A.b
# since lists are mutable and updates are made in place.
# This is unexpected for beginners in Python since it
# deviates from how int, str and tuples behave as "+="
# can be used for int, str, tuples and lists.
>>> print(A.b)
[1, 2]
Edit: formattingNo, they don't.
obj.b does not exist, and no assignment to it takes place, and there is no “indirect effect”.
A.b exists, and because class attributes can be accessed as if they were members on instances, and because A.b has an implementation of in-place addition, that implementation is called on A.b instead of an assignment to obj.b. That is the only thing that happens as a result of the statement that looks like it might be an assignment to obj.b. The effect on A.b is the only action, not an indirect effect of an assignment.
I thought of this as though the list A.b was marked static - then any modification of b through any instance of A modifies the shared static member
>>> class A:
... b = [1,]
...
>>> obj = A()
>>> obj.b.append(2)
>>> A.b
[1, 2]
>>> foo = A()
>>> foo.b
[1, 2]
>>>It seems to, but, actually, it doesn't.
> but wtf'ingly also alters the value of the class attribute "A.b" which completely alters the behavior of A.
Actually, all it does is call the __iadd__ method on the list that it is the value of A.b (which modifies that list in-place.)
The value of A.b can, incidentally be accessed via obj.b because obj.b does not exist and the lookup path for instance attributes includes the class.
There are two things going on here there are potential sources of confusion, because they behavior differently in different circumstances: object member access (which can get members from the object or, if they don't exist on the object, from the class) and the += operator, which calls the __iadd__ method on the left value or acts as add and then assign to the name on the left, if that name doesn't reference a value that supports __iadd__.
That is not correct. += will always perform the assignment, hence the classic oddball concatenation:
>>> v = ([],)
>>> v
([],)
>>> v[0] += [1]
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
TypeError: 'tuple' object does not support item assignment
>>> v
([1],)
So `obj.b += [2]` desugars to something along the lines of: l = obj.b
l.extend(2)
obj.b = l
Hence if you call vars() on the instance or access its __dict__, you will see that the instance does have an intrinsic "b" attribute: >>> class A: b = [1]
...
>>> a = A()
>>> a.b += (2,)
>>> vars(a), a.__dict__
({'b': [1, 2]}, {'b': [1, 2]})
And if you re-set the attribute on the class, it won't affect the instance, which it would if the instance delegated to the class: >>> A.b = [5]
>>> a.b
[1, 2]It necessarily alters the value of A.b, since `+=` is specifically overridden to update the list in-place. With lists, `a += b` is specifically not an alias for `a = a + b`, instead its behaviour is closer to
a.extend(b)
a = a
The issue here is the first part, I absolutely hate this override, and sadly the core team has not learned a thing there as `|=` was overridden the exact same way on dicts when `dict | dict` was added to the language.Ya, this is because the initialized object is stored as the default parameter, and is not initialized every time the function runs.
def functy(dicty={}):
print('dicty=', dicty)
return dicty
output = functy()
# dicty = {}
output.update({'key': 'value'})
functy()
# dicty = {'key': 'value'}Note to readers unfamiliar with Python: class attributes is a special thing, different than instance attributes.
Definitely one of the trickier things with Python for sure.
> # tuples are immutable and don't have "+="
What do you mean? Because it looks to me tuples have "+=":
a = (1,)
a += (3,)
# a = (1, 3)
I get that tuple is immutable so you're actually creating a new tuple and shove it back into variable a, but it does not conflict with "having += operation".Thanks!
Python is more of a lisp. The assignment semantics are different to C. Actually, you're letting the name "a" point to the newly created object.
In C, a variable is a memory position you can fill and point to.
In Python, a variable is a name that points to a memory position.
Similarity to C is no accident here, since the lexical scoping concepts in Lisps like Scheme and CL, as well as in C, both trace back to Algol.
Classic Lisp global/dynamic variables may store a global value a "value cell" which is closely tied to the symbol itself (perhaps stored in it).
And in that sense, C assignment and Lisp assignment are different beasts. Mainly in that - again in the semantics of the high-level description - "memory reservation" in C happens when we declare the left-hand side of an assignment and in Lisp it happens when we construct the right-hand side.
Though if you can show me a mainstream Lisp where the assignment works like setting a memory cell and not like a bind, I'm happy to be proven wrong.
The right hand side reserves memory only in the sense that some heap object is allocated (if that is the case), but that is secondary to the assignment; and that is the same as C also, as in:
f = fopen(...); // right hand allocates stream; pointer moves into variable.
> if you can show me a mainstream Lisp where the assignment works like setting a memory cell and not like a bind, I'm happy to be proven wrong.The mainstream Lisps clearly separate binding from assignment.
(let (x) ;; x refers to a freshly allocated cell (nil-initialized, in Common Lisp)
...
(setq x 42) ;; cell is clobbered, replacing nil with 42.
...)
Some of the terminologly used in the specifications is a bit confused. For instance, Common Lisp says say that:1. A variable is a "binding in the variable namespace" (Glossary)
2. A binding is an association between a name and a value (Glossary)
3. Yet, the description of SETQ uses language like: "First form1 is evaluated and the result is stored in the variable var1".
4. LET is described like this "let and let* create new variable bindings and execute a series of forms that use these bindings". No binding creation semantics is mentioned for SETQ.
Results can only be stored in storage places; nothing can be stored in an "association between a name and a value", unless that association is actually set up through a memory cell: the name refers to a location and the location holds a value. SETQ isn't described in terms of breaking an old binding and setting up a new one..
In the case of dynamic and global variables, there is a term "value cell" which the Glossary defines: like this: "The place which holds the value, if any, of the dynamic variable named by that symbol, and which is accessed by symbol-value. "
Let's turn our attention to Scheme. R7RS says in 3.1. Variables, syntactic keywords, and regions this:
"An identifier can name either a type of syntax or a location where a value can be stored."
and:
"An identifier that names a location is called a variable and is said to be bound to that location."
"The value stored in the location to which a variable is bound is called the variable’s value."
The next sentence is a kicker, and can be regarded as a criticism of the ANSI CL definition of binding:
"By abuse of terminology, the variable is sometimes said to name the value or to be bound to the value."
If you think that Lisp variables are bindings to values, then according to the Scheme maintainers, you've fallen victim to abuse of terminology. :)
But if l does support __iadd__, then l += r becomes l.__iadd__(r).
This behavior for list and tuple is familiar from regular local variables. Together with the fact that obj.x = ... always means "create or rebind an instance variable" never "rebind the A.x class variable", I think the example is less surprising.
>>> print(obj.a, obj.b)
(1, 2) [1, 2] class A:
a = (1,)
b = [1,]
def __init__(self):
self.a = (1,)
self.b = [1,]
A.a # this is the class attribute
(1,)
A.b # class attribute
[1]
obj = A()
obj.a # instance attribute
(1,)
obj.b
[1]
obj.a += (2,)
obj.b += [2,]
A.a
(1,) # still the same class attribute
A.b
[1]
obj.a
(1, 2) # instance attribute appended
obj.b
[1, 2]If I had had magic wand, I'd make operations on class attributes from an instance a syntax error, and only allow it from the class name, and perhaps with special syntax like in C++, e.g.:
A::a = (1,)
A::b = [1,]
*edit: operation on obj.a or obj.b... async function totalSize(fol) {
const files = await fol.getFiles();
let totalSize = 0;
await Promise.all(files.map(async file => {
totalSize += await file.getSize();
}));
// totalSize is now way too small
return totalSize;
}
You get an overly low totalSize. It's caused by 'a += b' expanding to 'a = a + b', and the double-mention of 'a' creating a concurrency issue. If '+=' were a single operation with the right-hand-side being calculated first, it wouldn't be an issue.That must have been a fun one to track down!
To fix this, simply await the result of getSize first, and then add the result to totalSize using +=
That is the sole source of the issue, and exactly what the comment you're replying to talks about.
The problem is that `a += b` desugars to `a = a + b`, so if `b` is an await you get
totalSize = totalSize + await file.getSize();
Since javascript evaluates left to right, it first evaluates `totalSize`, gets zero, then `file.getSize()`, then suspends waiting for the result... at which point the handler for the next file can run, doing the exact same thing, repeat for all files.I guess the problem in python is different though, it's about having some data structures that are immutable (like tuple) and others that are mutable (like list)
If anyone's willing to go through those examples with an interpreter on the side, you can check out https://www.wtfpython.xyz/
It's built with pyiodide, the only limitation is you may not get correct results for the examples that are version-dependent. The UX might not be great, especially on mobile (happy to hear ideas on how to improve it), but it does the job for now :)
can we use a different shade of grey for code blocks in dark mode?
I've found senior developers to have polarized opinions about the usefulness of the collection, so always good to see a review in favorable direction :)
Speaking of directions, how about the bidirectional use of `yield` and `yield from` ?
https://stackoverflow.com/questions/9708902/in-practice-what...
# Sending data to a generator (coroutine) using yield from - Part 1
I also learned that there is python code which causes a rather violent "Don't you ever get that near a code base I maintain"-reaction, which I much rather associated with perl.
Worse, no one uses it. I've yet to come across anyone that advocates it or remembers it.
>>> if m := re.search(r'(.*)s', 'oh!'):
... print(m[1])
...
>>> if m := re.search(r'(.*)s', 'awesome'):
... print(m[1])
...
aweFeature creep is programming language's worse enemy after a certain maturity level.
I absolutely love Go in this matter. They took forever to add Generics and generally sides with stability over features.
let some_func () = 5
let a = some_func () in if a < 5 then "<5" else ">=5" values = [
value
for line in buffer.readlines()
if (value := line.strip())
]
Previously, I would have needed to either duplicate effort like: values = [
line.strip()
for line in buffer.readlines()
if line.strip()
]
Or used a sub-generator: values = [
value
for value in (
line.strip() for line buffer.readlines()
)
if value
]
Or rewritten it altogether using a (slower) for loop calling append each time: values = []
for line in buffer.readlines():
line = line.strip()
if line:
values.append(line)
The assignment expression is perfect for this sort of use case, and is a clear win over the alternatives IMO.Edit: fixed initial example
values = [
value
for line in buffer.readlines()
for value in (line.strip(),)
if value
] nonblank = [value for line in file for value in [line.strip()] if value] open("file", "r").readlines.
map{|line| line.strip}.
filter{|line| line != ""}
or some smarter but less readable ways.I prefer the left-to-right transformations style to Python's list comprehension and inside-to-outside function composition. The reason is that it reminds me of how data flow into *nix pipelines. I spent decades working with them and I've been working with Ruby for the last half of that time. With Python in the last quarter of my career.
It's a matter of choices and preferences of the original designed of the language. Both ways work.
If function is None, the identity function is assumed, that is, all elements of iterable that are false are removed.
So it just removes false-y values.
Very handy I've used it a ton
filter(None, xs)
is equivalent to: filter(lambda x: x, xs)
That is, it will return an iterator over the truthy elements of the passed iterable.Certainly, I agree; I would usually use:
(x for x in xs if x)
Or, if I know more about the kind of falsy values xs actually needs removed, something more explicit like: (x for x in xs if x is not None)
Because Python’s multiplicity of falsy values can also be something of a footgun (particularly, when dealing with something a collection of Optionals where the substantive type has a falsy value like 0 or [] included.)Instead of:
filter(None, xs)
Which is terse but potentially opaque.Though it's additional syntax, I kind of wish genexp/list/set comprehensions could use something like “x from” as shorthand for “x for x in”, which would be particularly nice for filtering comprehensions.
stripped_lines = (line.strip() for line in buffer.readlines())
non_empty_stripped_lines = [line for line in stripped_lines if stripped_lines]Breaking down things in clear steps is underrated I think.
start = 0
while (end := my_str.find("x", start)) != -1:
print(my_str[start:end])
start = end + 1
vs start = 0
while True:
end = my_str.find("x", start)
if end == -1:
break
print(my_str[start:end])
start = end + 1
I'm still on the fence myself so I sympathise with your view, but the first version is certainly a bit tidier in this case.Many of my code were like this:
foo = one_or_none()
if foo:
do_stuff(foo)
Now I have the following: if foo := one_or_none():
do_stuff(foo)
This kind of code happens quite frequently, looks nicer with walrus operator to me. if one_or_none() as foo:
do_stuff(foo)I am not fully up to speed with 3.10, but quickly checked the docs and it doesn't appear to have been added in 3.10 either.
Let me know if I'm missing something.
for foo in one_or_none():
do_stuff(foo)
If you don't control one_or_none, but it returns an Optional, you can wrap it with something like: def optional_to_tuple(opt_val: Optional[T]) -> Tuple[]|Tuple[T]:
return (opt_val,) if opt_val is not None else ()In both cases, "foo" continues to exist after the "if" even though the second example makes it look like "foo" is scoped to the "if".
So to my eye, the following would look super weird (assume do_more_stuff can take None):
if foo := one_or_none():
do_stuff(foo)
do_more_stuff(foo)
whereas the following would look fine: foo = one_or_none()
if foo:
do_stuff(foo)
do_more_stuff(foo)I don't want to use Python for a software more than a few dozen lines or involving multiple developers. This language was designed to unintentionally shoot your foot so many times.
This is based on 3 different companies, of which all had decent developers who did follow good practices.
This is 100% true for medium to large codebases. Also because of the dynamic typing just looking at the codebase, it is very hard to understand the "shape" of variables and data. Of course this is improving now because of the type hints, but still it is comparatively hard.
And is it in, let's say Java, with several inheritance levels and very opaque types? It's one thing I struggle with
Sure you know a TypeA has method do_stuff() and returns a TypeB but what actually is happening, beats me. Then you chase down TypeA and find out the actual implementation is spread across TypeA0, TypeBaseA and you can't make sense of anything
That was about plain Java. Where it stops working is some "dependency injection" frameworks which jumped the shark and stopped being about dependency injection (cough Guice cough). A fracking argument to a method comes from god-knows-where because The Framework injects it based on a combination of its class and and annotation? Yes, now it is as bad as Python or maybe even worse :)
All that being said, the by far most common cause of bugs in Python is None in my experience, not any of the shenanigans the language allows you to do.
Mid-sized (5k < loc < 100k?) can absolutely survive a complete "hostile takeover" from new developers, with all the same potential (but not mandatory) issues and pitfalls that pretty much any such codebase invites, in particular in dynamically typed languages.
Also, being boring to write means you seek ways not to write so much, which is also good :)
The "worst" that I can remember was actually a Django quirk - ORM result sets (or QuerySets as they call them) execute their query at object evaluation time rather than at instantiation, so unless you cast your resultset to a list it won't actually talk to the DB just yet. Now attach a GUI debugger such as PyCharm/IntelliJ - it will internally evaluate every expression immediately (to populate its GUI) and cause the behavior to diverge (it will "fix" the code and behave as you'd expect when ran under the debugger).
I guess there could be certain contexts (system programming? low-level libraries? etc) where these are going to be an issue but when it comes to web/API or business logic development I haven't encountered these.
"Where is this called from?", "Where is x being set/modified/read?" - such questions are hard to answer with a large Python code base and I'm not even talking of code that abuses the dynamic nature of Python.
But yes, the moment someone tries to be too clever (metaclasses, dynamic *kwargs parsing...), all of that falls apart and you're back to reading the docs.
However, I do think that the pythonic way leads people to write terse clever code that's needlessly more complex.
At least now the easy_install vs pip contention is dead.
About 75% of the bugs I write seem to be about mutable state, and python has nothing but mutable state.
It doesn't even get scoping right.
Reading through the documentation in WTFPython I get some kind of affirmation of my bias against python. I hate it, but I also love it, but I hate it. It is just so "je ne sais quoi"--
The docs would benefit from some explanations on whys.
Ok that is infuriating. They dragged their feet through the dirt on how the switch statement is useless and pointless, but then they add some random bullshit like that that nobody ever asked for.
The existing "as" would have worked instead of walrus. I've never seen anyone use the additional flexibility that walrus affords because the code becomes too complicated and better at that point to just use two lines.
Also the discussions in comments makes me wonder if HN should implement syntax highlighting feature.
Random example I'm currently reading is: x, y = (0, 1) if True else None, None. The interpreter considers the last None outside the if statement. Gee, so you thought to need parentheses at first but not for the second one, and now it's a pitfall that parentheses are missing? The readme goes so far as so say "I haven't met even a single experience Pythonist till date who has not come across one or more of the following scenarios". Uh huh. Not me.
Add parentheses when in doubt, and when not in doubt, add parentheses for readability. In any language. Not excessively, like everyone knows what "if x == 2:" does without them in python, but for anything less obvious or more complicated than a=2*2+3, just do it.
"is" is not the same as "==". Confusing the two is a rookie mistake, yet is treated like a "gotcha".
For what it's worth, it took me a long time to internalize this (more than ten years after first learning Python), I think partly because the way most people teach Python is to say what scopes Python has instead of saying what scopes it doesn't have. While it's bad style to abuse this, I would definitely consider this a part of Python that a lot of working developers probably don't understand well.
class Foo:
a = [1,2,3]
b = [4,5,6]
c = [x * y for x in a for y in b]
that leads to NameError: name 'b' is not defined https://bugs.python.org/issue3692 (a++ if a==b else b++) + 1
Or something like this (borrowed from https://stackoverflow.com/questions/65024477/walrus-operator...): from datetime import datetime
timestamps = ['30:02:17:36', '26:07:44:25',
'25:19:30:38','25:07:40:47']
timestamps_dt = [
datetime(days=day,hours=hour,minutes=mins,seconds=sec)
for i in timestamps
day,hour,mins,sec := i.split(':')]
But this type of stuff is always greeted with Python's inflexible syntax. Recently, I learnt about Rackets's idea that "everything is expression" on (https://beautifulracket.com/appendix/why-racket-why-lisp.htm....) and was blown away to realize these are valid Racket codes: ((if (< 1 0) + *) 42 100)
or (+ 42 (if (< 1 0) 100 200))
The second one has an equivalent in Python: 42 + (100 if 1 < 0 else 200)
But the first one does not: 42 (+ if 1 < 0 else *) 100
> SYNTAX ERROR.Maybe it's time for me to get my hands dirty with Lisp.
from operator import add, mul
(add if 1 < 0 else mul)(42, 100)>
> 42 (+ if 1 < 0 else *) 100
Of course it does:
from operator import add, mul
(add if 1 < 0 else mul)(42, 100) a = 1
b = 2
((a := a + 1) if a == b else (b := b + 1)) + 1
although that's a lot to cram into a single line IMO.The closest I could come up with for the second example is this (noting that the datetime constructor requires both the year and month):
from datetime import datetime
timestamps = [
"2022:04:30:02:17:36",
"2022:04:26:07:44:25",
"2022:04:25:19:30:38",
"2022:04:25:07:40:47",
]
datetimes = [
datetime(*map(int, timestamp.split(":")))
for timestamp in timestamps
]
That said, it'd probably be better to use datetime.fromisoformat() or datetime.strptime() rather than manual parsing.And you could maybe use the operator module[1] for the last example, like this:
import operator as op
(op.add if 1 < 0 else op.mul)(42, 100)
[1] https://docs.python.org/3/library/operator.htmlPresuming this is indeed what behnamoh meant by a++, that `(a++ if a==b else b++) + 1` would be better written `++a if a==b else ++b`, which would then become the clearer `(a := a + 1) if a == b else (b := b + 1)`.
Why? This is absurdly unreadable. Why not create a function, give it a name, and use simple if/else constructs. Help the next developer out so they don't need to figure out your tricky expression.
Please rewrite this as
if a == b:
a += 1
else:
b += 1
and don't leave the other developers (or yourself two months from now) scratching their head about what that code does. You have more important things to do than trying to understand your code again and again.Furthermore, does that + 1 at the end run before or after the ++?
> (a++ if a==b else b++) + 1
I had to try it out (with a syntax that can run)
>>> a = 1
>>> b = 2
>>> (a+1 if a == b else b+1) + 1
4 # the expression evaluates to 4, wtf!
>>> a
1
>>> b
2
But this is not a += 1 which is syntax error and the point of your post.Simple example: any sort of graph algorithm that needs to know if two nodes are the same node.
https://stackoverflow.com/questions/72314247/how-can-ruby-co...
https://stackoverflow.com/questions/71439853/why-ruby-doesn-...
public class Main{
public static void main(String[] args){
Integer a = 300;
Integer b = 300;
System.out.printf(" %s == %s: %s\n", a, b, (a == b));
a = 100;
b = 100;
System.out.printf(" %s == %s: %s\n", a, b, (a == b));
}
}
// output
300 == 300: false
100 == 100: true
When combined with Java's auto-boxing, this can create some serious confusion >>> a = 300
>>> b = 300
>>> a is b
False
>>> a = 100
>>> b = 100
>>> a is b
True
>>>
CPython has a small set of canonical objects for integers.I once ran into a bug which had been there for years, since some Integer values had never crept over that magic limit, and auto-boxing allowed the numbers to be treated as int's all over, so it really took me some time to figure out what was going on...
>>> a = "this is a long sentence in a string literal"
>>> b = "this is a long" + " sentence in a " + "string literal"
>>> a is b
False
>>> id(a)
4386376272
>>> id(b)
4387721264
>>> a == b
TrueIf you test at the shell it apparently misses it (though it works for smaller strings e.g. "abc" / "a" + "b" + "c"), but if you put the same thing in a file and run it, it'll tell you the two strings are identical.
This right here is why it's critical to have developers of all levels across a team. Those of use who have been doing this long enough don't remember what caused us issues as juniors. We need the mid-levels to translate and remind us of these things lol
They are both objects; they only differ in size.
There is an enormous amount of quality training material available as well as excellent source code to major projects.
Most people would be better off just spending a hour or so with the tutorial at docs.python.org.
I get people coming to me asking for a job teaching Python when they don't know of some the material in the tutorial. Most have never read the FAQs and aren't aware of common solutions to common problems. To me, if you want to build your skills, start there. Get in the habit of reading docs, even boring ones, and occasionally read some of the source code from your favorite libraries.
That said, if you've already a very strong Python programmer, one possible use for these snippets is to help you root out any last, misconceptions of minor details.
But, you should resist the urge to show-off or use almost any of these "skills". There are a lot of cute things we can do with chained comparisons, but the only sensible things are: lo <= x <= hi or f(x) == g(x) == h(x). Anything else is too weird for communicating with the human beings. Likewise, you should only use "is" for None tests until you have strong understanding of identity guarantees.