Self Hell in Python
kmkeen.com
kmkeen.com
I do agree that OOP approaches can be a bit overused (especially if someone's been coding too much Java), and I've always appreciate that Python doesn't shove it down your throat. I like that you can lean towards more functional code like this example, or towards simple imperative code. But when I do want some classes, I find the Python mechanism mostly ok.
In JavaScript one would have to write: var bound_method = object.method.bind(object);
Python:
func = Foo.method
JavaScript: var func = Function.prototype.call.bind(Foo.prototype.method); " I don't really know much about Python. I only stole its object system for Perl 5. I have since repented."
Larry Wall
http://www.perl.com/pub/2007/12/06/soto-11.html self.minimum, self.maximum = min(self.minimum, self.maximum), max(self.minimum, self.maximum)`
This line is needed to 'defensively protect the input'), i.e. try to compensate when minimum and maximum arrive out of order. Firstly, I can't understand exactly why this defensive protection is not done before assigning the attributes, i.e. simply: # not too hellish, more readable on two lines
self.minimum = min(minimum, maximum)
self.maximum = max(minimum, maximum)
Secondly, I'm not very up on my Python Zen, but I think this naming of arguments followed by defensive protection isn't particularly Pythonic anyway. How about something like this? class Window(object):
def __init__(self, x1, x2):
# pass your dimensions in any order
self.minimum = min(x1, x2)
self.maximum = max(x1, x2)
def __call__(self, x):
return self.minimum <= x <= self.maximum def __init__(self, min=x, max=y): def make_window(min, max):
min, max = sorted([min, max])
return lambda val: min <= val <= max >>> timeit.timeit('x=4; y=3; x, y = sorted([x, y])')
0.49517297744750977
>>> timeit.timeit('x=4; y=3; x = min(x,y); y = max(x, y)')
0.2774021625518799What is being called "defensive" I would rather call "permissive". If at all possible, I would prefer raising a ValueError if x1 >= x2.
Of course, good luck having someone else understand your code...
So it can be called anything you want
BUT DON'T DO IT! Because you'll break the convention and anyone who maintains your code after that will hate you.
* As a throw-away variable:
>>> a, b, _, c, d = f(t)
* As a gettext call: >>> print _('Welcome to Python')
* And at the interactive prompt as the result of the previous expression: >>> 3 + 4
7
>>> _ * 10
70I would recommend to the author of the article one of those things:
(a) either find a programming language, where you like everything (I guess, unlikely)
(b) or write your own programming language
After some years in the business, I have learnt, that every programming language has its own "hell".Some hells are hotter, some are relatively mild. When you try to avoid every hell, you will end up writing your own programming language ... and be very lonely for the rest of your computer-life.
Pythons oddities are one of the mildest I have found yet, IMHO.
Author could have written
class Window(object):
def __init__(self, min, max):
if min > max:
min, max = max, min
self.min = min
self.max = max
to avoid repetition. If many args are needed, def __init__(self, a,b,c,d,e,f,...):
for k, v in locals().items():
setattr(self, k, v)
Etc etc there are endless idiomatic ways to avoid boilerplate in Python class Window(object):
def __init__(self, minimum, maximum):
self.minimum = min(minimum, maximum)
self.maximum = max(minimum, maximum)
Why set attributes twice? AFAIAC you only need to write "self." once for each attribute in the constructor. class Window(object):
def __init__(self, minimum, maximum):
self.min_func = min
self.max_func = max
self.tuple_type = tuple
self.minimum = minimum
self.maximum = maximum
self.minimum, self.maximum = self.tuple_type([self.min_func(*self.tuple_type([self.minimum, self.maximum])), self.max_func(*self.tuple_type([self.minimum, self.maximum]))])
def __call__(self, x):
self.x = x
result = self.minimum <= self.x <= self.maximum
del self.x
return resultI would propose making your inner function into a classmethod and moving away from the state of the class. This structure is well defined and is easily testable.
Alternatively, assuming you feature doesn't get more complicated, you can make your min-max ordering happen in one line:
self.minimum, self.maximum = sorted((minimum, maximum))
I think you can avoid self hell through careful programming.http://stackoverflow.com/questions/7054228/accessing-a-funct...
(Not technically true - deep code inspection is possible. That wouldn't be the same as unit testing, though.)
For example, given:
def add(a, b):
def inner():
return a + b
return inner()
You can write tests like: assert add(1, 2) == 3
But you can't write something like: assert add.inner() == 3
If inner() were doing something nontrivial, you'd probably want it to be standalone. closure = OuterFunc(var1, var2)
self.assertEqual(closure(var3), result)
inst = SomeClass(var1, var2)
self.assertEqual(inst.somemethod(var3), result)
When you think about it, there isn't much sense to calling an inner function without having called the enclosing function (the inner function typically depends on data trapped or closed by the outer function).Likewise, there isn't much sense in calling a method without having instantiated a class (methods typically depend on data stored in the instance).
> The double checking costs a few CPU cycles but saves a lot of keystrokes.
Not sure I agree with this, idiomatic python should be explicit, even if it means more keystrokes. def make_window(x1, x2):
lo, hi = sorted((x1, x2))
def window(x): return lo <= x <= hi
return window
and decided to give the window a custom repr. You could write window.__repr__ = lambda: '<%s..%s>' % (lo, hi)
but this sort of thing gets less fun the further you take it. So you actually rewrite all of this code with 'self' and '__init__' and it gets much longer and if you're me you conceive an antipathy to classes. What I wish you could write is def make_window(x1, x2):
lo, hi = sorted((x1, x2))
def window:
def __call__(x): return lo <= x <= hi
def __repr__(): return '<%s..%s>' % (lo, hi)
return window
BTW, 'foo hell' is too strong a term for most real foo in programming -- we don't need more encouragement to split into warring tribes. def make_window(x1, x2):
lo, hi = sorted( (x1, x2) )
window = type('Window', (object,), {
'__call__' : lambda _, x: lo <= x <= hi,
'__repr__' : lambda _: '<%s..%s>' % (lo, hi)
})
return window()
That's actually making a new class called 'Window', inheriting from object, with a call and a repr special method. I still haven't personally decided whether it's gross or not, but I am pretty sure it's not pythonic.Plus, if we really wanted to, we could do something insane, like
def window_repr(self):
return '<%s..%s>' % (self.__closure__[0], self.__closure__[1])
... but we are trying to program in python. We have standards.Python with nested scopes and classes has two ways to make an object. (Originally it had only classes.) A Python-like language with just one way of defining objects, using nested scopes like my "def window:" above, would be simpler and more concise and more Pythonic (in the sense that there's one and only one obvious way to do it). It's obviously too late to call that language Python, which disappoints me.
When the obvious way to do it is very different depending on how many methods an object gets (a function for one, a class for more), refactorings like my example get kind of annoying.
I do not think, that you should use nested functions that often, as already mentioned, they are hard to test.
You get more readable and testable code, when you put the code in a new method or in another class instead. You end up with less lines per method and you can pass the data explicitly, getting rid of the self.
I don't get this. Do you write test for every private methods in your class? That's totally unnecessary.
self.min, self.max = sorted((min, max))
This may be what was meant with the sentence that claims "but eventually you'll need a line like this outside of init" but I'm not sure what that part means. (I generally support functional programming practices, which Python leaves a lot of room for).
Or, you could subclass tuple.
Now, if you're doing something more complex, a class complete with self parameters is simply unavoidable.
[1] https://github.com/papers-we-love/papers-we-love/tree/master...
class window:
@staticmethod
def __init__(min, max):
window.min, window.max = sorted((min, max))
@staticmethod
def __call__(x):
return window.min <= x <= window.max
Or if you are just bored of self: class window:
@classmethod
def __init__(cls, min, max):
window.min, window.max = sorted((min, max))
@classmethod
def __call__(cls, x):
return window.min <= x <= window.max
But my favorite is: window = lambda *args: (
lambda (min, max)=sorted(args):
type('', (), dict(min=min,
max=max,
__call__=lambda _, x: min <= x <= max))()
)()I kind of like what Ruby did with "@" instead, but dislike some of the other syntactic choices.
Now, say, Python metaclasses -- there's a language feature I don't like.
https://gist.github.com/skatenerd/72281cadd2ae3e44d6cf
For some reason, stashing the variable inside of a list will fix the variable resolution..?
http://stackoverflow.com/questions/4851463/python-closure-wr...
nonlocal was added to fix this problem.
lists make a difference because they are a mutable container for other values, not because name resolution changes when you use them.
Python will by default shadow variables in an enclosing scope instead of overwriting them during assignment, to prevent accidental introduction of global state. You can use the "global" keyword to force overwriting the outer variable. The list hack also works because it is a mutable data structure, but I wouldn't recommend it.
This means that in the following code, "x" will refer to a local variable for the whole function body:
x = "foo"
def f():
print(x)
x = 42
f()
...so instead of printing "foo" (the value of the global variable) the print statement will raise an exception: UnboundLocalError: local variable 'x' referenced before assignment
If you meant it to refer to a global variable or something in the parent scope, simply put "global x" or "nonlocal x", respectively, at the start of the function. Again, that applies to the whole function body, so the assignment in this example would change the value of the global variable.There are lots of other awful hacks:
def a2(start):
def a(v):
a.v+=v
return a.v
a.v=start
return aTurns out the problem is the +=, perhaps it would be a bit clearer if written as:
grand_total = grand_total + to_add
which introduces a new local variable called grand_total shadowing the outer one. So it's straightforward enough, the pitfall is missing the shadowing.Note that you could use the outer grand_total without problems if you don't shadow it, e.g. using it as an rvalue (is that a pythonic term?) , and also note Python 3 provides a 'nonlocal' keyword to let you pull the outer variable into scope so it won't be shadowed.
def Window(x1, x2):
minimum = min(x1, x2)
maximum = max(x1, x2)
class _(object):
def __call__(_, x):
return minimum <= x <= maximum
return _()
The trick here is that the object instance takes values from its creating context as closure variables, rather than constructor parameters. Hence, it doesn't need a constructor at all, or any self references. It doesn't even need to use the self parameter it implicitly receives - hence why i can call it _ even though Peter Norvig couldn't! I've also called the class _, partly because it's sort of throwaway, in that it's never referred to after the instantiation which immediately follows its declaration, and partly just to wind people up.You can do this in Java too, using an anonymous class to avoid having to use a throwaway name, as long as the object you're creating implements an interface:
public static Function<Double, Boolean> Window(double x1, double x2) {
double minimum = Math.min(x1, x2);
double maximum = Math.max(x1, x2);
return new Function<Double, Boolean>() {
public Boolean apply(Double x) {
return (minimum <= x) && (x <= maximum);
}
};
}
Neither of these are outlandish hacks; rather, they illustrate the fundamental connection between objects and closures.In particular, the Java version is extra comical because the interface being implemented is functional, so there is a mechanical transformation (that my IDE will do for me!) which evaporates the anonymous class and uses a naked function, in the form of a lambda, just as the OP advocates:
public static Function<Double, Boolean> Window(double x1, double x2) {
double minimum = Math.min(x1, x2);
double maximum = Math.max(x1, x2);
return x -> (minimum <= x) && (x <= maximum);
}What exactly?
That you can write code that isn't OO? That's hardly a surprise, given that code has been written in other programming paradigmas for 50 years, before and after the invention of great OO languages like Smalltalk.
Or that you can write good Python code without "self"? Also hardly a surprise, as the OO-ness has been tacked onto Python afterwards.
That's exactly the reason why there is the "self hell" in the first place.
This is a myth, if I recall correctly.
Python had the OO-ness from the very beginning. It might seem "tacked-on" if you dislike Guido's style of OO, though.
I'm no expert in the internals of Python and its evolution, so I actually don't know how much that influenced the actual design of the classes later on but it always seemed plausible from what I've seen and also given the fact that Guido himself often says that the very first version didn't have a class statement.
Or here in this also quite interesting talk about "21 Years of Python" hhttps://www.youtube.com/watch?v=ugqu10JV7dk&t=49m50s