Curing Python's Neglect
zedshaw.com
zedshaw.com
Fos most of the rest of his issues, I think Zed's problem is that he thinks there's one obvious answer, but lots of other people would disagree with his one answer. Python tends to be either very opinionated, or very agnostic. Where there's One Way To Do Things, that's what you must do; except where there's obvious disagreement in which case it doesn't bless a single answer.
API doc autogeneration; I can guarantee that if anyone came up with a default tool with default output such that "doctool module html_directory" was all you needed (and indeed people have) then there would be widespread, perfectly legitimate, disagreement on any number of choices in the design of it.
(Prsonally, I think API doc autogeneration is entirely pointless, and indeed fundamentally the Wrong Thing To Do for most Python libraries. When I come across pages of JavaDoc-ish API documentation, I groan internally and just look at the source instead.)
You're absolutely right, but that doesn't mean Zed isn't on to something. How many different ways are there to open a subshell in python? There's popen, popen2, popen3, popen4 (I think), plus more (subprocess?). Yes, the past can't be changed, but it's still a good idea to stop and ask why there are so many mistakes like this.
Why is there a time module, and a date module, and a datetime module, and yet all three are broken?
The main question is What factors allowed the stdlib to get so quirky, and what can be done to fix that? Pointing out the stdlib is quirky is the first step in fixing the problem.
proc = subprocess.Popen(['tidy', '-q'], stdin=PIPE, stdout=PIPE, stderr=PIPE)
(tidyout, tidyerr) = proc.communicate(input=s)
if proc.returncode != 0:
...I think that a lot of this quirky stuff is due to the "uncool" factor of cleaning up such infrastructure level stuff. It's much cooler to be working on a proxy framework that talks to some new protocol. Making sure dates and times work well seems pedestrian. But it is actually tremendously important.
I was under the impression that Python 3 broke a lot of backwards compatibility, so could they not introduce new APIs with that? Or is that in fact what they did, and Zed is just complaining about 2.x?
I apologise if I'm mistaken on this, I've hardly used python at all. (though I keep meaning to learn it so I can finally stop writing bash scripts)
Get crackin'. It's a breeze to learn, and you'll be writing more robust scripts than in bash in no time.
Of Zed's complaints, I feel that two are misguided (time handling seems quite reasonable to me, and I've never seen an easier documentation system than Sphinx).
I think applying the "del" keyword to objects in a container is a mistake in the language, but that's how it was designed, so it probably wouldn't have been changed in the 2 -> 3 transition. It certainly won't be changed now -- the best that can be hoped for is for its use to be discouraged in the documentation.
Install/Uninstall of modules is fairly easy, but requires all concerned parties to work together. Otherwise you end up with half-installed, broken modules. There are PEPs in progress to work on this, but it's a social/political problem rather than a technical one.
Python's file-manipulation APIs are an ongoing calamity.
This is simply false. Modules were added, renamed, and modified. Module names now conform to pep 8 (Queue -> queue, SocketServer -> socketserver, etc). cPickle, cProfile and cStringIO are now accessible through their non-c counterparts instead of individually. Packages have been grouped - for example the http libraries are now in http.* instead of the top level.
Examples of API changes include the removal of sys.maxint and sys.exitfunc and friends, many libraries returning unicode strings instead of byte strings by default, and lots lots more.
Consult the "What's new in python 3.0" for a high-level overview: http://docs.python.org/3.0/whatsnew/3.0.html#library-changes
In my experience, the opposite actually happens. When something is written, people either use it and customize it over time, or they just find something else and say it didn't suit their needs.
It's when an extremely useful/general library is PROPOSED that it becomes a huge disagreement. This is, for instance, why C++'s Boost still doesn't have logging or a unified XML library, and won't likely have a garbage collection library before a C++ TR defines the interface: everyone argues over the color of the bikeshed, and the people who try to write the libraries get caught up in infinite microdiscussions over details.
Conversely, someone like Linus Torvald's success is that he just makes the things he needs/wants (ignoring bikeshed discussions), and people generally end up using his creations. When it comes to software design, a dictatorship is often more productive than a democracy!
"When I come across pages of JavaDoc-ish API documentation, I groan internally and just look at the source instead."
Amen :)
"The only thing democracy has ever really been good for is making sure that things people don't like get vetoed. Democracy is no way to plan; it is merely a way to restrict one's choices, possibly to none at all."
"No easy_UNinstall" - I'm one of the lucky few who has access to an extensive and sensible archive of easily installable and uninstallable Python library packages; I call it "Ubuntu". From what I can see, distutils is a good metadata format that helps real packagers like Debian and Ubuntu create real packages - but trying to write a cross-platform installer is just asking for trouble. Just unpack the package somewhere and set $PYTHONPATH in your startup scripts; that's what your startup scripts are for, anyway.
"rm -rf" - my understanding is that the contents of the "os" stdlib package are a thin wrapper around POSIX, and rmdir(2) won't remove a non-empty directory either (although Zed says Python's rmdir will remove a directory with subdirectories, but not files... that's odd). "shutil.rmtree" isn't in POSIX, so it isn't in "os" either.
"Time Converstion" - Zed says "If all they did was give me the exact same POSIX C API I’d be happy.", but so far as I can tell, the 'time' stdlib module basically is the POSIX C API, with braindead awkwardness fully intact. I'll grant that "calendar" is stupid and "datetime" is crippled, although the third-party "dateutil" module fills a lot of the holes in "datetime". (also, for as long as I've known of it, mx.DateTime has been freely available)
"API Documentation Generation" - It's not in the standard library, but I've been quite happy with the third-party tool "epydoc" for Python API doc generation, and it pretty much is as easy as "epydoc path/to/package" (and predates Sphinx and a lot of the other tools).
I have to agree with the rest of his examples, though - Python's had a long, rocky road from "procedural Unix scripting language" to "Object-oriented, Internet service-providing language", and although Python 3.0 has cleaned up a lot of the cruft there's still some oddness that remains (like the len function, or the del keyword). A lot of the standard-library crud has come from people saying 'here, I've found/written an 80% solution to this particular problem, let's put it in the standard library since it's better than the solution that's in there at the moment' rather than 'here's a problem, let's design a 90% or 95% solution for the standard library'.
This is always a cultural/community issue. It would need a cultural/community fix. And, given what I've seen in other communities, it should be fixed, as there's a lot to be gained by doing so.
> Then when you are told about it, you’d make up excuses trying to explain
> why it is totally normal.
That's a pretty incredible rhetorical device. "If you disagree, you're fundamentally incapable of reasoned thought."If I disagree with someone, then I often still grant that their opinion is the result of reasoned thought. In such a case it's the underlying assumptions that we disagree on and those assumptions are usually not amenable to reasoning.
A programming example is bracing styles: I have my preferred one and a colleague has his preferred one. I have my arguments and he has his arguments. In the end, it comes down to what each of us considers to be 'best readable', which is an entirely irrational consideration. We get along very well, despite our differences (and of course, for each project, we settle upon a style, based on exterior considerations like: what would be consistent with this clients' codebase).
Another example is religion: I'm a staunch atheist and my girlfriend is a Christian. 'nuff said.
A programming example is bracing styles
That's not an example of what he's talking about. He's talking about people creating reasons to explain away a real problem in a way that preserves reputation or personal peace-of-mind.What he's talking about is the pattern you get when you speak to an Alzheimer's sufferer who is trying to pretend they have nothing wrong. They can be extremely convincing to you and themselves for a while making up excuses for why things are the way they are. Eventually they exhaust (rapidly, if you make them reload context) and the facade falls away.
Programmers do this stuff all the time. "Oh it has to be this way because ... [some bullshit reason that doesn't explain why a customer has a reasonable objection to the software cracking its head doing a double blackflip when they asked it to step forward]". Is it reasonable to the educated impartial observer that a piece of enterprise software should crap itself on a null pointer exception and start silently failing? Not remotely. Will you find programmers defending their software when it exhibits such behaviour? All the time.
>> Then when you are told about it, you’d make up excuses trying to explain
>> why it is totally normal.
> That's a pretty incredible rhetorical device. "If you
> disagree, you're fundamentally incapable of reasoned
> thought."
No, that's not what Zed's saying at all. He's saying that people often respond to valid criticism of things they know well by justifying the behavior being criticised.Ie, people mold their way of thinking to suit their existing tools. Since they're familiar with how things are, they have problems seeing what could be.
It's not an ad-hominem attack at all. Think of the academic who has no idea of how to teach because he knows his subject so well he can no longer see it from the perspective of an outsider.
>>> exit
Use exit() or Ctrl-D (i.e. EOF) to exit
Enough said. If you bloody well knew what I typed, just exit FFS.Python's design is full of shortsighted, dogmatic decisions like this - don't even get me started with functions versus methods and __len__-esque line noise, argggg! And there should only be one obvious way to do something? Give us a break - that's just not how life, mathematics, or anything should work.
>>> exit = True
>>> exit
True
>>> exit()
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
TypeError: 'bool' object is not callable
The only way to then exit is Ctrl-D, but the shell will not prompt you to enter it!Nope.
>>> __builtins__.exit()"exit" merely references a function. It doesn't call by name alone. It can thus be passed to a function (a callback for example), saved to a variable, etc for later calling. Maybe not so useful for the exit function in the interactive shell but very useful for functional programming. This is an example of consistency, sometimes at the expense of convenience; not great for hacking, awesome for collaboration and large projects (imo).
>>> dir(exit)
['__call__', '__class__', '__delattr__', '__dict__', '__doc__', '__getattribute__', '__hash__', '__init__', '__module__', '__new__', '__reduce__', '__reduce_ex__', '__repr__', '__setattr__', '__str__', '__weakref__', 'name']
(pardon the dunderware)
Here's another way to exit the interactive shell ;)
>>> exit.__call__()
In CPython, exit is actually an object of type site.Quitter (the actual class is implementation dependent). site.Quitter overrides the __repr__ special method (which is called by the interpreter on any expression typed in the shell to print the result). site.Quitter also overrides the __call__ special method, so when an object of type site.Quitter is called, the overridden __call__ method invokes system exit.
>>> exit_class = type(exit) #gets a reference to the class
>>> my_exit = exit_class('bye') #the arg is used to print the message
>>> my_exit
Use bye() or Ctrl-D (i.e. EOF) to exit
>>> my_exit()
<python shell exits>
Minor inconsistency: typing "bye()" doesn't work so technically the message is incorrect. But I suppose they don't want you to be hacking exit() in the first place. >>> class ExitFunc:
... def __call__(self): __builtins__.exit()
... def __repr__(self): return "That's how!"
...
>>> exit = ExitFunc()
>>> exit
That's how!
>>> exit()His point was that the user already typed in "exit", and the interpreter recognized the user's intent to exit, so it should just shut up and exit already, instead of telling the user to exit in a different way.
>>> type(exit).__repr__ = lambda self: exit()
It will exit then the way you want :)Nobody do it for the sake of overall consistency. Yet you are free to change it with the one line of code.
del x unbinds x from whatever namespace it resolves to. How would you say c.remove(x) if you don't know which collection c represents? Now, of course you could argue that del x is too different from removing something from a collection to have it use the same syntax. But that's not about neglect. It's just a different opinion about what is consistent.
The reasoning is probably that removing a variable should always be done in one way and that is del x no matter if x is inside a collection we know or not. It's a very conscious attempt to make it consistent and that's not neglect. Whether it's a good idea I don't know.
Maybe a global del function would be better. If x is inside c and we know what c is we would say c.del(x) and if we want c to be resolved we say del(x). But is that really so different from what we have now?
More generally, my opinion is that the way functions are used in OO is in itself very inconsistent. If you have a function f and variables x, y, and z, it's mostly a consideration of implementation details that leads to a decision of whether it should be x.f(y, z), y.f(x, z) or z.f(x, y). Why does the user of an API have to think about or remember which one it is?
There are definitely some very nice things about Python. I highly recommend that all rubyists do a nontrivial project in it asap.
I wonder what a whitespace sensitive ruby would be like :)
%w(a b c d e f g).each do |i|:
puts i
foo = lambda do |x|:
x + 2
foo 3 myFunc = do
this
that
other
Or equivalently: myFunc = do { this; that; other; }But Haskell can pass around code blocks as well, with or without significant whitespace.
For example, I can define a function called 'forever' to run a sequence of actions forever like so:
forever x = x >> forever x
Where 'x' is any sequence of actions. You can pronounce '>>' as "followed by". (These are monadic actions, actually, but nevermind that.)I can invoke it like so:
myFunc = forever (do {this; that; other})
or: myFunc = forever $ do
this
that
other
The dollar sign is a function that lets you dispense with some parentheses. Think of it like an opening paren that closes at the end of the subexpression it's in.If the block is a function rather than a sequence of actions, Haskell uses '\' as a lambda for introducing anonymous functions.
For example, the function 'map' takes a unary function and calls it with each value from a list. Below I'll call it with an anonymous function which returns double its argument:
map (\x -> x * 2) [1,2,3]
This results in a new list [2,4,6]. forM_ [1..9] $ \i -> do
this i
that
other bla bla foo:
x
y
z
It's the same as this: bla bla foo
x
y
end
z
i.e. insert `end` when the indentation ends. foo(a, b, lambda x: x+2)
or foo2(a, b, foo):
foo() # this is like js
The lambda is closed by the next ) or , def test():
def boo(x):
x += 2
x += 3
return x
something_else(boo, 33, "hi")
test()
This is equivalent to how procs are often used in Ruby... collection.each(lambda x:
...
...
)
Or if the language doesn't require () to call a method, and uses `do` instead of `lambda`: collection.each do x:
...
...
If it's not the first argument: map(lambda x:
...
...
, collection)
Yes this is ugly, but using a multi line anonymous function in the middle of an argument list will always be ugly.a=[] dir(a) 'append', 'count', 'extend', 'index', 'insert', 'pop', 'remove', 'reverse', 'sort'
help(a.remove) remove(...) L.remove(value) -- remove first occurrence of value
help(a.pop) pop(...) L.pop([index]) -> item -- remove and return item at index (default last)
append(...) L.append(object) -- append object to end
insert(...) L.insert(index, object) -- insert object before index
And apt-get, pacman, rpm are not excuses, because I like to use Slitaz as my OS
If virtualenv (which I love) is something that plays havoc with package management - they should be merged, rather than not do anything about it.
Get pip and be enlightened.
(I liked your article.)
Python-core does not ignore bug reports, especially ones which really fix things, and come with a patch. Asking him to help isn't out of bounds.
Nit: Sphinx is a system for long-form, separately written documentation. Like manuals and tutorials; think stuff that requires indices. That's why autodoc is an extension, why you usually set up a separate directory, why you have the flexible build system, and so on. (Although all you have to do is create a file and type `automodule` and then the module name.)
Sorry - people misappropriating math jargon into the mainstream is a little pet peeve of mine. I agree with your main point though.
In IT, I've heard the phrase "mutually orthogonal requirements" for years now to indicate requirements that do not overlap.
In mathematics, orthogonal means perpendicular in a geometric-sense (think two vectors) but can also be used in other contexts with a different meaning.
This is why the word "perpenidcular" is not used, as sometimes your vectors don't really have "directions" in the intuitive sense of the word (e.g. the inner product space of functions).
When you decompose a goal into several linearly independent sub-goals - those sub-goals are said to be 'orthogonal' since they don't have any interaction/interdependence with one another.
It means at least:
1. having a zero inner product
2. (a square matrix) that is the inverse of its transpose
3. linear transformation that preserves angles
4. statistically independent
But you are correct, sir ;-)
Used to describe two things that are independent of each other. One does not imply the other.
"Common sense and intelligence are orthogonal. I've seen plenty of smart people with no common sense."
ie, if you had an enormous immutable value (in Python, a tuple with lots of entries), then in a hypothetical Python implementation which did call-by-value, it might be significantly slower than call-by-reference, on account of having to copy the entire tuple.
But this falls squarely into 'implementation detail' - if you're so inclined, you could implement either CBV or CBR in hundreds of other ways. Certainly from the perspective of time-independent program behaviour you shouldn't be able to tell the difference.
Maybe some of us already know what hemispatial neglect is, Zed.
mystuff[4] = 'apple'
del mystuff[4]
i'm not saying it's very elegant, but it makes sense.
It appears that this condition is more serious than we originally thought.