PEP 8 Modernisation
hg.python.org
hg.python.org
The default wrapping in most tools disrupts the visual structure of the code,
making it more difficult to understand.
It's 2013. Let's fix the tools.You think we would have gotten past the point of having to manually figure out "where should I break these lines for the best readability." Code is meant to be consumed by machines, and a machine should be capable of parsing the stuff, figuring out visual structure, and wrapping dynamically to account for window width and readable line-lengths in a way that preserves the visual structure.
If I'm working with code on my phone, I'm going to have different preferences than on my tablet, or my laptop, or doing side-by-side comparisons, or my 30" display, or a projector.
Probably the best current examples of what can be done with this are docco and it's friends. iPython notebook is another example of what we can do when we decide that hardware improvements in the past couple decades may be taken advantage of to improve production and consumption of code, rather than declaring that no progress shall be made after 1976.
Progress sometimes requiring taking a couple steps backward if we've reached a local maximum. An unwillingness to compromise backward compatibility - a perennial commitment towards designing for the lowest common denominator - limits the progress that can be made.
Are you referring to the unwillingness to bump up the maximum line length? I don't think it's because of backward compatibility. I limit just about every line of code I write in any language to 80 columns, and my horizontal screen real estate is 5760 pixels.
I hope you're joking...
I was not intending to downplay the significance of code being intended for human consumption.
Anything from 45 to 75 characters is widely-regarded as a satisfactory
length of line for a single-column page set in a serifed text face in
a text size. The 66-character line (counting both letters and spaces)
is widely regarded as ideal.
— Robert Bringhurst, The Elements of Typographic Style, 3rd Edition, §2.1.2Code is simply easier to read by humans when it isn't stretched out obtusely in one direction. We should be making code easier to be read by humans. Machines don't matter.
Newer versions of emacs do nice code wrapping (dynamically, to the width of the buffer, without modifying the buffer text). Xcode does nice code wrapping. I've stopped manually wrapping my code, and it's great.
Edit: I know emacs can probably view diffs. That's not the point, unless you expect every window to be the same width, and every tool to wrap everything exactly the same.
Mentally account for one more line or two?
The same diff tool works for text files, latex and other markup files, for any computer language, and so on. I can write downstream tools that take the output of the diff and do further processing, without worrying about what the output might look like for brainfuck. And how long would we have to wait for Torvalds to add a C++ or Java extension to git? :)
Raw text is a limited medium, and working at the level of plain text analysis rather than semantic analysis does limit how much our tools can achieve. The UNIX philosophy of having many small, text-based tools and chaining them together represents a common platform, but it’s also a least common denominator platform.
As long as we allow that glass ceiling to remain, the best our tools will ever do is push it incrementally higher, one tiny step at a time. If we want to make big leaps, we’re going to have to sacrifice some of that generality so we can use more powerful but specific tools. Unfortunately, that means that any new programming language wanting to take on the established standards needs not only a compiler/interpreter, but also the rest of the ecosystem: a comprehensive tool chain, ample library coverage, documentation and training resources, and so on. It is also, almost inevitably, going to need some standardised way of bridging the gap to today’s established languages for interoperability and backward compatibility purposes.
This is why I think the view that programming languages should be designed optimally for humans is short-sighted, and I suspect most of the big success stories in the coming years will be languages that were designed with clean semantics and easy parsing in mind. Those languages will better support building that surrounding ecosystem, and a good language with a good ecosystem is more practically useful than a slightly better language with a limited ecosystem.
We can add semantic analysis. I have no problem with that. But removing or harming the ability to do text analysis might not stop us from using more sophisticated tools, but it does stop us from putting tools together quickly in a way not previously predicted. This is why Unix is so powerful. If you haven't read The Art of Unix Programming, I urge you to do so. It puts the argument and examples forward far more convincingly than I can.
Would moving to semantic tools harm the ability to use our existing text-based tools? I'd say so. A simple tool like diff works better, for example, if you take your big list of Python imports and put each one on a separate line, and keep them sorted. Patching works smoother this way too, reducing the likelyhood of the need for manual conflict resolution. If we eliminate doing this kind of arrangement by hand, and instead start relying on a semantic editor, we'll lose this ability. I have yet to see a tool that does semantic diffs, patches and merges better than diff, patch and git do.
Well, OK, but we’ve been using these text-based tools for a year or two now, and I don’t see many radical advances taking place in how we use or combine them. Are you sure you’re not chasing an illusion here?
The freeform text-based tools represent a great deal of flexibility, to be sure, but they were conceived at a time when flexible text manipulation was about as much as one could hope for. Today, we can do more.
I have yet to see a tool that does semantic diffs, patches and merges better than diff, patch and git do.
As long as everything is limited to manipulation of freeform text files, perhaps you never will. That doesn’t mean better tools aren’t possible; it just means they aren’t possible within the constraints you’re choosing to impose.
I'm not claiming recent advances. I'm claiming existing power that has been around for decades, which we would lose if we compromised the text tooling available today.
Have you read TAOUP? Do you understand the extent of the power that existing text tooling gives us today? Are you experienced in the advanced use of the existing tools, so you are able to make comparisons about their power?
> As long as everything is limited to manipulation of freeform text files, perhaps you never will. That doesn’t mean better tools aren’t possible; it just means they aren’t possible within the constraints you’re choosing to impose.
This is backwards. We move forward when people show how it can be done. Please show us how we can improve diffs, patches and merges by moving to semantic data structures over text, without compromising any existing capabilities. Even just illustrating specifics of how these tools might work, rather than implementing them, will do something for your argument. The onus is on you.
Why would we lose it? The power of those tools isn’t in a particular executable, it’s in the algorithms they embody. For example, it is useful to compare two text streams reasonably efficiently and identify differences. How those differences are then presented obviously matters, but if you’ve got the algorithms and the ideas underlying them, producing a new tool to apply those ideas in a different context is the easy part.
The only significant difference I see is that if you made a major change, for example adopting a more structured storage model or using some sort of action/history analysis to better capture a programmer’s intent, then you would have more data to use in your algorithms, and you might therefore be able to present more interesting results.
Have you read TAOUP? Do you understand the extent of the power that existing text tooling gives us today? Are you experienced in the advanced use of the existing tools, so you are able to make comparisons about their power?
Yes, though I find your emphasis on that one book a little surprising. For one thing, the UNIX philosophy was established for decades before Raymond wrote that particular work. For another, I seem to recall that he gives examples of both text and binary formats being useful in the book. I don’t think his point was that text formats were good and non-text ones bad; I think he was arguing that things like adaptability and composability were good and that flexible and standardised formats helped to achieve those things.
We move forward when people show how it can be done.
Right, so why aren’t the programming language community picking up on decades of research and industrial progress with databases and HCI? Programming languages and the related tools are, fundamentally, just a user interface to design and control a complex, highly structured set of data.
Please show us how we can improve diffs, patches and merges by moving to semantic data structures over text, without compromising any existing capabilities.
You’re begging the question, by starting from the position that having an equivalent to today’s text-based diffs, patches and merge tools is desirable. I don’t think that is necessarily true.
As a programmer, I want to be able to specify how my software should work, and I want to be able to explore and modify that specification effectively, and I need to be able to do these things in collaboration with others. My claim is that to do those things much better than we do them today, we may need to move to a different representation than freeform text and then build new tools that are designed to solve our problems in terms of that new representation.
The problem is that there is so much momentum behind text-based formats today that we are effectively stuck around a local maximum. No one individual could possibly meet your challenge today, and I’m sure you were well aware of that when you made it. That doesn’t mean the community as a whole couldn’t do it, but it would need some serious collaboration, which realistically means one of the heavyweight organisations with the resources to bootstrap a whole new software development ecosystem would need to throw its weight behind such a project. Unfortunately, most if not all such organisations are commercial in nature, and the commercial incentives don’t align with moving in that direction.
Can you explain what problem you think might occur? Its not clear to me what problems you anticipate.
An example: I type "git diff" into one window, and look for some particular change area in my editor in a different window. If the wrap and alignments depend on the widths of my windows, then they won't match unless the widths of my windows also match. And if I have to make them match, then the original ideal of "make your window however wide you want and it'll just look right" is defeated.
Furthermore, I don't think this is true: "And if I have to make them match, then the original ideal of "make your window however wide you want and it'll just look right" is defeated." Being able to resize both, with the constraint that they must both be kept in sync (which to be clear, I don't consider necessary), is better than the alternative: no effective resizing at all.
You should probably look into integrating git-diff with your editor though (use git-difftool).
If I'm diffing in an external tool (like gitx, or just on the command line), then I just see (possibly) long lines. So far, it hasn't been an issue for me. Obviously, diff, as a line-oriented protocol, breaks down as lines get really long -- but it's still just code, so it's not like my lines are ever insanely-long. :)
I also use a mode that highlights my current (physical) line in the file, so I still tend to have a very good sense of what the physical line in the file is, despite the wrapped visual display. E.g., http://imgur.com/a/wAXHJ
One of the oft-touted benefits of wrapping lines at (say) 80 chars is it makes it easy to do side-by-side viewing of files -- using dynamic wrapping gives you this same benefit, but even more so, since you can heads-up different files at whatever width your current display happens to have. (Or however many files you want to have side-by-side.)
Also, there's a nice side-benefit to diffs, which is you don't get the noise that comes from a change that forces a manual rewrapping -- e.g., maybe I decide my variable "id" was too generic, so I change it to "frobnackId", but then this pushes some function call over the 80-character limit. If I'm manually rewrapping, I get a weird diff of multiple lines changing over the rewrap.
(On the other hand, a definite down-side is if you do find yourself using an editor that can't wrap nicely, code with long lines can be quite annoying.)
Why can't my editor tokenize my code and then show it in the format I want? Why do I, as the programmer, have to worry about whether one whitespace convention works for you versus me, when you could just come up with whatever scheme you like and view code that way?
Because editors are written by humans. And as of now, we haven't sold a whole lot of problems with their use.
For example, you get an IDE with lots of features, but a subpar editor, crammy UI and GC pauses (3 of the five most popular are written in Java).
Or you get something like VIM, with a great editor, but subpar compatibility with the rest of your system, not very good understanding of the code (refactoring, tokenizing etc). '80s GUI capabilities etc, ad-hoc collection of plugins to fix basic pain points (like file navigation).
Or you get Emacs, with millions of configurable options, a subpar extension language, script in various stages of great and rot, '80s GUI capabilities, etc.
In general, we lack tools that run the whole gamut: great editor, fluent shortcuts (either Emacs or Vim style), 2013 GUI capabilities, refactoring and intimate knowledge of program syntax (to the level of understanding the AST, no BS regex used for syntax highlighting), embedded REPLs and terminals, etc.
Something like a Lisp Machine + Smalltalk UI + Light Table + Visual Studio + Vim/Emacs combo.
Excuse me, but vim has much better compatibility/interoperability with the rest of my system than all the IDEs written in Java.
An in-editor debugger?
Does it offer a REPL?
Does it work with your build system and your SCM system without some ho-hum third party plugins?
In-editor debugger? I admit this is less than ideal.
A REPL? Languages suited to REPLs have REPLs, and a lot of them support SLIME, allowing me to interact with the REPL from vim. For everything else there's tmux.
Build system? :make and errorformat
SCM integration? ^Z. Or fugitive I guess. (a third party plugin, yes. But if you think it's ho-hum, you probably haven't used it).
The question then becomes, why configure a transformer when you can just format it right the first time?
http://www.reddit.com/r/java/comments/1j7iv4/would_it_not_be...
These are useful for static code analysis and finding congruence with typesetting conventions:
https://pypi.python.org/pypi/flake8
Well, partly. Code is actually a piece of writing intended for two very different audiences - one human, one machine. The machine doesn't actually care how long your lines are - this is an optimization for humans.
And I hear you about fixing the tools. Wouldn't it be easy if everyone used the same editor. But alas, this isn't a problem with the tools. It's an inherent limitation of text. That "figuring out visual structure" bit only works if your editor deeply understands your language. And different editors understand languages to different degrees and in different ways.
They say parsing is a solved problem, but I dare you to parse Ruby or Python correctly.
A friend at an old job had a piece of perl code where one of the lines would be a comment or not when you ran it, depending on user input. He would bring it out whenever someone suggested writing something in perl.
For example:
# Is this some_method with an empty block, or an empty hash?
some_method {}
# Is this a regex or a comment?
#
# It depends how many arguments "whatever" takes, and you
# don't know that until you run the code that defines it.
#
# Interpretation is implementation-defined. It just so
# happens that the MRI determines that the below ISN'T a
# regex because it has a SPACE character just after the "/".
whatever / 25 ; # / ; raise SystemExit.new;Plus, just getting the code to do the thing I want is enough work; let's not also start worrying about how to write code in such a way that most tools will understand how to wrap it. That sounds like bringing all the joys of CSS to my plain text. Because the tools are _not_ going to be perfect. Have you ever seen Perl code with punctuation in a comment at the end of a line, because some idiot syntax highlighter got confused about where a regex ended?
1) Lines longer than 80 are an indicator that your code is getting too verbose, at least for Python. 2) You should be taking the time to figure out how to maximize readability of your code. And breaking the lines is damn near instantaneous when compared to the time taken to pick a good variable name or decide on overall design structure. 3) Not everyone will have the same editor. Your code should stand on its own and be readable in its own right.
I completely disagree, partly because I prefer to err on the side of verboseness, and partly because 80 is just ridiculously short. Look at the examples in PEP 8:
def __init__(self, width, height,
color='black', emphasis=None, highlight=0):
What about that code is "too verbose"? It's an extremely simple constructor with 6 arguments (5 explicit).I understand that readability is subjective, but I find it extremely hard to believe that anyone honestly finds that more readable on two lines than one.
FWIW, I do. Especially since the new line is a clear point of separation between required and optional arguments.
But I also find HN extremely hard to read because of the ridiculously long lines, and often find myself jumping to the wrong line when I reach end of the screen.
Whom should I listen to... Hmmm, tough choice...
I mean, I've just seen Gerry Sussman in a recent video saying he doesn't know how to compute, so you must be a winner.
> Aim to limit all lines to a maximum of 79 characters, but up to 99 characters is acceptable when it improves readability.
When applying the rule would make the code less readable, even for someone who is used to reading code that follows the rules.Yeah, I could have 300 char wide terms, or 600 if I really wanted to. There was a reason for 79. There isn't for 99.
Edit: Glad to see that section might get an update http://bugs.python.org/issue18472#msg194086
Edit2: Gudio van Rossum proposes a patch https://codereview.appspot.com/12269044/diff/1/pep-0008.txt
I think I'm fine with it provided that readability doesn't suffer if my editor or viewer wraps line >79 characters. I don't want to be forced to make everything wider to see that extra readability. But if making it wider makes it more readable even if my editor wraps the line, then that makes sense.
In vim, I'm thinking about using:
set colorcolumn=79
set textwidth=99edit: you're correct, that's the rationale in the pep. It's not why I stick at 79 though, I do it for git log -p, etc.
OTOH, 0x0d definitely counts toward the limit!
In other words, it's a line ending, not a line separator.
I've changed my mind because I've started doing a lot of code-reading on my phone. It's got a big screen (Galaxy S3) but even still, 80 columns is the perfect width to fit all the code on my screen without needing to scroll.
What's old is new again. We went from tiny monitors to huge monitors and now we're back to tiny "monitors" again. Anywhere where I am able to make the determination, 79 columns will be the standard.
I read a lot of code.
While we're at it, can we ditch the convention of indenting arguments to line up with the method name? A single tab per-nested level will do.
https://gist.github.com/radiosilence/6133713
I mean how can anyone say that the former is better than the latter?
If you have a line that would suffer under 79 chars, let it go a little long, but no farther than 99.
http://bugs.python.org/issue18472#msg194089
--- a/pep-0008.txt
+++ b/pep-0008.txt
@@ -159,12 +159,11 @@
Maximum Line Length
-------------------
-Aim to limit all lines to a maximum of 79 characters, but up to 99
-characters is acceptable when it improves readability.
+Limit all lines to a maximum of 79 characters.
EDIT: formatting/description
+Some teams strongly prefer a longer line length. For code maintained
+exclusively or primarily by a team that can reach agreement on this
+issue, it is okay to increase the line nominal line length from 80 to
+100 characters (effectively increasing the maximum length to 99
+characters), provided that comments and docstrings are still wrapped
+at 72 characters.
+
+The Python standard library is conservative and requires limiting
+lines to 79 characters (and docstrings/comments to 72).
+
https://codereview.appspot.com/12269044/patch/1/1001There's a reason newspaper stories are in columns.
I get it, it looks nicer with 4 spaces. I'm still not convinces it's a straight win.
(Of course, you should never mix tabs and spaces, and I'm glad Python 3 enforces this)
before, some people criticized my choice believing that the PEP8 favored one of the 2 approaches, when in fact it just acknowledged that spaces were used more often, and in no case that was an excuse to change pre-existing code... there was some confusion
At least now I can/have-to concede that there's one preferred way
OTOH: I still find subpar the handling of space indentation in most editors nowaday (emacs, vim, gedit... you name it: they all behave similarly)... I could write some emacs-lisp to make the editor behave as I wish, but I never found that hitch worth scratching enough
(I know, if vim was second nature to me, this point would probably be moot... but I'm still using the arrow keys a lot :) )
Thanks for clarifying!
Why have one of the most common semantic units in source code (indentation) represented by four characters when there is an otherwise unused character that has already great editor support. Tabs work well in every editor I've used (except notepad.exe, which can't change their display width). Spaces only work well in smart editors.
Editors aren't the only place code appears. I see it in my terminal, in git, in various web-based tools, in email, in browser textareas, etc. I don't want to reconfigure all of these things to make your code look like you see it (if they're even possible to reconfigure), also potentially breaking other applications that take for granted that a tab character aligns to 8.
set tabstop=4
set shiftwidth=4
set softtabstop=4
set expandtab[0]: http://www.python.org/dev/peps/pep-0008/#maximum-line-length
In fact, I always found space-alignment cumbersome and sometimes ugly... always indenting with tabs was thus perfectly consistent
But that's the past, now :)
Indeed, a quick test with python -tt shows that the tab-indent/space-align method does not throw an error, while using tabs and spaces to indent does.
Both Emacs and Vi correctly allow you to indent by 4 spaces rather than foolishly inserting a TAB and indenting by the standard 8. Other editors like Sublime Text stupidly insert a TAB and expect you to muck with how TABs are rendered in order to indent correctly.
Since many people don't understand that distinction, they map TAB to 4 columns which is incompatible with most formatters (including all printers and HTML). Switching to spaces is the only solution which really fixes this.
"translate_tabs_to_spaces": true
to your Preferences.sublime-settings to have spaces instead of tabs in Sublime Text.And because emacs and vim are still "dumb" editors in some ways: They still conflate the arrow keys with "move one character", when they should really mean "move one character, unless you are in non-alignment indentation, in which case move one level of indentation, which happens to correspond to four spaces since this particular file is python." They have similar problems with backspace and delete, regexes and search/replace, etc.
Further, what if you prefer 2 or 3 space indentation? If you use semantic tabs, you can have that simply by changing the display width of the tab character. It needn't be equivalent in width to 8 characters. To my knowledge, emacs and vim are both too dumb to display code indented with spaces at a different indentation width than saved in the file.
Note that when using semantic tabs, tabs only represent semantic indentation. A new block adds a tab. Further indentation for the purpose of aligning multi-line statements must be done with spaces. Tabs following spaces on a line are always wrong. Tabs are not to be used to align to columns. If you want to do that, use spaces.
For example:
def foo():
--->if (long_named_function_that_returns_a_boolean() and
--->....thing2() and thing3()):
--->--->do_stuff()
And again, with your editor reconfigured to display tabs as two characters wide: def foo():
->if (long_named_function_that_returns_a_boolean() and
->....thing2() and thing3()):
->->do_stuff()
Basically, tabs get you several niceties with partially-dumb editors like emacs or vim. Indenting with spaces gives that all up and all you get in return is the ability to view code with notepad.exe.If your indentation scheme only works when configuring your editor to mark whitespace, we might as well go back to MUMPS and indent with dots.
This, alone, is why I always use tabs, no matter what any style guide says. People have different tastes. Using tabs allows any editor to indent according to personal tastes.
[1]: http://www.python.org/dev/peps/pep-0008/#pet-peeves/pedantic
go to the "Maximum Line Length" section. Now try to understand the changes wit the side-by-side one.
Now go to the original article link and do the same. It's much easier. At least for me.
"Imports should usually be on separate lines ... No: import sys, os"
Why not? I don't really spend a lot of time studying stdlib imports, nor need them on separate lines. Pyflakes will yell at me if something was not imported.
After all, one does (and is legal in the pep): from subprocess import Popen, PIPE, ... So how is that different conceptually than doing from stdlib import os, re, sys, ...
I certainly agree with separating stdlib, 3rd party libs, and your own modules by blank lines.. but why should usually trivial stdlib imports need so many lines?
If I see `import os, sys, foo, bar...` then I don't get that information without actually inspecting the rest of the line.
The difference is minor, but skimming is valuable.
"You should use two spaces after a sentence-ending period.":%s/\. /\. /g
Anyway, in the grand scheme of things this is really insignificant :) The fact that PEP8 exists at all is just one of the things making Python great!
Use a single word space between sentences. In the nineteenth century,
which was a dark and inflationary age in typography and type design,
many compositors were encouraged to stuff extra space between sentences.
Generations of twentieth century typists were then taught to do the
same, by hitting the spacebar twice after every period. Your typing
as well asyour typesetting will benefit from unlearning this quaint
Victorian habit. As a general rule, no more than a single space is
required after a period, colon, or any other mark of punctuation.
Larger spaces (e.g., en spaces) are *themselves* punctuation.
— Robert Bringhurst, The Elements of Typographic Style, 3rd Edition, §2.1.4(Some people go as far as to recommend less than word space after a period or comma, since the blank area above the punctuation mark contributes to the total visual white space.)
I want to set for myself how much spaces represent an indentation, not someone else, it depends so much on your setup, the monitor, resolution, font settings, editor setup, if 2 spaces are fine, or 3 or 4...
Plenty of C++ code uses UPPERCASE_SNAKE_CASE_NAMES for macros, but that doesn't seem so bad because macros are ugly and should be used less frequently than classes and functions. (I hope!)
Python code does tend to use UPPERCASE_SNAKE for "constants" as well, probably inherited from C (sometimes via Perl or Ruby).
Or is there already a Ninja-IDE plugin that formats code?
pip install autopep8
Just pip search pep8 to see all the tooling.
It's not a "Ninja IDE" thingjobber, just a command line tool.