The Hardest Program I've Ever Written
journal.stuffwithstuff.com
journal.stuffwithstuff.com
My feeling though is the problem is they have a line limit. Maybe they should rethink their style. I'm serious.
Before I worked at Google, in 30 years of programming I never worked at a company that had a line limit. Adding a line limit at Google did not make me more productive. At first I thought "hey, I guess 80 chars makes side by side comparison easier" but then I thought back, hmm. I never had problems comparing code before when I didn't have a line limit.
Instead what I found was that 80 character limit was a giant waste of time. The article just pointed out a year of wasted time. I can point to searching and replacing an identifier and then having to go manually reformat hundreds of lines of code all because of some arbitrary style guide. I also had code generators at google that had to generate code that followed the line limit. I too wasted days futsing with the generator to break the lines at the correct places all because of some arbitrary line limit.
That should be the real takeaway here. Make sure each rule of your style guide actually serves a purpose or that its supposed benefits outway its costs.
There are a few reasons why we do 80 columns:
1. Human eyes have an effective line limit. The longer a line gets, the harder it is to scan back to the beginning of the next line. This is why paperbacks are taller than they are wide and why newspapers use several short columns instead of one wide one.
2. Being too narrow hurts readability, sure, but being too wide does as well. Also, even though many developers have giant monitors now, we also spend a lot of time on laptops, doing side-by-side code reviews, looking at code on blogs, etc. 80 columns is pretty friendly towards all of the various and sundry places where a user may be looking at some code.
3. We found it encourages better code. Dart is syntactically kind of a superset of Java, which means you can write Dart code that looks like Java. In particular, you can revisit some of the egregiously verbose naming practices that infected that community in the 90s. I see a lot of code like:
LoggedInUserPreferenceManager preferences = new LoggedInUserPreferenceManager();
A shorter column limit has worked as an effective nudge to get people to do: var preferences = new Preferences();
> The article just pointed out a year of wasted time.We'll get the time back. It's amortized over the amount of time saved by running it x the number of engineers using it.
> I can point to searching and replacing an identifier and then having to go manually reformat hundreds of lines of code all because of some arbitrary style guide.
The problem here is that you had to manually reformat it! Refactoring is a key goal of automated formatting. You can make a sweeping change to the length of an identifier and automatically fix the formatting every where it appears.
> I also had code generators at google that had to generate code that followed the line limit.
Code generators are another explicit use case. We have a lot of code generators now that produce completely unformatted code and the run dartfmt on it.
> I too wasted days futsing with the generator to break the lines at the correct places all because of some arbitrary line limit.
Should have used an automated formatter.
Yup.
My feeling is that this keeps gofmt much simpler, but it kind of punts the problem onto users. I wanted a more complete solution, even though the result is a lot more complex.
Often I want to format things that are more readable for me. Example (yea, not Google style guide example. Too lazy to dig one up)
uint32_t rgba8888 = ((red & 0xFF) << 24) |
((green & 0xFF) << 16) |
((blue & 0xFF) << 8) |
((alpha & 0xFF) << 0);
vs uint32_t rgba8888 = ((red & 0xFF) << 24) |
((green & 0xFF) << 16) |
((blue & 0xFF) << 8) |
(alpha & 0xFF);
I think the first is objectively more readable than the second. An auto-formatter is unlikely to ever format that in a "readable" way.Another simple example. If I have a many argument function
ctx.arc(xPosition, yPosition, radius, startAngle, endAngle, clockwise);
If I break that line I'm not going to break it between xPositon and yPosition, nor and I'm going to break it between startAngle and endAngle. An auto-formatter will never know that semantically those things are more readable when they are on the same line.Similarly you claim the short length encourages shorting names but you still run into plenty of situations where the code is far far less readable because of the line limit.
Example, assume 40 char limit
int w = desired_width *
scale_factor + padding;
int h = desired_height *
scale_factor + padding;
glBindTexture(
GL_TEXTURE_2D, someTexture);
glTexImage2D(
GL_TEXTURE_2D, level, format,
w, h, border, format, type, data);
glTexParameter(
GL_TEXTURE_2D, GL_TEXTURE_WRAP_S,
GL_REPEAT);
vs int w = desired_width * scale_factor + padding;
int h = desired_height * scale_factor + padding;
glBindTexture(GL_TEXTURE_2D, someTexture);
glTexImage2D(GL_TEXTURE_2D, level, format, w, h, border, format, type, data);
glTexParameter(GL_TEXTURE_2D, GL_TEXTURE_WRAP_S, GL_REPEAT);
Chrome in particular is full of code line wrapped into what is effectively obfuscated code. So no, I don't agree what a line limit has any point.You claim human eyes have a line limit. I don't disagree per say, but in my years before Google I never found anyone seriously abusing line length. Then again I never had to use Java but that's a separate issue. I hope Dart is not targeting Java's verboseness. I could make the same claim that unbroken lines, up to a point, are more readable, understandable and I will go on to claim that the 80 char limit at Google breaks that rule and ends up cause 20% of Google's code or more to effectively be obfusticated.
int w, h;
w = desired_width;
w *= scale_factor;
w += padding;
h = desired_height;
h *= scale_factor;
h += padding;
glBindTexture(
GL_TEXTURE_2D, someTexture
);
glTexImage2D(
GL_TEXTURE_2D, level, format,
w, h, border, format, type, data
);
glTexParameter(
GL_TEXTURE_2D, GL_TEXTURE_WRAP_S,
GL_REPEAT
);
That's assuming C, for C++ I might start it off like so: int w = desired_width;
w *= scale_factor;
w += padding;
int h = desired_height;
h *= scale_factor;
h += padding;
I might be tempted to use a macro but likely pass and just use shorter names all around.If it was up to me, programs on disk and in version control would basically be an abstract syntax tree, and any time a user viewed them they would be formatted using whatever style rules and screen size the user liked.
The way we do things at the moment seems like the equivalent of those marketing e-mails where all the text is made into a big image to "make it look right".
If they want a different limit, it's configurable.
That said, long lines in an existing project aren't a dealbreaker, if I have other reasons for contributing. (The one I'm working on now has them.) But if there's no reasonable way for me to write code in a language without breaking 80 chars, I'm not likely to choose that language for my own projects.
I've seen it happen - you'll get that one guy who likes to zoom a single editor to the full size of his monitor and then just type, type, type all the way to the far edge, making it impossible for anyone else to figure out what the hell he's doing without switching to his idiosyncratic editor setup. He'll see nothing wrong with nesting control structures fifteen or twenty levels deep, because it will all look completely reasonable on his screen, and he'll happily glom up absurdly complicated 30-40 character long identifiers because he has no taste and autocompletes all his identifiers anyway. An organization like Google can't tolerate that kind of crap and a strict code formatting guideline provides a simple first line of defense toward keeping it under control.
Limiting line length makes it easier for developers to work on each other's code, because they won't have to go resize all their terminals or change the length marker on a per-file basis. It's also a mechanical way to push developers toward shorter parameter lists and shorter, less deeply nested functions, which are good practice anyway.
(For what it's worth, I too have been programming for around 30 years, and Google is also the only place I've worked that had an official style guide with line limits. My reaction is the opposite of yours: I loved it, and I've tried to lobby for the practice everywhere I've worked since. I don't care what the details of the style guide are so long as they are enforced consistently - it's a really great feeling to drop into some far-distant file written by people you'll never meet and still find that the code is clear and consistent with the code you work on every day. I like 80 columns because it lets me fit three full size editors on each monitor, but I would be just as happy with 96 or 110 or 132 or whatever as long as there is some consistent limit I can count on.)
The width of a developer's editor window is so fundamentally a presentation issue - I have trouble imagining anything more so. Having line length limits is like mandating editor color schemes. How about I set my soft-wrap preferences the way I like and you can do the same?
Personally, I'd never know. I assume my editors can soft-wrap, but, IME, in terms of being able to easily work on code, lines that fit the window are better than long lines that aren't soft-wrapped, and long lines that aren't soft-wrapped are better than long-lines that are soft-wrapped, so I don't ever use soft-wrapping features.
I'm personally not that tied to 80 characters as a perfect line limit, but its a not unreasonable general guideline for most code in most languages. Like most guidelines, there's times when its inconvenient as a hard limit.
That's an unusual requirement for a code generator. What is the purpose?
Generated code generally is not read by humans or checked into source control. I suppose in the rare event where you want to read the code you could rely on your editor's line wrapping feature.
I have wasted far too much of my life arguing with people about how code should be formatted. Ideally I would just have a rule set up in source control to format the code on checkin and be done with it.. Then if I want a different style when I come to edit a file I can run whatever formatter I want on it and it won't affect anyone else.
I maintain a source beautifier too [1], and it's not as nice as I would like. One of the issues I run into is that the correct indent on a broken line is context dependent. For example:
while (someReallyReallyReallyReallyLongFunction() &&
anotherLongFunction()) {
loopBody();
}
is a nicer indenting than: while (someReallyReallyReallyReallyLongFunction() &&
anotherLongFunction()) {
loopBody();
}
In the first case, the two conditions are aligned which makes the code clearer. Does dartfmt handle this? If not, do you have ideas on how it might?Also, how does it handle invalid input? I may want to reindent my code before it's correct.
Also, did you explore constraint solvers instead of a graph traversal? It seems like they would be a natural fit.
[1]: fish_indent, https://github.com/fish-shell/fish-shell/blob/master/src/fis...
Either:
conditionWithReadableName: function() {
return
someReallyReallyReallyReallyLongFunction()
&& anotherLongFunction();
}
...
while (conditionWithReadableName()) {
loopBody();
}
Or if you don't want to make up a good name for your conditional: while (true) {
if (!someReallyReallyReallyReallyLongFunction()) break;
if (!anotherLongFunction()) break;
loopBody();
}
Or, in a more rules-based fashion: conditionWithReadableName: function() {
if (!someReallyReallyReallyReallyLongFunction()) return false;
if (!anotherLongFunction()) return false;
return true;
}
...
while (conditionWithReadableName()) {
loopBody();
}
I find it troubling to make up "good" indentation rules when the code to indent isn't well-written in the first place. Multi-line conditionals are an anti-pattern in itself (no matter if they appear in "while", "if" or "for").When all your while conditions are inside the block, I have to actually look at it to figure out that it's really just a standard while loop.
Good point! I didn't think of that.
Yes, it does! Correct indentation based for nested expressions is a vital feature for helping the reader understand the structure of the code.
The basic idea is fairly simple. Any place a line break may appear, you mark it with a number representing how deeply nested in the expression it is. So in code like:
function(outer(inner(first, second), third));
You would get chunks whose expression nesting level is like: function( 1
outer( 2
inner( 3
first, 3
second), 2
third));
When you line break, you ensure that its nesting level gets assigned a correct indentation level. Deeper nesting means deeper indentation: function(
outer(
inner(
first,
second),
third));
There are a bunch of interesting edge cases, though. Some times a nesting level doesn't happen on a line break, so it doesn't need to get an indentation level associated with it: function(outer(inner(
first,
second), third));
Here, levels 1 and 2 don't appear in line breaks, so we give the first indentation level to nesting level 3. There are stranger cases where you may assign a deeper level before a shallower one like: function(outer(inner(
first,
second),
third));
Shaking out all of the bugs requires a lot of work and a really big test suite.> Also, how does it handle invalid input? I may want to reindent my code before it's correct.
If the code doesn't parse, it just exits with an error. Being able to run it on incomplete input would be useful, but I decided it was out of scope.
> Also, did you explore constraint solvers instead of a graph traversal?
Good question! I have a little experience with them. I think the line between the two is sort of fuzzy. I didn't approach it directly like a constraint solving problem, but it does have some features in common.
while (first && second) {
third
}
We might assign nesting levels like this: while( 1
first && 2
second) 2
third 2
which could lead to this line breaking: while (first &&
second) {
third
}
This is bad because 'second' and 'third' are visually aligned, and so one might think that 'second' is in the loop body instead of the condition. We want to indent 'second' more than 'third', even though it has the same nesting level: while (first &&
second) {
third
}
This is what I meant by "context dependent:" here we have two chunks at the same indentation level but that want different numbers of spaces. Does dartfmt attempt to handle this? while (first &&
second) {
third // <--
}
So the body is indented less than the wrapped condition. Expression nesting is considered different from block nesting.So, the biggest problem is in assigning weights to the break candidate instructions is this stream. This can only be done by a lot of experimenting, I could not find any formal method or a passable heuristic.
On each opening parens put the current line length (=indentation) on a stack. On closing parens pop from the stack. On newline use the top of the stack as indentation.
We're all just doing the same thing in one way or another :) Good work and nice article.
Interesting, I figured Google was handling data serialization in their own way, and now I know[2]
[1]https://developers.google.com/protocol-buffers/docs/overview [2]https://github.com/google/protobuf
For anyone that hasn't tried it, grab the Dart SDK and the IntellIJ Dart plugin. Takes less than 5 minutes to setup. It's been a great platform for building server side stuff - I haven't tried it for front end web stuff. It took about 3 reads of the language tour (https://www.dartlang.org/docs/dart-up-and-running/ch02.html) and about a week and I already felt very comfortable with the entire platform.
\o/ It improved a lot in the past two months. That's when rules and the new splitter landed.
The most complex single piece of code I ever wrote was a scheduler. The user could specify a pattern of when events should be raised (eg on this date, at this time, every other hour on the last day of every month, at midnight for me in this TZ on a server in another TZ, etc), and the scheduler would raise the events at the prescribed instant(s).
That took about 9 months, and my biggest takeaway was that how humans measure time is completely f*ed up!
The Tidy.pm module is 1.1M in size, and over 30,000 lines long. I have much respect for formatters now, I thought the job they do was an easy one.
Fantastic looking sourcecode, btw,
https://metacpan.org/source/SHANCOCK/Perl-Tidy-20150815/lib/...
From memory there were other requirements indenting the first word of each paragraph things like that.
As the article alludes to it was a surprisingly complex problem - we also had to worry about memory allocation as we were using C. I remember I was quite proud when I got the sample text (which was a few paragraphs from "The Hobbit") to render correctly.
I've never thought about writing a code formatter I just trust emacs to format my code for me. I'd be interested in digging up my old code and seeing how easily I could modify it to operate on source code.
For example, if there's a common pattern among a set of lines, I'll often line them up vertically to make the repetition clear and focus attention on the differences rather than the commonalities; for example:
if (foo ||
quux ||
baz) {
....
}
let foo = 10
quux = foo * 2
baz = quux + 1
in baz * 2
fields = ['name', 'address', 'country',
'dob', 'status', 'salary']
To me, those few extra spaces make it easier to glance over the code than without: if (foo ||
quux ||
baz) {
....
}
let foo = 10
quux = foo * 2
baz = quux + 1
in baz * 2
fields = ['name', 'address', 'country',
'dob', 'status', 'salary']That really struck me back then, and I've kept it in mind whenever I hear about code beautifying/indenting.
[1] http://steve-yegge.blogspot.com/2008/03/js2-mode-new-javascr...
Similarly, covering all your food in Doritos dust doesn't always taste great but it ends the interminable soul-crushing arguments about what flavor things should have.
Flavour, texture and aroma are learned appreciations, and are not universal. So I think the comparison with code style is quite apt.
My point being that I don't really see the persuasive value of the analogy. It's a false-equivalence. The point was that yes many people often care deeply about the formatting of the code (myself included), but discussions around formatting are almost always a form of bike-shedding. A better analogy would be "randomly selecting the restaurant doesn't always lead you to your favorite place, but at least it prevents the interminable discussions about where to go."
I disagree with the comparison, but go ahead with it, what's your point?
Isn't it annoying when a globally optimizing tool switches back and forth between "all arguments on one line" and "all arguments on separate lines"? E.g. producing overly complex whitespace changes in diffs for small "triggering" changes?
My first thought was CSS.
Wish I could gain more context on how big an arena of these types of programs. I'm a bit lost as to how important code formatters and beautifiers were until reading more on the difficulty of writing such a program by Mr. Nystrom.
JSCS [3] added autofixing a while back for most whitespace rules, and ESLint has just begun autofixing as well [4]
[1] http://jsbeautifier.org/ [2] https://github.com/rdio/jsfmt [3] http://jscs.info/ [4] https://github.com/eslint/eslint/pull/3635
It is fascinating that we don't have a definitive method for formatting, yet.
In this case, we don't know where in the solution space the best solution will be. It's not even that easy to tell if we've found it. So that rules out simple pathfinding algorithms like A*.
Even if the output of the formatter isn’t great, it ends those interminable soul-crushing arguments on code reviews about formatting.
With a configurable formatter you just move those arguments into the style-guide discussions and still have disagreements between people and teams using different configured values.