Writing Good Code Is a Lot Like Writing Prose
alexkras.com
alexkras.com
One of the key pieces of advice I give to new recruits is "write a paragraph about your intent, what do you want this code to do?". This, it turns out, is non-intuitive for folks as it requires first understanding what is there (OP notes this well) before jumping into coding. I think the root of this is that you're paid to write code and you want to impress right from the start.
Thus I found it very interesting when a few Q4's ago I was tasked with creating a new recruiting site for engineers ( https://engineering.twitch.tv ) to better understand our culture. The central focus of this site was to be a series of blog posts which highlighted core work that had recently taken place. My job quickly became editor in chief; working with engineers of various levels to understand what they've been doing, what was cool about it, and get them writing about it. Seven blog posts were identified and I made calendar events to check in seven days from then to read drafts... seven days later no drafts existed. As I dug into this it was apparent that they'd all assumed they could jump into the writing prose much like they do code, they all remarked that writing prose is actually really hard. I provided a template of my own pieces over the years and by the following quarter we had three pieces written!
I credit TPW for giving me words to describe this to engineers ( http://tom.preston-werner.com/2010/08/23/readme-driven-devel... ) and my high school English teachers for giving me the love of words!
It functions a lot like BDD, or at least my version of it, of keeping me on track for the problem I'm trying to solve at the moment and felt like it led to generally better code quality once I started doing it.
This so much. I find that my strong liberal arts education in high school has made me a much better developer. Things like meter and pacing are quite applicable to writing code, and like you I have only my high school English teachers to thank for that.
So many of the processes around writing (good) code are similar to those around writing (good) prose, it amazes me that non-technical writing and literary analysis aren't more strongly preferred.
It takes some time, rereading a bit. discussing it with other people who are reading it, making the right amount of notes in the margin, consulting the dictionary when you don't understand a term. If it's in a language you don't know, you can translate it slowly, but knowing the basics goes a long way.
I like this way of looking at it because it has more to do with gaining understanding than with promoting some semantic style in regards to documenting
I find when I am writing code, I try to use the same principles. I don't want anyone to look at two isolated pieces of code and think "what the hell, these don't even belong in the same universe, let alone the in same repo." Whether through comments or just good self-documenting code, I want things to connect in some way.
Jeremy Ashkenas made the connection in a piece of his writing about coffeescript as well. Kind of an ah-ha moment for me.
However, this article leaves me wishing he'd show examples of what he's talking about.
I'd like to see his code to determine whether it reads more prose-like.
When I want the constant for the user's language, everyone knows they're looking for LANGUAGE, in principle. And long identifiers aren't quite the problem any more with autocomplete. But some library- and language authors insist on saving 4 bytes by naming it LANG, or even just L. Or they somehow end up with this abomination:
> locale
LANG=en_US.UTF-8
LANGUAGE=
There are better examples where more variations can be found in the wild. And there are some abbreviations that somehow just feel terribly ugly to me, but I've apparently succeeded in erasing them from memory. They sometimes remind me of people who use a plus sign instead of the word "and"–it betrays a certain lack of respect for the language,You have your primitive ideas for some chosen rules, chosen such that the desired product is a valid idea in those rules. Then you combine your ideas through concatenation and composition to get more complex ideas. Sometimes you realize your rules aren't quite right so you change them some. You keep doing this until you have a big idea that fits your definition of a "solution" (whatever that means) and then you're done.
I built a prototype for an Atom extension a while back that rendered Comment blocks with a markdown parser, inline, right in the editor.
I'm hoping that the VSCode team gets around to implementing the extension hooks necessary to implement this at some point, since I've switched away from Atom.
It would require us to move away from strict file-and-text-editing as the end all of programming, which makes many uncomfortable of course. And you get into lock-in issues.
Actual comments already do that if the editor hides them by default; you don't need different syntax or abandoning pure-text representation, just the right editor features to do what you want.
(ROOT (SINV (VP (VBG Writing) (NP (JJ Good) (NNP Code))) (VP (VBZ Is)) (NP (NP (DT a) (NN Lot)) (PP (IN Like) (S (VP (VBG Writing) (NP (NN Prose))))))))
(ROOT
(SINV (VP (VBG Writing)
(NP (JJ Good) (NNP Code)))
(VP (VBZ Is))
(NP (NP (DT a) (NN Lot)) (PP (IN Like)
(S (VP (VBG Writing) (NP (NN Prose))))))))From the code you can tell what the program actually does. From a comment you can get some idea about what the authors intended, at least at one time in the code's history.
https://github.com/illumos/illumos-gate/blob/master/usr/src/...
Defining the objects, the actors, and their interactions is part of writing prose, but there is more: you have to present them in a logically consistent manner, and furthermore in such an order that the reader is not confronted with an idea she has not yet been prepared for. The way that the components of a program are assembled has analogous consistency and ordering requirements.
Writing a program is a bit like writing a specific form of prose: an argument for a proposition. In the case of programming, the proposition is that the program will do whatever it is that you require of it.
That's a big assumption. I assume the OP meant not quite literally 'prose' but 'good, clear writing' which can certainly be achieved using the built-in keywords of a language combined with well-chosen identifiers.
Many comments that are written could be better handled by restructuring the code they reside in the midst of, but many code comments can’t be done otherwise, and to attempt to do otherwise would be convoluted.
To take an example of some code I wrote today:
// Object.values is from ES2017 and Overture uses it; anything that has it will
// have the other things from ES6 that we need, so we don’t need to add the
// weight and load time of core-js for them. Browsers that thus need core-js
// are approximately: IE, Safari≤10.0, iOS≤10.2, Chrome≤53, Firefox≤46,
// Edge≤13?. No up-to-date browser needs it (for better or for worse).
if ( !Object.values ) {
dependencies.O = [ 'core-js' ];
}
This comment is about explaining why a thing is as it is. Self-documentation by naming tends to only take you as far as what; for a case like this, it would have led to something like this: var needCoreJs = !Object.values;
if ( needCoreJs ) {
dependencies.O = [ 'core-js' ];
}
This wouldn’t have explained why core-js is needed—or why I considered Object.values a suitable test.I tend to draw a distinction between single-line comments which are only sentence fragments, and prose. Prose is beautiful in code.
A part of the illumos codebase was linked to elsewhere in these comments, and I think it satisfies my point also: https://github.com/illumos/illumos-gate/blob/master/usr/src/.... 600 lines of documentation up front is more than I’d typically expect, but it’s well done in this case. After that, there’s liberal application of explanatory comments where they can help.
The rustc borrow checker is another example that I like: https://github.com/rust-lang/rust/tree/master/src/librustc_b.... What is now README.md used to be somewhat shorter (though still long and involved!) and in mod.rs. Look at both files. The documentation here is all about implementation detail; without this extensive prose, I’d say: good luck understanding borrowck. With it, it’s involved, but quite attainable, and you can approach the code without needing excessively deep and intricate understanding of everything about rustc.
Comments can be done badly, but they can also be done really well.
> And good code will contain a lot of prose.
I disagree. Good code can also contain no comments. It's not an either/or thing.