The Shapes of Code
fluentcpp.com
fluentcpp.com
How much longer will this rumor persist that "good code" doesn't need commenting? Good comments explain things that aren't immediately obvious from the code.
So that leaves the “why”. To my mind, code comments are the worst place to record why a particular implementation exists. The why often needs collaboration with non-coders.
The worst is finding some snippet from stack overflow or GitHub without a reference to the issue to back up why this thing is here or perhaps a TODO: with an improvement
Commenting the purpose of the code and maybe a link to some documentation allows a future reader to see if the code still fulfills the requirements.
Writing the rationale is the most important comment you must not skip.
And there would be links to design documentation so that it can be kept up to date. (For non-programmers.)
This isn't always possible, but it's far more possible than many developers seem to think.
When I read code I care about what the code is supposed to do and why it's supposed to do that.
There is code that passes the current tests and is logically sound but no longer fits the business requirements. Knowing that it was implemented to solve a certain use case helps the reader / reviewer see whether the code and the test still fit or if they should be updated.
The best comments I've come across comment the intention and any context that isn't immediately obvious from the function.
def clean_start_year(x):
# Account for the 7 year offset in the database
return x + 7
def process_record(record):
record['start year'] = clean_start_year(record['record'])
It's perfectly obvious what this code does, but without the comment it's totally unclear why it does what it does. int frob(const int x, void *context)
Performs a frob on x in the given context.
Parameters:
x - int param
context - a pointer to context
Returns an id of the frob performed.The better a variable / function is named (and the better its scope or purpose is limited to one thing), the less it needs comments. So there are a lot of self explanatory variables or functions with doxygen comments that state exactly the same thing the name already does, without adding anything besides noise. And then there are some poorly named things where the exact same thing happens: comment re-states what the bad identifier states. Sigh. Genuinely useful comments are rather rare.
Use your imagination. The "cleaning" routine might not be a one-liner. Maybe it depends on multiple functions from another file/package/library.
const database_year_offset = 7;
fix_start_year(year):
return year + database_year_offset;
process_record(record):
record['start year'] = fix_start_year(record['record']);
In this version with one less magic constant (and renamed function), the comment would look very redundant.DougBTX had the same idea, at the same time. I don't think their version needs the comment either. EDIT: Dang you removed a perfectly fine post :)
Informative comments cannot always be "factored out" like this. Do you at least agree that a ticket or bug #, or a link to an issue in an issue tracker, would be appropriate?
Could I have instead massively restructured the feature to prevent the possibility of the bug the comment was explaining the workaround for? Probably. But that would have increased the number of lines of code to make the change by a factor of something like 20x-200x and would have required far more testing. It also probably would have resulted in something overall more complicated.
My process of writing code is to of course code a proof of concept in order to get a feel for the problem domain this could be considered a spike in agile parlance
Then the primary and alternate flows can be coded against a basic test suite
followed by a refactor and then a comments pass
Depending on what comments arise perhaps another refactor may be in order
I typically try to follow the spirit of TRUE software
> ...to explain each line of code.
Commenting every line is a hint to a more pervasive problem. This section has nothing to do with "comments are bad".
Indeed. I recently came across an article[0] which does a good job of debunking this idea, and describing different types of comments and when they can be valuable. Have a read if you're skeptical about comments.
> My second remark is that our intellectual powers are rather geared to master static relations and that our powers to visualize processes evolving in time are relatively poorly developed. For that reason we should do (as wise programmers aware of our limitations) our utmost to shorten the conceptual gap between the static program and the dynamic process, to make the correspondence between the program (spread out in text space) and the process (spread out in time) as trivial as possible.
While Dijkstra later talks about using procedures to index into the call stack, needlessly adding to the depth of the callstack for single use functions seems excessive.
With this logic, I actually think "The paragraphs with headers" is the way to go since as TFA states: "You know that the algorithm operates in steps, and you know where the steps are located in code."
http://www.u.arizona.edu/~rubinson/copyright_violations/Go_T...
Refactoring has rediscovered this trick, among others.
It'd be better to ask if the code is more complex than it needs to be.
I liked the terms John Ousterhout uses. For some module of a program (e.g. a method, class, package, etc.), it has an interface and an implementation. The interface is "what you need to know to use the module", the implementation is "how it's done".
Complexity comes from dependencies (e.g. more interfaces you need to know about), and from obscurity (e.g. things you need to know about, which aren't obvious from the interfaces given).
In some cases, many small methods would be an increase in complexity overall. In some cases, one big method may be an increase in complexity.
Code is a set of instructions for the computer, and for the next person who has to maintain it (“ Programs are meant to be read by humans and only incidentally for computers to execute.”)
Cooking instructions are split into steps. Assembly instructions are split into steps. LEGO instructions are split into steps. When drawing instructions aren’t split into steps the Internet turns them into a cultural phenomenon (Step 2: draw the rest of the fucking owl).
Only programmers think they are immune to this ridicule, smash all the steps together, and then get salty when others complain, ignoring a series of luminaries who beg in every format of media available, and for decades, to do otherwise.
Seriously, organize your code, separate the steps. You’re killing us.
Why have a function that prints a whole line? Isn't that just smashing together a bunch of small steps to print one character at a time?
https://webcache.googleusercontent.com/search?q=cache:http:/...
More functions => More complexity => Bad.
Also: decreased legibility, less opportunity for compiler optimisation.
https://docs.microsoft.com/en-us/previous-versions/windows/d...
Apparently 13 different indentation levels is good.
But sometimes, the shape can be deceptive:
Not saying I go by shape alone, but if it looks messy, that's a warning flag.
If the bug wasn’t in the ugliest part of the code, then that code likely had two bugs in it.
I found examining code shape to be an effective method for assisting in algorithm recall in some college courses. Now working with larger codebases in industry, I find it really useful for navigating around large files or remembering where to go for particular snippets of code.
I wonder if this phenomenon has anything to do with different spatial tricks people use for memorization and recall (the “memory palace” technique). I’ve never thought of myself as having good spatial reasoning at all, but this kind of thing makes me question the boundary between “spatial” and symbolic / verbal thinking.
It's easy to recognize, and the importance of recognizing it is right in the name.
I'm not sure what happens in a functional language. Maybe you could just think about numbers of functions. In Haskell sometimes it helps to write things in a pointfree way and you can spot some more general function to replace some noise.
Someone made an Atom plugin: https://atom.io/packages/squint-test
In the computer field, for the first time, it was realized that the unification of hardware engineering and software engineering on the logical model.
It has been extended from Lisp language-level code and data unification to system engineering-level software and hardware unification.
and it brings large industrial production theory and methods to software engineering. It incorporates IT industry into modern large industrial production systems, This is an epoch-making innovative theory and method.
This is the [Pure Function Pipeline Data Flow v3.0 with Warehouse/Workshop Model](https://github.com/linpengcheng/PurefunctionPipelineDataflow).
This is an interesting article, and I intend to go listen to the podcast, too.
In the past I've made efforts to try to speed up the process of learning a codebase. I'd do things like copy the code into a text processor, then shrink the font so I could see which files (classes) were biggest, and look for repeating shapes like the author mentions. I'd also use code cleaners to point out troublesome classes and UML full-trip tools to try to get a good sequence diagram out of a piece of code.
It was all cumbersome, unfortunately. Most days I just use 'grep' to help me figure things out. I'm still looking for helpers, though. I'm hoping the podcast helps.
Wouldn't this invite consternation during code reviews because it doesn't have anything to do with the current feature? Or are you one of the lucky few who gets to exercise professional judgment when working?
There is an intrinsic and deeper geometry that textual programming hides.
More indentation equals more mental context required.
But although these shorter, less-indented functions are inherently easier to understand, they're not necessarily freer of bugs...
When it happens to me in Python or Ruby, I return immediately from the smaller branch and move the other branch out of the if, to the main level.
When it happens in Elixir... It doesn't happen because I don't use ifs there. I write two versions of the same function. One matches the condition for the then branch, the other the condition for the else one.