The Documentation Triangle, or, why code isn't self documenting
sourceless.org
sourceless.org
I prefer to think about documentation in terms of four audiences:
- the student (tutorials)
- the chef (how-to guides)
- the theorist (explanations)
- the technician (API references)
Divio [1] has an excellent set of articles that explains this approach to documentation.
We want multi-level docs combined/spliced/interwoven together so we can get our work done without wasting gandalfing the docs.
> While redesigning the Cloudflare developer docs
...oh, so this is the reason for CF's docs totally uncomprehensible and confusing mess. They've made getting to anythin so comfusing that AWS seems a breath of fresh air in comparison.
I dunno, maybe other peoples' brains are wired totally different than mine, but boy, these new fever of "writing better docs" seems to produce the exact opposite effects whenever applied.
The few enojoyable docs nowadays seem to be written by old-school unixy geeks and/or people that are not programmers or worse, "documentation developers": physicists, mathematicians, biologists, quants, data analysts etc. seem to do a nice job when they have the time.
Further it's explained that this is where the "understanding" bit comes in. Ugh, so this school of thought is where some of the terrible modern documentation comes from...
I want the explanations/understanding first COMBINED with a minimal how-to guide serving as illustration for the explanation.
Then a quick link to an API reference that INCLUDES snippets of explanations and usage examples (that you'd find in tutorials or how-to guides).
Separating thinking about documentation this way leads to horrible documentation where you start with dumbed-down tutorials missing lots of crucial information, and then you have to put in tons of unnecessary effort putting disparate things toget ther in your head. FFS, programmers are rarely students since most of what we learn is learning-on-the-job and never chefs or theoreticians. Tutorials always waste your time so you'd prefer to start with something much denser... only that the only thing denser you easily find are the API docs, and you lack the explanations to understand them and usage exmaples so you have to hunt them for yourself... yuck.
As an example of GREAT documentation that combines things well (the API docs contain concise math theory / formular / explanations AND usage examples), take a look at some of the Pytorch docs: https://pytorch.org/docs/stable/generated/torch.nn.Conv2d.ht...
That said, I'll agree that these should be interlinked. A tutorial that doesn't ever link to practical how to's is terrible, and please include references to theory in your api docs so I can gain a deeper understanding when the design doesn't make sense to me, etc etc. It's easier to write and read when you are assuming the audience of a particular work; you can switch roles when you click on a how-to article, but trying to inhabit all personas at once is exhausting and unproductive.
There's no personas, developers don't have three different brains they keep in the fridge and swap the ones in their head with.
Just admit that sure, interweaving information to get good quality docs is exhausting. And that we invented this "write for an audience" crap to simplify the work of the writers - "we're lazy or we have limited resources for writing docs, that's why we're doying it this way".
It's the same for writers and journalist, the "write for a specific audience" trick is way to lower the effort needed to get reviews / publishers' attention / clicks / views / shares etc. You do it when you're getting started if you ever want to get started I guess. But large scale it produces shallow literature, hard to fact check or contextualize journalism etc.
You do it for yourself (the writer) or to save resources for your companies (sure, interwoven/mixed docs are hard to maintain - I'm not even sure how eg. Pytorch's team manages to do that, but probably being the leading ML frameworks attracts top talent from the whole planet to one open-source project so it's doable for them...). But don't sell the bs that you're doing the reader/consumer any service, keept it hones to yourself and others!
> a single disorganized pile
Nobody proposes a "single disorganized pile". Sure, it's hard to eg. put bite-sized examples + condesed 'why' explanations into API docs and to maintain such docs, but it's doable - in the ML community some people use notebooks for docs wiriting and sometimes even for development by generating the code from them (sure, this is utterly extremist and I don't like it myself, but it works for some), some kind of renaissance of "literate programming".
I don't say we revive literate programming, but going haf-way between that and "a disorganized pile" is what I meant.
IME it's mildly obvious when the comments are obsolete or out of whack. It's testing one hypothesis. Whereas with no comments, there are basically an infinite number of possible hypotheses as to what the code does.
1. Some high-level documentation explaining how the key systems work and interact conceptually, to give context.
2. Comments for anything non-standard (i.e. we're using some weird encoding here instead of JSON because of XYZ reason)
3. Comments for the tricky extra complicated bits
For everything else, well-written code following standard conventions and good naming practices is enough.
Also choose a statically typed language, so that your function signatures will give you a useful contract to work with, and it will mostly be pretty easy to trace the intent of your code.
Furthering this, avoid:
- "Stringly typed programming": where everything has a static type, but they're all `String` (or `Int`, `Float`, etc. instead of something with domain meaning). This gives very little information to the reader, and makes it easy to mix-up values.
- "Boolean blindness": where we query a bunch of booleans to determine what data we can extract, e.g.
Result::hasUserId : Boolean Whether this result contains a UserId
Result::getUserId : UserId Returns a UserID if hasUserId; otherwise error
log("Request came from " + myResult.hasUserId? myResult.getUserId : "anonymous guest")
Compare this to e.g. Result::getUserId : Option[UserId]
log("Request came from " + myResult.getUserId.orElse("anonymous guest"))You can even go through a particular file’s associated changes to see how it’s morphed over time.
Reading code was always strictly better. Truthful and often easier to parse and understand.
And with code you can actually converse by putting in breakpoints, disabling parts of it and observing it what it is actually doing, how and infering why.
Documentation was always dead almost always outdated prose.
Even when it was up to date and accurate I often understood what the person writing a comment actually meant only after I read and understood the code it was referring to.
If you want to read it breadth-first, you'll want some context as to what the three children do. If you want to read it depth-first, the comments are somewhat redundant, because you're going down to the source anyway.
I rarely have a desire (and even less often a strict need) to understand the whole call graph of a function. If it works fine, tests pass, and I want to understand / change one specific behaviour, I need to grok a slither of the entire possible call graph. Comments help guide where I'm going then. If grokking the whole thing, I see how comments are at best a waste of space.
To me, this is where the most valuable comments are. The more that your code is a representation of business logic, the more you have weird "whys" to fill in.
Any kind of plan that involves everyone on the team inferring the same "why" is a terrible plan in my book.
Which is important context for figuring out how it works now.
Similarly, in a production environment, I found that when there was no time to keep the docs up to date with the latest changes, I marked them on the cover page as PRELIMINARY just to put readers on alert. And so, similarly, should a document not getting the attention it needs be marked as NOT WELL MAINTAINED ? I get the feeling this idea would not fly. Not in a commercial environment anyways.
Lying comments have often helped me because it's often useful to know that a particular person at a particular point in time believed something to be true, even if it isn't true any more and even if it were never true. In a pile of spaghetti code, a good lying comment can point to the needle in the haystack where the root cause of a bug is hiding.
Just because the comment is at odds with what the code is doing doesn't mean that the code is right. The code wins by force because the comment isn't executable.
Sometimes the comment is at odds with the code because the code should be the way the comment says. It's not that way because (1) someone changed it. Or (2) it never was that way; the comment is just wishful thinking.
If you want to know which of those two it is, you do the archaeology.
Using something like ADRs[0] can help provide the structure the team can look for when it comes to "why" documentation. I recommend a standardised README across projects. If it's "what" comments look for language tooling that will execute documented examples as test cases. If it's "how" then make sure there's value in it - do you have disaster recovery exercises, does your incident response team use your "how" documentation to triage issues, do stakeholders outside your development team get consulted on contextual changes in your application, etc.
[0]: https://brunoscheufler.com/blog/2020-07-04-documenting-desig...
For example: sometimes you may have a codebase where whenever "X" is done, there is an automatic retry mechanism of failure. But then you have one spot where there is no automatic retry. And it turns out there is a very subtle and non-obvious reason why an automatic retry here would be a problem.
The code then deserves a comment explaining what this subtle issue is, and and why it means automatic retries must be avoided here. Without such a comment, the next person to touch the code may well just assume the automatic retrying was forgotten, and add it, causing say the painful data corruption bug to return.
Some tips for better habits:
- whenever changing existing code, read the comments
- if you know the changes are going to be big, just delete the old comment first
- if the changes are minor, but change the code add a TODO: check comments on top of the existing comments
- try to keep comments tied to the thing you are commenting, comments that describe interactions between components go to the outside scope, this way there are less surprises where comments need to be updated
- consider writing comments describing what you are going to do before writing the code. On top of writing the comments this acts as an additional reality check and quite often leads to better code as well
Teams of two or more individuals can establish norms with only a little more difficulty than individuals, and can provide accountability to those norms better than an individual.
Have editor plugins that highlight if the hash no longer matches, then you can evaluate if the comment needs to be updated or not and update the hash.
This would eliminate the problems raised about working on teams and people not following procedures.
Nobody wants to memorize your code. And you shouldn't want to either. You can't memorize code that other people are contributing regularly to. So either you won't know it when you come back, or you're subtly pressuring everyone to move their changes outside of your code, and leaving your code alone. Even if it has a choke hold on data flow in the system.
But that almost misses the point, which is that yes, code that doesn't look like what it actually does can cause people to miss the real problem. But code that looks like it's a potential source of the problem being investigated steals attention from the real culprit. And no amount of memorization is going to completely solve that problem. I think people mistake the notion of 'code smells' as the act of an overly fastidious mind, but the neat freak is forever asking the slob "how do you find anything in this mess?" and that problem plays out in code, but amplified. Like if you were looking for your credit card and it wasn't a matter of whether you put your wallet where it 'belongs' or if it was in your bedside table, but instead you have 9 wallets scattered around your room full of customer loyalty cards.
Next up is the "how", although it's not as likely as the "what" since it's not as common (as it should be - you don't need to write a "how" about every list iteration).
The least likely to be a lie is the "why", since the circumstances leading to a "why" rarely change, and when they do, you usually end up replacing the whole section.
And so, your "what" should be almost entirely in self-documenting code, your "how" should be used judiciously, and your "why" probably won't be a problem.
The more unnecessary comments you write, the greater the risk of them going stale due to maintainers glossing over them. When they're a rarity, you take more notice.
// Because of [sane business reason] we need to remove all 89's
something.removeall(75);
I found this because people in case 75 were logging a never ending stream of tickets for months that the program was doing weird things. No shit, sherlock. Talk to business, everyone agrees 75 makes no sense at all and it really really should be 89. So before we change this back to sanity, let's check the git history: Date: 3 years ago
Message: Urgent fix for major downtime. (No detail of course)
Commit by: someone who left a little bit later
Change: 1 line patch, changes 89 to 75
Asking around, nobody technical knows anything anymore, the whole team got laid off around that same time 3 years ago and the new team basically rediscovered everything from scratch. Some business people do remember it, it was Very Bad, they don't know exactly what it was or how it got fixed, but anything that might cause it to happen again is strictly forbidden.What would you do? I got reorganized to another team, lucky me.
In fact, even the fact that the comment is now wrong tells you that the change was ill-considered and needs to be revisited.
Any fix should produce at least a test to guard against regressions of the bug it's trying to fix and encode the correct business logic.
The test name would then explain the why. The test variables present the story of the what. And finally, the API calls tell the story of how how.
Yes a unit test will show you that something is changed, but often it's the comments that don't get updated.
Hell maybe just let us flag the comments that will need to be changed if X or Y unit test fails. Just some way to better link comments to code adjustments.
Definitely put the comments in the docsting or equivalent for your language, not in the middle of the function.
The best place for comments is nicely grouped under (or above, an elegant C++ convention) the interface. If the function signature is 200 LOC away... the problem isn't the comments.
But they should be there.
2. People, need, diagrams. People NEED a picture.
At the minimum, the bare minimum, your application needs a high level and low level diagram explaining how it works.
It your first week onboarding significantly easier in my experience, the first time I encounter a codebase. I'm sure people make do without.
Lotta science behind it, humans work much better if there's a diagram with coloured shapes/lines involved.
I've found that if I'm struggling to create a diagram for something in Mermaid then I've gone too deep in my explanations and it's time to step back.
High level architecture diagrams, user journey diagrams, low level diagrams, are probably all distinct topics.
I feel all three are extremely important, but that's bordering on my opinion.
> but I find that low level diagrams usually tend to be a lot more trouble than their worth
Out of date low level diagrams are surprisingly useful.
One of the things that I've found is that, with extremely high LOC projects (LOC definitely measures complexity), you'll find that low level diagrams are worthwhile, at the cost of a time investment in generating them.
When/how to do them is a certainly a fine art. There's no right answer.
Perhaps there is a correlation between complexity due to a high number of lines of code and benefits provided by low level diagrams. If there is then it does pose the question of when you should invest time in creating and maintaining low level diagrams and when it would be better spent reducing the size of the codebase.
Would agree that the higher level documentation (diagrams, process interaction, etc.) can be written down once the dust settles.
And if the team is properly collaborating, they are linking issues and merge requests with commits, providing historical snapshots of the ever evolving context!
They're just not designed for pedagogy in the same way docs should be.
And yes tests are not the best candidate to learn how to do stuff. The best candidates for this are out of band documents / media / training.
For example, Rust has documentation which is executed as tests (doctests). That way, when the code changes, the docs change automatically.
To your mocks counter-example: How exactly would you propose to do it otherwise? If you write down an example which works now, how would you guarantee it still works if i.e. you don't use mocks but a say, a real database, which you then change.
You would be in the case of 3 green parallel lines which are red and intersect. In a powerpoint/doc you can write whatever you want. Doesn't actually need to be true.
Naming is hard. So is documentation. I did that just to trigger fallacy trolls, but sometimes smart naming is documentation. Maybe that hard work of naming is worth it.
Verifying that code flow makes sense via code review guarantees that code is documented well enough.
As does handing it to the intern, without any "training," and asking for a new feature. Or an expert in a different part of the stack but not this code. As does playing "spot the bug" and having most people understand the code, even if they don't see the bug. Heck, just passing through some linters and style analyzers can get you pretty close to not needing documentation.
There are many things that count in different ways. What does your team want/need/desire? Does the answer change when it's someone else's code and their own?
The flip side is a lot of prose in the human lingua franca wherever you code. (Probably American English but that's another screed.) That means people with coding chops also need literary chops. Or a friend who has them. It means that you're not only judged on how well you solve problems with code, not only with how easy that code is to inherently understand, but also judged on how well you can explain to someone else what it is and how to use it.
If you really need that heavy documentation, hire a tech writer. Just like you probably specialize in a kind of dev, you can find specialists who specialize in making things make sense to devs.
(I often recommend diagrams just to make boundaries clear to everyone "at a glance" but don't think those really count as documentation in this context.)
In general, after you get to a piece of code you need to have thre questions answered:
- what problem does this solve? - this is actually the problem how-to guides solve, in practice your particular use case is NOT even close to the example use cases, so you'd just read the how-to guides as the no-bs higher level explanations because you know anything higher level is often bs
- how do I get something done with this? - API docs combined with dense explanations combined with dense usage docs (I don't want to have to piece them together myself through my own effort, I want them all in one place)
- what is the reasoning for it being implemented/architected/etc. this way? - design docs + "philosophy" section of docs + architecture diagrams
# using heapsort instead of quicksort
# because we care about the worst case
heapsort(data)
Code alone cannot tell the reader why heapsort is chosen, nor whether that's how the code should be. And tests are unlikely to capture this either.Code can't lie (although it can be misleading with improper naming or bad abstractions).
Comments can be completely false (although they're easier to catch because the comments sit alongside the code)
Design docs almost always lies and it's hard to catch the lies because a design doc written 2 months ago is never going to match the resulting code perfectly. In my experience, people rarely update documentation unless there's a dedicated employee/developer advocate who is helping enforce and drive the culture.
That depends how you're writing design docs. If a design doc contains arbitrary flowchars about what happens in the code, it will change.
But if you come up with designs similar to design patterns, you can document the pattern once, and use it in your code the way you use a design pattern.
Concepts that span a large codebase is also enormously paintful to figure out from reading the code. For instance object models / microservice architectures, complex state machines or security frameworks are hard to catch implicitly from reading the code.
But when you too lazy to update design docs, code and docs get out of sync. And it's always easier to just leave docs unchanged, because "hey, code works".
It's a bit similar to word processor compatibility issues on linux and windows. You can have perfect implementation of specification on linux, but Microsoft does some changes that are not specification compliant and now a doc created on windows does not open on linux. And from users perspective it's linux fault that file can't be opened, even though it's actually microsoft that is not following specs.
We're talking about a line of code that has 2 words here. "A method to sort something" and "something to sort".
I agree with you but in my experience unless you are writing this documentation in a regulated industry (with external incentives to keep the documentation up to date) the technical docs will eventually go stale or out-of-sync with the actual implementation.
I've worked on code older than I am and those little comments that probably should have been design docs were a godsend. Nothing else had survived.
Describing what a single code is doing is only useful, if the comment is easier to read than the code itself. That can be useful for arcane cases, such as dirty regular expressions, bash scripts with several layers of pipes, or non-standard pointer manipulation in C.
Much more important, is the why, as in your example.
But there is also the what when it applies to code design that spans more than a few lines of code. This requires documentation. As long as we follow a design pattern, the documentation is covered by the pattern description.
But if we invent a novel design, either a one-off or a new design pattern, it needs to be described in text and figures, not just inline in the code.
Similarly, and even to some extent when staying inside a design pattern, object models need a conseptual design, where the ideas and purposes behind each class or interface is explained, as well as where the relationships between classes are visualised.
I still maintain we're awaiting another generation of version control where commentary on code is cumulative, instead of just set at the time of commit.
import heapsort as worstcase_nlogn_sort
worstcase_nlogn_sort(data)For example, you could express the same information that the comment in your example provides in the following manner with self-documenting code.
enum PerformancePrioritization {
FastWorstCase,
FastAverageCase
}
function sort(data: List<TypeOfData>, performancePrioritization: PerformancePrioritization) {
if (performancePrioritization == FastWorstCase) {
heapsort(data);
} else {
quicksort(data);
}
}
...
sort(data, FastWorstCase);But a sibling comment [0] did provide an example that would give a much simpler approach to accomplishing the same thing in this particular instance, with zero extra lines and runtime complexity.
/*------------------------------------------------------------------------
; For the RR OPT, the class field is actually the size of the UDP payload,
; and the TTL field are a bunch of flags. We have a separate field for
; the UDP payload size, so here we make the adjustment behind the scenes
; so you don't have to know this crap.
;-------------------------------------------------------------------------*/
if (answer->generic.type == RR_OPT)
{
answer->opt.class = answer->opt.udp_payload;
answer->opt.ttl = ((data->rcode >> 4) & 0xFF) << 24
| (answer->opt.version & 0xFF) << 16
| (answer->opt.fdo ? 0x80 : 0x00) << 8
| htons(answer->opt.z & 0xFFFF)
;
}
I'd like to see how I could do away with the comment and make this more "self-documenting".[1] Code to encode/decode DNS packets: https://github.com/spc476/SPCDNS/blob/05aead581acf050edf610c...
I have been picking at some thoughts on the ~naming side of this recently... https://t-ravis.com/post/doc/what_functions_and_why_function...
Problem with documenting in the code itself: The "Why" is further from the code than the "What", and the "How" is further away still.
Explaining "What" a function does is close to the code, I just put some comment lines before the function/type, etc. and maybe some into the definition. They are just to make the code better readable for the human.
But where do I put the "Why"? I could put it with the "What", but often the "Why" depends on somewhere completely different in my code (or someone elses code). eg. "Why am I using a map here instead of a list? Because the remote API expexts a JSON object and the map is easier to Marshal into that".
But what if the remote API changes? Okay, I change the code that deals with it, and update the comments for that code. Do I also change the comments everywhere else where this map is used? For that matter, is it explained anywhere else in the code why that is a map instead of a list? The "What" says "this code iterates through the maps values and appends an underscore to each", but is the "Why", which explains why that is a map there as well? In every function that works on that map?
My point is, not only is it difficult to explain everything in the code, it's also very very difficult to keep this documentation up to date, because the relationship between a comments location, and the code it belongs to, becomes more complicated the further up we go from the "What".
Code absolutely can be written in a self documenting way, just that most people don't structure their code like English/natural language. Also some languages have too much syntax bloat to be written effectively in a self documenting way.
I agree with the principle that some things are better expressed through documentation, such as motivations and higher level diagramming. But the average code base is so far from being self documenting, that the gap there should be closed first before moving on to external documentation.
External documentation creates a second source of truth that often falls out of sync with the code, so becomes a form of technical debt/maintenance burden.
The best way would be exporting docs from your codebase, and have all documentation related artifacts expressed alongside the code. Diagrams expressed in a programmatic DSL etc
I LOVE Golang's approach to this [1]:
- Documentation generation from code is built in to the language toolkit
- There is a standard for writing the comments so all of them look the same across every project
- It even supports code examples which are visible as a REPL in the generated docs
When people ask "what's so great about Go?". It's stuff like that which is hard to succinctly describe but makes a huge difference in overall quality of Go code in the wild.
Many of the Go features are only a win when the competition is C.
More to the point, just because some platforms provide equivalent functionality doesn’t make the GP wrong.
In what way would the JavaDoc @param/@return/etc annotations actually help? What Godoc has is more than suitable for the vast majority of cases.
That crap allows for tooling that Go will never have thanks community culture.
Keep building strawmen, I'll be building software.
“That crap” is at least in part a reaction to the endless stream of DocTagFactoryFactoryFactories patterns that came out of the Java world. While I think Java is a decent language with a good ecosystem, I’m completely over the idea that the only way to do things is the Java Enterprise(tm) way.
What is on that documentation site you linked is just a text formatting guide, it's absolutely not enough to get a powerful documentation system. But yeah, it's simple.
What kind of function documentation are you reading inline lol? Does it have like 30 function arguments?
With some error handling code, for example, briefly commenting what conditions lead to what exceptions goes a long way.
But of course it depends on the context and what you're trying to do.
React/Vue is a good example. You have some Application component that has a set of sub components representing different regions of the page, which have their own subcomponents and so on.
Makes it very easy to dive in from top down and drop into lower level abstractions as needed.
It's like walking through the forest and wondering where you are, but every tree has a sign on it that says "tree".
The main problem with comments is they make code less readable because they interrupt the flow of code-reading and thus code-understanding.
But now, I write all my multiline comments in this style:
/*-------------------------------------------------
This is a comment
*/
In the IDE I am using (WebStorm) there are keyboard shortcuts for collapsing and expanding sections of code. For comment-sections like above it collapses them into:/*--------------------------------------------- ... */
What this does is it turns a code-reading distraction (the original multi-line comment) into .. a SINGLE LINE SEPARATOR.
Single-line separators do not negatively disrupt the flow of reading the code. In fact they actually help code-reading by dividing the program into semantically coherent code-sections separated by what looks like a simple horizontal line.
I'm no longer very concerned about out-of-date comments, I accept them as fact of life. They are ok as long as they don't impact my code-reading negatively. I don't need to ponder: Is this comment too long? It is not when it is collapsed.
And comments will still contain much useful information. If I don't understand why some section of code is there or is written the way it is, I can expand the comments around it and maybe I find some helpful information there. Often I do.
I do wish the IDE would offer more support for this style of commenting. It would be great if all multi-line comments would be collapsed by default. I haven't figured out if that is possible in my IDE. I do know how to "collapse all sections" which is what I frequently do. It is a good way to start reading a code-file anyway.
So, the trick is really simple: Start every multi-line comment with a line that consists of a line of dashes.
import random
comments = [' ',' ',' #this is a loop',' #is this tehcnically a quine?',' #small optimization']
code = ['import random','comments = ','code = ','code[2] = code[2] + str(code)','code[1] = code[1] + str(comments)','for i in code: print(i + random.choice(comments))']
code[2] = code[2] + str(code)
code[1] = code[1] + str(comments)
for i in code: print(i + random.choice(comments))Interestingly, two of the comments are descriptive (answering "What"). The one asking whether this is a quine is in fact a type of communication widespread in code, when authors use comments for communication.
Sure. Obviously. But if we don't neglect the comments, is it technically a quine?
If one program is different from another in a way that no one can discern by looking at the output or runtime, is it actually different? Are comments part of the code per se? Does "don't interpret this text as commands" count as a command?
I started my career in a pretty low-comment Ruby codebase, but now count myself as a person partial to adding lengthy 'why not how' commentary in programs.
I think the triangle metaphor muddies the water a bit.
But... I agree with most everything else the author says. Context is important and is not present in the code (usually.) The developer's intent is generally absent from the code and should be supplied somehow.
I might even go further to explicitly mention documentation at the top of a function traditionally describes that function's context in the greater scope of the application or module or class while documentation inside the code body often describes the programmer's intention with respect to the function's implementation. (This is just an observation from code I think is well documented. Maybe there are counter-examples out there.)
I have a bit of an architectural bent, so I might have talked about documentation and it's relation to interface/implementation and problem domain and solution domain. Which is to say, documentation should be clear about which domain you're documenting.
Though I don't understand why the triangle is important, props to the original author for saying "no. code is not self documenting."
See, OOP worldview is a priori based on the idea that all programs can be written as collections of small, independently reusable, objects. In that world, there is no place to put any high-level explanation of how the parts interact as a whole, or what are any emergent properties of the system, because it is (on the ideological assumption of object independence) not needed.
It's very visible from e.g. Java package system design, there is no place to document what a particular package is for, or what classes are part of its public API and what are internal. And in Java in general, every entity must be assigned to some object, the things that lie outside (and can be shared) are frowned upon.
Java has had a well-defined place for package-level documentation for close to 20 years, since 1.5. It goes in package-info.java (docs: https://docs.oracle.com/javase/6/docs/technotes/tools/solari...).
As for documenting which classes are part of the public API, you've been able to make classes package-private since Java's initial release. The module system introduced in Java 9 in 2017 takes that further; you can explicitly control which classes are exported to consumers.
"Everything must be an object" is still more or less as true as it ever was, though.
Yes, the module system is an improvement. Java is becoming more pragmatic as the ideological dominance of OOP is dying out.
> Failing to recognise this is the number one leading cause of new users and junior devs being unable to do 'this simple task' (don't quote me on this, I have no data).
This. At all three startups I've worked at, founding/senior devs err on the side of assuming that any process not involving writing code is trivial and self-explanatory when in fact those things are some of the biggest time-eaters. Big organizational wins can be had by documenting them well.
It's extremely straightforward to read... once you understand the requirements and limitations, and some intricacies of Excel format.
So there's a 5-page accompanying document that explains those.
Code is very, very rarely self-documenting on its own. Because more often than not you need to know most of the code, most of the context and history of the code, and the business domain this code is intended for to understand it.
I assume it means: What does the code do, why does it do that, and how do we leverage it.
These perspective issues are tricky. There's a difference between "what does the code do" (low level) and "what does the program/command/API do" (high level). There's a difference between why we need the code to do x, and why the code does x in one way and not another.
Done correctly comments should be redundant.
That doesn’t mean actual documentation isn’t useful, but it only captures a point in time, rather than the actual living code.
As an aside, I think the article has what & how swapped. The code tells you how, but not what (is the correct thing to do).
That's what we do in my current company, but I vastly preferred what we used to do in my previous company, which was:
> Why (commit description)
which prevents comments obsolescence and code dirt
In this company noone cares and we even have squash & merge mode with description truncation, but methods comments can be longer than methods themselves.
It's a question of team mindset I guess
IME it's not like that at all, commits like "shifted all indentation per linter" are not so common. It's also easy to skip them.
> Useful "Why" comments
"Why" comments are not so useful to me, I guess that is the point. The add noise in the code that can be moved elsewhere.
Consider that I'm not talking about specific implementation choices comments, like
# using heapsort instead of quicksort
# because we care about the worst case
heapsort(data)
which is perfectly fine to be in the code, I'm talking about "why" details that may be not present in the Jira/Whatever issue because they are related to e.g. why a bug raised up from a technical point of view.Why?
It's faster, but more importantly the documentation says what should happen. The code says what does.