Minimalism: Practical Guide to Writing Less Code (2002) [pdf]
two-sdg.demon.co.uk
two-sdg.demon.co.uk
I partially disagree. Some of the worst code I’ve seen was an attempt to reduce duplication.
It’s easy, especially in UI code, to mistakenly identify duplicate code, and prematurely build a DRY “solution”. Premature DRY is the root of much evil.
I try to follow the rule of three before trying to identify duplication. Even then, I’ve become much more cautious than in my younger days.
Duplication for < n is good in principle until you realize some coworker decided to make their own system or went way past n with end result being it'll take more effort to refactor than to have initially designed something robust.
For example, suppose you have two promise based functions with error checking; one of which is doing a fetch and the other is performing a complex calculation. The error checking blocks might be the same since they are mechanically performing the same tasks; however, since the errors they are likely to see are so different a generic version could make the code more difficult to understand.
I certainly wouldn't want to wait to 3 if I am copying a block of 10+ lines that requires just a single int parameter to abstract into a function very local to the 2 uses. And I may use "the rule of never" on a 2 liner that requires a callback to abstract.
I'd claim that _on average_ DRY-ing is easier than unDRY-ing. This assumes the repeating code is intentionally kept simple and linear.
The absolute worst code is a mix of repeating code and half-assed abstractions, where the "perfect refactoring" would have to "move" or "redo" abstractions. This to me is even harder than DRY or unDRY.
For me personally, this was one of the biggest lessons I learned that made me see large codebases and working in a team in a new light.
And there are definitely cases of duplication too small to matter - which as you note is often the case for UI and UI-like things. You often end up in a position where the most reusable code is, in fact, the copy-pasted code, because the only thing it does is describe a permutation of abstractions and assets glued together. Sometimes this code is not the optimal code, but that's a problem that can be returned to down the line.
The easy picks for DRY are one-liner functions that describe a preconfigured intent, like "draw a centered bounding box". That's something your glue code will crave, and there's little issue with deduplicating it later. But even there some tension arises since you can always decouple a little further by having those functions rely on an imperative context where most of the state(e.g. the size or color of the box) was previously configured, versus specifying it at the callsite. After certain thresholds, there's a flipflop between wanting to code it and wanting to configure it.
Likewise with continue and break. I think this advice is very bad. All the other recommendations seemed fine.
Take this contrived example:
function foo(bar)
{
var result, error;
if(bar is null)
{
error = "{NameOf(bar)} is required";
result = null;
}
if(bar is not null)
{
... Main function logic ...
}
return Tuple(result, error);
}
If you can return multiple times instead you'd be able to write: function foo(bar)
{
var result, error;
if(bar is null)
{
return Tuple(null, "{NameOf(bar)} is required");
}
... Main function logic ...
return Tuple(result, error);
}
Assume "Main function logic" has several of nest of its own (branching, loops, etc) it quickly gets difficult to read. success = doOneThing()
&& doSecondThing()
&& doThirdThing()
...
&& doLastThing();
and the more common "traditional" way: success = doOneThing();
if(success)
{
success = doSecondThing();
}
if(success)
{
success = doThirdThing();
}
...
if(success)
{
success = doLastThing();
} Promise.start()
.then(doOnething)
.then(doSecondThing)
.then(doThirdThing)
.then(doLastThing)
.catch(handlesomeerrors)Instead of loading up every language with stupid keywords that are meant to simplify monadic code (but only for one domain), language designers should really take a page out of Haskell or Scala's book and think seriously about unifying on a do notation/for comprehension type construct.
That way regardless of if you're working with promises or not, you don't have to worry about hacking it with && or nesting 10 layers deep with if/else.
Having said that, I'm with the "guard statements" crowd, too, but it would probably be better if that were handled by other parts of the language syntax. Contracts, for example.
The slide refers to the inversion of control principle but can easily include other important aspects of including only what is needed.
I would like to expand on this for discussion. Consider the challenges of not over-engineering solutions to problems you and others don't fully understand, yet. Unless you work for NASA, you won't fully know what you need until you need it.
How do programmers here account for the unexpected?
"Always implement things when you actually need them, never when you just foresee that you need them"
https://en.wikipedia.org/wiki/You_aren%27t_gonna_need_it
EDIT: The way I think about this is that software development has a very large degree of unknown unknowns. More than maybe other domains like architecture/chemistry/...
I'd add that, while you shouldn't change the complexity of your implementation very much, you should look at potential future scenarios and migration paths. Sometimes you can make small changes to the plan that make migration paths easier. Optimize for ease of change I suppose.
Obviously the effort you put into this should be proportional to the complexity/importance of the problem.
This works fine for known future changes, say if you’re implementing X, and know that a similar Y and Z are down the road.
Beyond that it leads to the usual problems of crystal ball based planning.
As usual, in engineering, needs create capabilities. It is very expensive and risky to go fully determining your needs before execution, but they pay for it.
Not all feedback is good, not all of it is applicable, requests for marginal features that benefit only few are exceedingly common. All feedback should be taken strictly in advisory capacity and assessed critically. However if you did miss something useful or dearly needed, it will surface almost immediately.
We've been practicing this for ages and it works extremely well.
You don't. At least, I don't.
When I'm writing code for an application, I'm doing the following...
"Hmmm, I need a function to perform <certain thing>"
Function is written... tested... debugged until <certain thing> is achieved. Job done.
Whilst writing said code, I do not think to myself "Hmmm... y'know in the future it might need to also do <some future thing>" and proceed to add additional function parameters and write additional code which would only be useful sometime in the future, just in case.
No. I write sufficient code for what is required of said function at that moment in time. It is only if and when required would I then add any additional parameters and code to that function. Or write new functions.
So that is what I try to teach junior developers today: Only write what is required now and don't think too much about what the function might be expanded to do in the future.
Those are my trigger words. My experience is that about 97% of the time somebody says those words (in the past it's been me) in the context of writing software and it turns out that:
* <some future thing> is never required.
* <some future thing> is required but it ends up being needed in a completely different form, rendering the preparation work pointless.
* <some future thing> was nice to have but it ended up so far down the list of priorities that it might as well not have been.
This applies to every level of the stack from one off functions to grand features, to refactoring, testing and even to things that many people consider "best practice".
I don't think it pays to be an extremist about many things in software but "if there isn't a glaring need for it then don't implement it" is one of them.
So, yeah, DON'T just add arbitrary stuff because you'll think it will be useful to someone; but DO add stuff to maintain the symmetry and understandable abstraction of what you've made.
Including/writing a fully featured Vector library when you need a Vector library is perfectly fine; but if you're writing some code to handle those Vectors you should stop after it's done and not generalise it to also handle a bunch of other types in case you might need it later (because you probably won't).
Honestly, it sounds like you're hourly, and not at all related to what the submission is talking about...
I'll add that sometimes the more general solution also happens to be the simpler one, and in that case you should always take the simpler one even if it's more general, because then you have both the advantage of simpler code now and flexibility for future change.
To put it succinctly, "increase generality only when it decreases complexity."
I work on systems that need to run reliably. However, they depend on data collectors (like sensor data or web scraping) that are not known for their reliability or scalability.
We do lots of retries, log errors and failures, and patch problems until things work as expected. One thing that helps is the use of sentry.io . It alerts us about exceptions and detects regressions/regression fixes by checking if there were any commits that might have caused or fixed issues.
This is where good architecture and design come into play.
Design for what you know and need now but use good principles (OO, etc.) so that code can be modified in a sane way.
The situation to avoid is to pay a complexity tech debt now, for something that may not occur for some time, possibly never with all the maintenance of current features upon that complexity.
A common reoccurrence is generalized methods. There was a PATCH endpoint for main entity which I'm sure started out innocently enough with maybe just two change sets. Following the intents through the layers of this choke point implementation is unpleasant. We could achieve the same by composing variants at a higher level so that each concern is separated from the others, but then we introduce this 'machinery' that may not be warranted. If there are many or a dynamic components then it may be the best solution but not until.
I recently fell for this myself. On recognizing that a particular service was just a combination of two state machines, I abstracted the state machine operation and applied it to each subproblem. It worked like I expected but the result wasn't right. The state machine machinery was more visible than the core logic. I eliminated the concrete abstraction and created two implicit state machines that invisibly did what was required for each subproblem. If we already had a common state machine form that we were familiar with, things could be different but not when introducing it into a codebase for the first time.
It is fine to write towards future problems, but all to often trying to be generic explodes the codebase and leads to bugs.
There is a tension between dependency and redundancy. Avoiding redundancy often introduces dependency as code becomes shared for several purposes. Knowing which is preferable (dependency or redundancy) in a given case requires a lot of judgement and changes as the code evolves.
Redundancy is bad when it introduces too much code and too many chances for error or unnecessary divergence of implementation.
Dependency is bad when two things that start out closely related diverge in purpose, straining an implementation to handle both cases.
Also, previous discussion: https://news.ycombinator.com/item?id=5024221
Without fail every suckless person I've talked to either has no idea what they're talking about, or in the off chance they do they're incredibly shitty far right wing types who gripe on about women/minorities/CoCs ruining tech.
Having said that, my general impression was that there was an inordinate amount of Plan9 cargo-culting (with some DJB fanboys mixed in). Young, idealistic people without much experience but lots of admiration for the more Bauhaus part of their elders.
I used to make this same argument, and then took it overboard by never commenting. After reading some literate codebases, I’ve changed my mind.
Code is (almost?) always obvious to you as you are writing it. So, you never recognize code that is in need of a comment until you come back to it and struggle to understand its purpose and the context that gave it birth.
I now follow a rule: a meaningful comment at the top of each file. A comment on every exported function / value. Often, in the process of writing the comment, I realize I’ve poorly named something or that I’ve failed to handle some case. Comments help guide the reader, but also the writer.
Example: https://github.com/Convery/Desktop_cpp/blob/VersionDD/Source...
In a team setting, comments like these are even more important, as it gives context and higher level meaning to the code without forcing you to jump all over the place through small functions (which would be an alternative way to make this self documenting.
I strongly disagree specifically about comments like these throughout an active code base. A well named variable or method can act as a much better descriptor of what’s happening and doesn’t have the same maintenance cost.
I think we both agree that comments are valuable, just not the scope. Comments are valuable when you’re doing something unexpected or where the code fails to explain what’s happening.
For what it’s worth I mostly work with Ruby, JavaScript, and TypeScript which definitely color my views.
Well, it can be both. Less code is easier to maintain than more code, and that applies to comments too. Finding the "goldilocks" point is the challenge of much code design, and it applies to comments too. There's such a thing as both more comments than you need (increasing maintenance cost and risk of outdated comments without providing enough value to justify), as well as less comments than you need (to decrease cost of understanding the codebase).
// Have the window rendering in the highest allowed state.
I'd argue that comments like these are fluff. It says basically the same as what the code does but doesn't say why it does that.In general I find comments like that suffer from bit rot faster than anything else (perhaps with the exception of commented out code) as it's tempting to leave the comment changes until you've got something working. And a wrong comment can lead you down the wrong path very quickly.
(0. The debugging approach can depend on a person's mental model)
1. Readability - even for a beginner, the method name should give a fair indication of what the code does. If the method name doesn't give a good indication, then the name should be changed. Otherwise, the comment is redundant.
2. Staleness - imagine a commit which changes the behavior of Engine::Compositing::onFrame(Deltatime) to only notify some of the components about the new frame (comment currently states "// Notify all components about the new frame"). Presumably, this commit should only change the internals of the onFrame method. However, because of a redundant comment, the person who changed the onFrame method now also has to find all its invocations and possibly update many comments.
My personal preference is commenting only domain-specific code chunks (example: [1]). These code chunks are usually put as close to the implementation as possible.
This way, whenever someone wishes to change the code, the person (1) immediately notices reasons behind the implementation (and may decide against modifying the code if he/she was unaware of these details) and (2) can modify the comment right away in case domain-specifics have changed in the meantime.
I'm curious to hear your or someone else's opinion against this argument?
[1] https://github.com/tkukurin/lesshint-intellij-plugin/blob/ma...
It depends on who you're trying to reach. Since the GP wants his open-source commits to be "learning opportunities for others," it might make sense to comment in that level of detail.
That particular comment does contain semantic information which is not expressed in the code: that the thread is being prioritized for rendering purposes. It's possible you could communicate that through variable names, though.
Generally, comments should focus on the "why", not the "what".
I think what is intended in general are comments in the form of:
// Start the loop.
while ( have_posts() ) : the_post();
[..]
// End the loop.
endwhile;
And yes, that is an actual example from a real codebase.This is an extreme example, but quite a few of comments I see are very redundant. Sometimes comments are useful to split stuff up, but stuff like "start loop", yeh nah. It's just a random stray comment, probably as an artefact from when the author was gathering their thoughts and/or took a break.
Even worse are comments that are just confusing, or don't match the code.
I've done some literate programming, and I find it interesting for certain types of programs and/or if you're struggling with something. But I think that for a lot of stuff it's redundant.
So often when I see this I wish people would have used another function. Its hard to argue that a function which requires splitting things up using comments is "doing one thing".
Way too many people here are confusing comments with documentation.
A great approach to this is Python's docstrings, which parse to the naked eye as comments but can be rendered into proper documentation.
In my highschool introduction to programing class, we were for a long time required to comment every line of code. Examples given by the teacher were all this type of comment.
I've no doubt it put the wrong idea into some students' heads.
When one of these rules is "thou shall comment all declarations", it results in the developers polluting the code base with verbose drivel, to avoid the review-loop.
// FooBar - Foos the Bar
// @param baz The baz.
// @param quux The quux."Foo is the second argument and should be a string
Bar is the second..."
It blew my mind the first time I encountered it. A group of adults who just... gave up to the process.
In the same way, I think the problem with "prefer code to comments" is that people often miss that it actually requires even greater focus on communicating your intention to others (e.g. via methods like declarative programming and meaningful names).
// OLD CODE WITH COMMENTS:
// check ...
4 lines of code
// drop ...
4 lines of code
// read ...
4 lines of code
// do xxx ...
4 lines of code
In total, that method was like 16 lines of code, with some short commented sections. And the sections used variables from before.Situation: New coder comes in, sees comments. Says "comments are bad, they get out of date, yadda yadda". So proceeds to change comment names into method names:
// PROPOSED: COMMENTS => METHODS
check()
drop()
read()
doxxx()
What coder forgot was the variables and context. As said before, the total was 16 lines of code with simple variables used between it. Not a big deal. But now when it has to be split into functions, it became like this: // ACTUAL 1: COMMENTS => METHODS
a, b, c = check()
d, e = drop(b)
f, g = read(a, d)
i = doxxx(e, g)
And this was a language that didn't support multiple return types. So the code was actually: // ACTUAL 2: COMMENTS => METHODS
class X { a, b, c}
class Y { d, e }
class Z { f, g }
X x = check()
Y y = drop(x.b)
Z z = read(x.a, y.d)
i = doxxx(y.e, z.g)
So now coder added 3 more classes, with more lines of boilerplate. Then the coder decided, to manage this problem better. So they created some more interfaces and did some DI. The result was around a 250 line PR, which I cannot ping here anymore. When we pointed out the same logic was now 250 lines, the answers were: "It's inherent complexity which we didn't know how to manage. His/her solution scales better, and he/she has shown us the way".All for what? Because comments were considered a smell. Go figure.
I guess this is why this other HN post trended – "Please do not simplify this code": https://news.ycombinator.com/item?id=18772873
Because the original code, while preferable to the version with 3 classes, is not good either. Indeed, as this makes clear:
// ACTUAL 1: COMMENTS => METHODS
a, b, c = wait()
d, e = drop(b)
f, g = read(a, d)
i = execute(e, g)
There are complex dependencies among the subroutines which can probably be simplified with refactoring, and which should be explicitly called out, but are allowed to "hide" in the original code where, locally, we have 7 global variables shared by 4 inline subroutines -- not a good situation.> There are complex dependencies among the subroutines which can probably be simplified with refactoring, and which should be explicitly called out, but are allowed to "hide" in the original code where, locally, we have 7 global variables shared by 4 inline subroutines -- not a good situation.
What is locally global supposed to mean? This piece of code can be written at least two ways: as a single function with comments or as a function calling four other functions. Just because splitting it up into multiple functions a certain way requires multiple returns or long argument lists doesn't mean that the original code is necessarily bad; It could be that the splitting points were chosen badly or that this code is simply better off as a single function. These "complex dependencies" you're worried about form a DAG in the example which seems simple enough to me.
Anyways, most of this is moot without real code.
Agreed, I hesitated answering because of this, but oh well.
> What is locally global supposed to mean?
It means within this local context, we have 7 global variables. Just because they aren't global to the entire program doesn't make them magically exempt from the problems of shared data. Specifically, it's too difficult to reason about logic and control flow, even in this small example. You cannot have strong confidence the code is bug-free just by inspecting it.
It also depends on logic reuse. If you are now able to use the `check` method in other places, then I say it's overall a successful refactor.
How short? 100 lines? 10 lines? 1 line? "It depends" => i.e. discuss forever in PRs towards a one-way "shorter and shorter" ticket?
More and more coders who join the project say the same thing, and reduce methods shorter and shorter and shorter, until it's dozens of classes with one method each having small lines like "twiddle-doo" and "fiddle-foo". Of course, now the method is easy to memorize, and supposedly it's "clean code" now. But the functional value and flow is just completely screwed up and can never fit in the head of anyone.
I think that's true if "short" is measured in terms of internal complexity. I think each function should try to manage a small bite size amount of complexity that represents a reasonable level of cognitive load for a brand new person to reverse engineer. It could be quite a lot of lines of code if they are very simple ones, or it could be a one liner if it's super complex / obtuse.
> Code is (almost?) always obvious to you as you are writing it.
I try hard to go back and reread code as I'm working on it, several times a day, as if I'm someone else, seeing it for the first time. It may help to pretend I'm showing it to a particular coworker.
It's a certain kind of empathy skill that it takes work to develop, but it's very useful both for coding and any other writing.
Perhaps the most common problem is what I call "assumed context". That is, you're so deep into the task that you take the major parts of it for granted and it never occurs to you to explain the most basic and fundamental parts of what you're doing. Instead you explain minor details.
Comments are failures on the part of the programmer to describe the implementation clearly in the language being used. The lesson you should have learned is to recognize that it's okay to fail in this way, but to still consider it a failure.
In a lot of cases I think it tends to be the opposite - adding more code just for the sake of 'explanation' bloats the code to much larger lengths and ends up making it harder to understand than if you just added a few lines of comments that explained the exact same thing. In that situation, adding comments is not a 'failure', it's just another option for making your code easy to understand.
In fact, it was reading those projects and then returning to my workplace’s uncommented source that convinced me to change my philosophy on commenting.
[0] https://sqlite.org/src/fdiff?v1=7c288b4ce309b5a8&v2=8efa2c81...
And Redis makes no such requirement. You are wrong.
If they can't understand why you wrote a certain code block or do what it does, then either re-write it or explain in a comment.
When I'm reading some API, then of course I want a detailed comment for each function and value, including valid/invalid parameters, maybe even some runtime behavior stuff if it's important (like whether some async method may contain some cpu-bound parts) any definitely anything unexpected or deviating from conventions.
The conventions part being important because usually I don't have the time to read all the comments and all the documentation, but instead I'm looking for some simple concepts I can learn once and then use to understand most things just by reading the identifiers.
If the project is written in a consistent way, it can be pretty large but require very few comments to be understandable. And anyone modifying it will need to know some set of rules for it anyway, so they should be made explicit.
(5) Manage Resources Symmetrically (10) Let the Code Make the Decisions
What does it look like when you do the opposite of these rules?
I think I'm unfamiliar with the problems these slides address, so I can't grasp the insight in them.