How to Write Readable Code
jeremymikkola.com
jeremymikkola.com
I used to work at a place where they had a substantial amount of legacy, they kept two systems up to date, a current one and a new one. In my time working there they still didn't phase out the old system and the new one was already a system full of legacy and they were planning a new system (I wonder if they're running three systems now).
Many of their functions had +10 arguments which were poorly documented so they had to be deciphered from the function bodies who were sometimes thousands of lines. I was fortunate enough to be granted my own project so I didn't spend too much time working with it, but when I did I had no idea what I was doing.
I was very happy when I left, they practically begged me to stay but I just simply couldn't deal with it (and there were plenty of other issues as well). It was very effective at teaching me ways of how you can bungle up every aspect of your IT.
And to top it off we did not have a versioning system or test environment, we did everything live, right there and then.
Just out of more curiosity, can you comment on the idiosyncratic code style? I saw in your other post that it was written in PHP 4. So that means it did have access to classes and whatnot, right?
Would not recommend.
Their main problem was that most of the people they hired took 6 months to become proficient in the language and then they quit because they hated it so much.
I actually managed to last a couple years without ever learning the language, by bouncing into a newly created research division and then getting put on a big deathmarch project that did not use the language to implement a governmental specification - but I did write a critique of some of the of the security issues / validation issues about the language in my first month there.
It was awful.
Apply the "Principle of least astonishment" aggressively in all aspects, variable/function/argument/class naming, arguments/mutation, return type variations. Make it so reading the call site of the function lets the reader accurately guess what it does and not need to look into how it does it, or if it does something else too/sometimes.
All of that is useful but doesn't help if the overall structure of decomposition from top to bottom is poorly chosen: reflect and reconsider. "Perfection is achieved, not when there is nothing more to add, but when there is nothing left to take away."
Specifically, I always have in mind that a function should "only do one thing", so having multiple layers of abstraction in the same function would be picked up by that rule generally, but I like the explanation given with "if a welcome email has not already been sent, send a welcome email". This approach also makes it easier to test parts of code in isolation.
Whose time/resources/energy? Not the software vendor's.
>most programmers are incompetent.
You can't afford the software written exclusively by very competent programmers, optimizing for efficiency/speed/energy usage instead of development time. That'll usually cost you far more than more compute resources/time in the end.
Also, paying an army of incompetents to modify a codebase isn’t a sustainable development process.
Eventually the code spins out of control. Worse, one of the incompetents eventually accidentally gets promoted, then rapidly hires and promotes more bozos. Ultimately, they eat the organization from the inside like a cancer. (This is well-documented, and has happened to many organizations. Search for “bozo effect” or “bozo explosion” for more information.)
Unfortunately it doesn't have to be sustained long term. Only until you get bought or whatever, the people at the top make their money, the army of incompetents cash in somewhat too if they are lucky, and some poor sucker is stuck with an unmaintainable mess that they just paid a premium for.
It seems therefore natural, that as hardware becomes stronger, developers feel less pressure to optimize. Naturally also, it is the developers fault when something is so poorly optimized that it affects the product. Because in the end, we are workers like all other kinds of workers, and we're looking to get a job done, and not done "perfectly". (unless you're hired to optimize servers at Google to the point of perfection)
Also, a term I like to mention is that "instant = instant * 0.5". Meaning, if your code is fast enough that it doesn't affect the user experience in the end, and feels "fast". Making it twice as fast has literally no impact. And I doubt server computing cycles is causing energy shortages around the world. We've got other problems causing that.
that makes the very flawed assumption that your program is the only one running on the computer. In practice even if your program does not appear slower, it uses more battery, other software have less CPU time, etc etc
Then it often becomes prohibitively expensive to just throw more servers/VMs at it.
And often the “expensive” engineer changing a nested loop into a map lookup can save years of compute time in a matter of hours.
Unfortunately all too often that only happens after scaling up/out without understanding why systems are slow.
Don’t get me wrong, I’m not arguing for needless early optimisation but an inefficient query/algorithm can waste a lot of computing resources...
The slowest code I’ve encountered was also unreadable and unmaintainable (otherwise someone would have already fixed the obvious performance issues).
The most significant efficiency improvements typically come from better algorithms or data structures in specific cases. Sometimes improvements to heuristics too. Lesser improvements may come from optimising hot code (eg making it better for the cache) or fundamental language changes (e.g. in a more statically typed language, field accesses may be pointer dereferences, or even better pointer addition, rather than hashtable lookups. But languages like python probably have better tuned general-purpose hashtable implementations than you’ll likely find elsewhere. Also changing default data structures from e.g. linked lists and binary trees to arrays (or something array like) and b-trees (or some other shallow tree) will likely be good)
On the surface that looks like clarity, but when do you have to look at that code? Often when you need to fix an issue or extend a functionality. You need to understand what's going on in detail.
When you debug a workflow or Algorithm in a code base like this you have to jump from one function to the next and suddenly you are 6 layers deep and totally lost the context. It is so much harder to grasp code structured like this compared to simple linear flows where each step is marked by inline comments.
This was in Java which is not even my primary language.
It was pretty impressive.
It breaks down to basically, when you can keep the whole system in your head, that style is needlessly complex, but once it exceeds that level, then it is better.
Extracting functions does way more than removing descriptive comments. When you extract a function, you're compartmentalizing code and adding scopes where none existed and limit contexts. When you extract a function, you're explicitly constraining a block to comply with a contract, which allows you to not care what goes below that point. You just care about pre and post-conditions, and that is more than enough to troubleshoot and fix bugs, and more importantly not add them.
Ah, I thought this would define "readable code", but it's "how to write" it. The title, literally.
missing step: understand WTF you are doing.
Many small functions also create complexity, though hadn't considered stacktrace documentation:
> [small functions] It’s easier to tell what the program was “thinking” when you look at a stack trace or run a debugger.
[don't mix levels of abstraction] also explodes the number of functions. I've seen this reasoning before, it makes sense, but I'm not convinced yet. I think if there's some substantial, genuine work done at each level, it's helpful. But just a sequence of calls doesn't help.
Nice bit on incidental duplication:
> The point of DRY isn’t to run a manual compression process on the codebase, it’s to avoid a dependency where two parts of the code need to be manually kept in sync.
In writing, clarity and simplicity go together. Of course, it takes longer to write a short letter. And, code is not writing.
I'm uncomfortable about this because abstractions are difficult to define and demarcate - and end up being leaky anyway. So you need to change several separate functions in parallel, instead of having them in the same place.
OTOH there are also some obviously different levels of abstraction, which it makes sense not to mix.
Further, some such levels can be established in a particular domain, so whether good or not, they are helpful for people already trained in or familiar with that domain.
As I’ve grown older I’ve come to hate functions of a certain length, and especially how conditions are used. I think the authors gets at this with the “avoid configurable functions” but if I were allowed to give it a new heading I’d choose “avoid conditions.” This is in keeping with the need for the law to be via negativa. But conditions are powerful language features, you say, and I totally agree. I think they should be used for the dual purpose of (1) dispatching commands (switch/case) and (2) eliminating/refining inputs. That is, you should always return in an `if` branch.
Interesting. This code style simply doesn't need 'else' at all. I can surely say, if I was reading a codebase that adhered to this discipline, it would be very easy to read.
What are your thoughts on how difficult it is to write code in this style?
I've had seasoned team members criticizing that approach as possibly the worst mistakes a developer can do.
Their rationale is that you needlessly increase the cyclomatic complexity of a module whenever you return early, and if you care about logging you'll already have to add multiple code paths just to emit them.
What are some good looking code one could read, regardless of the language?