What does code readability mean? (2018)
typicalprogrammer.com
typicalprogrammer.com
But there is a single “case” I’d never agree on: the code whose complexity and/or composition ends with nothing. The more experience you get, the more high-level structure you are able to pick on, but sometimes structures are just there, serving nothing at all. Overzealous (or compelled) decomposition into 5-line chunks over tens of files. Pointless renaming, re-exporting and encapsulation. Higher-order code golf. All this done to a finished non-extensible project makes no sense and makes it harder to read, for no reason.
This is a lenghty article, but it is readable in one chunk, a good piece of read. Now imagine the author split it into a number of submodules, then in each they’d give a new name to every phenomenon. Then instead of using English they’d construct a new sub-language (also in modules) to express meaning in a shorter way. E.g. this:
I can’t read the code because I don’t have sufficient experience or expertise (with the language or domain).
I haven’t spent enough time trying to read and understand the code (“it’s not obvious” or “it’s not intuitive”).
I don’t have much interest in understanding this code, I prefer to rewrite it in my own style.
…
turns into this: With code as object, let me as you negating:
frobnicate causality of read object over sufficient experience or expertise (with the language or domain).
frobnicate relation of enough time and attempt to read and understand object (“it’s obvious” or “it’s intuitive”).
have much interest in understanding object, no prefer to rewrite it in my own style. // “no” negates negation
include more items from [here|https://…].
See frobnicate in my frobnication article.A 1000 line function broken down into 100 different 10 line functions which are then each called in turn (and need lots of parameters to pass around the inevitable shared state) is actually often less readable than just a 1000 line function.
The way to deal with such a function to ask why there needed to be a 1000 line function. Why does so much apparently need to be done here and now and so much state shared across so many lines.
Code metrics tools would rather you break everything down into micro-functions which if you're blindly trying to follow the tooling recommendations to get it past a check-in gate will probably end up badly named and with little thinking into actual abstraction.
If you have two functions but whenever you call function foo = Bar() you have to then call function Baz(foo) then you don't really have two functions at all.
This would help testing and understanding. You can test and understand the parts independently, and only have to test the huge function in order to check that the parts are attached correctly.
https://rhodesmill.org/brandon/talks/#hoist
https://danuker.go.ro/the-grand-unified-theory-of-software-a...
But now imagine that it returns the complete text of Shakespeare if the input it 'bob' (why? customer requirement. their business literally falls apart if this doesn't happen). In this case you probably shouldn't pull it out into it's own thing because it's so specific (and weird) that you'll want it nearby it's use case (and not available for developers to accidentally use for the non Shakespeare uses).
* yeah this is convoluted. There are better ways to factor such a function. However the point is that something can be 'pure' but still not a good candidate to be pulled out of context.
I guess it depends on how much this weird Shakespearean to_lower is used around the codebase.
If there is one spot where it's used, I would probably factor out to_lower and add a special case check for "bob" in the business logic code (outside the scope of to_lower).
If the Shakespearean to_lower gets used in various places it actually might make sense to be factored into its own function.
Do these bits of pure functions appear anywhere else in the code? If so, sure, I happily pull them out into their own functions.
If not, I will leave them exactly where they are. It's easier to read a laundry list than to look up many different parts in isolation without context.
As for testability: If some bit of code only appears in that one long laundry-list function, what do I gain by testing it outside of that context?
1000x1 lines of code all at once have a lot of space for hidden non-obvious interactions. Maybe it won't have that, but I haven't ever encountered it.
Making it 10x100 lines of code won't necessarily make it better, but at least it should isolate chunks that can be understood with little effort.
Having to understand 1000 lines of code before being reasonably certain you can make a change is far harder than understanding some 20-50 lines.
Function extraction adds indirection that doesn’t necessarily pay its own rent. It’s not free.
> but at least it should isolate chunks that can be understood with little effort.
The chunks are easier to understand. But the laundry list function that calls them isn't because it's logic is now spread over 10 different functions.
That's... the point of abstractions and programming languages. At work you don't regularly care how System.out.println() is implemented, do you?
If we follow your way of thinking through all we would have would be binary. Maybe assembly.
The whole point of having programming languages and functions is having high level descriptions of what it is doing without having to worry about the implementation unless we have to.
Now, maybe the 1000 lines of code is actually ultra specific and has no bearing at all on everything else, but that's not really common. Could also be that the 1000 lines are more data structure than code.
There is a level of abstraction that's conducive to understanding stuff. Going at higher level costs obscuring implementation details, going at lower level costs making it hard to understand what it does.
One thousand lines of code is probably way too low level. Chunking those 1000 lines into smaller abstractions is already what your brain will do when trying to make sense of it (because the cache is small for abstractions and concepts that aren't already internalized)
Someone new started and... within 10 minutes commented in the code "150 lines is way too long, we need to understand why this is so bad and how to refactor this to be good - you should never do so much in a method".
Could I break it up even more? And have one method that composes 4 then each compose 4 that each compose 3? Possibly. But there's a tradeoff, and... what annoyed me more was that there was 1) no questions asked, 2) no view of the history of how it went from 800 down to 150, 3) no review of the tests in place which document and demonstrate that it does, in fact, work.
Pursuing brevity at all costs has a higher price than some people understand. Worked with the person a while longer, and they eschewed tests because... "hey, everything I write is already so small and understandable, there's no need for tests", demonstrating a pretty basic misunderstanding of tests (imo).
I have seen 1500 method in an app I was rewriting. The dev no longer worked for us. It was badly written, I told the management it wasn't optimal and there were some issues with it but I said I don't know under what constraints the dev was working. I simply wasn't there when it was written and have no idea if it was incompetence or some other factors.
The point is anyone who thinks certain way is BAD and needs to be immediately improved while barely knowing the code base is a bit incompetent. I honestly would wait weeks to start bringing up some issues I may see in the code base or sql. I would learn the team dynamics first and try to see if there is any other reason for those things.
let (x, y) = {
... (some code that only serves to create x and y)
};
I really really like that feature for controlling scoping - it's like a single use function that doesn't involve so much navigation.Other languages allow similar constructs with blocks too. For ones that allow inline lambdas, you can also get similar results by defining and immediately calling the function.
Aside from the 1000 line function (sheesh), this is a failure of the IDE, not of the practice. We have computers that can inline these functions so you could see the code in a big 1000 line thunk if you would like, but only a LISP IDE gives you inlining afaik.
> “Good code is simple” doesn’t actually say anything.
That's a bold claim and I will assert it is wholly incorrect. Good Code is simple and simple code is VERBOSE
> whether that translates to “simple” code depends on the programmers.
That is incorrect. It largely depends on the language abstractions. When the abstractions are familiar, it goes from incomprehensible (literally) to familiar. Code cannot seem simple if you do not have the familiarity with the abstractions used.
> Should we strive to satisfy the Shakespeare for Dummies demographic
Yes.
> Programmers seem to believe in a realm of beautiful, readable, easy-to-maintain code that they haven’t seen or worked with yet
Lots of programmers believe it because they have seen it. I have. It was embedded in hard-to-maintain code, but it was there.
It doesn't mean anything because both good and simple are subject to on-the-spot redefinition. At what skill level does code become "simple"? Does all code have to be ELI5 to be good? When a Node developer looks at CUDA implementation of k-means regression analysis and finds it "difficult", does that mean the code isn't good?
When the abstractions are familiar
So you're saying it does depend on the programmer, because the programmer should be familiar with the abstractions used?
Short answer: No. "good" is a different bar (whatever you mean by it). It's not possible for every function to be simple, for more than 1 reason (for sure, some calculations are inherently complex, some relationships are complex, etc). However, the vast majority of code can be.
> So you're saying it does depend on the programmer, because the programmer should be familiar with the abstractions used?
Correct. It is expected the viewer is familiar with the language, not additional abstraction on top of that. You have to assume there is a cost to those new abstractions that elevate it beyond simple.
This can be argued the other way too. Decent IDEs have code folding and other features so you can look at a long function at a higher level without needing to abuse language features for subjective purposes.
Function inlining is folded, moreso...
foo() { return bar(); // <-- bar can be clicked on and expanded in place }
becomes:
foo() { return {return "boo";} }
------
bar() {return "boo";}
Exactly.
For decades, people were told this is the way to go, and would somehow, almost magically, make code readable, maintainable, reusable. The result is the exact opposite, a gigantic pile of needless abstractions and spread-out functionality, most of which exist solely to satisfy some paradigms.
This might seem like a minor addition to the list, but I think it's an important one. And it has a deeper philosophical implication: if you end up with a "Rube Goldberg machine" chunk of code it is not simply stylistic preference, even if the code functions correctly. In other words: you should not make the claim that just because it functions correctly that it is "correctly written". You certainly can make that claim (and I have seen many do, and I have done so myself).
There is code that is complex because it needs to be. There is code that is complex even though it doesn't need to be. If you don't actively prune the second type, it will grow all by itself and take over your garden.
Code readability = syntax that make semantics obvious.
If a block of code is written in an imperative style or functional style doesn't matter. Ask yourself instead, are the semantics clear? Are we coping or moving the values? Are we cloning the references? Are we iterating over mutable references or copies of the value?
A good language, with good code readability, makes semantics obvious. A good programmer, encourages good readability. With good readability we don't have doubts about what a piece of code is doing.
My previous answer is that good code makes it easy to discover intent.
But I think the semantics are important too - for example in scala we use monad transformers to treat futures (async code) much like lists. But the problem with this approach is it's not clear without inspecting the type signature if you're doing something that has a high algorithmic complexity eg. spawning threads or not.
I think the best language i've ever read for semantics is elixir.
But i think the lack of static typing makes it harder to code review, especially on github. Maybe Gleam is the answer?
What language do you like?
Semi-related, can you have something like Haskell that seems hard to "read" eg.
() :: ->
but written in something easier to read like C++. Idk maybe I'll get used to it like the transition from old JS functions to arrow funcs
Readability is a very real concern and something every programmer should think about. It’s a form of communication.
It all starts with naming. Naming the entities, classes, methods and variables is probably 70% of making really readable code. In my experience, the difference between ‘spaghetti mess’ and ‘nice, clean code’ can often be achieved with an identical AST structure but with well-thought-out and consistently applied names.
The other 30% probably deserves it’s own book. It involves breaking complex portions of code up into well-named segments (ex: functions or variables), keeping related segments near each other (not spreading functionality out over dozens of files), using consistent patterns and much more.
There is a such thing as unreadable code. The author here has some nuggets of wisdom buried inside a lot of chaos. Maybe his style serves a good purpose for prose, but it isn’t an example of efficient and clear communication. Readable code should be efficient and clear communication.
> 1. I can’t read the code because I don’t have sufficient experience or expertise (with the language or domain).
> 2. I haven’t spent enough time trying to read and understand the code (“it’s not obvious” or “it’s not intuitive”).
> 3. I don’t have much interest in understanding this code, I prefer to rewrite it in my own style.
> 4. The code offends my sense of aesthetics; I would write it differently.
> 5. The original programmer didn’t know how to write code.
> 6. The code appears to violate some principles or patterns I believe in.
The rest of the article mainly deals with 1-4 here, and only selectively tackles 6 as a "performative" repeat of 4.
It does mention vaguely toward the end that there may be some objective sense of readability, but doesn't go into any technical analysis. Which is a shame as I don't think it's that complex.
A few sibling commenters have alluded to this using various terms: ultimately I think objective code readability is about context and locality.
Consider a static analysis security tool that produces a control flow analysis tree: often a long tree is a very good sign of bad readability (in practice a long tree can be caused by many things: excessive boilerplate, frequent variable mutation or reassignment, etc.)
These same tools often fall down analysing control flow of applications with weird state patterns: global state, "magic" model loading, etc. Things that further hamper a reader's ability to hold a complete picture of any given file they're reading in their head at once.
I think this is interesting because not only is it a fairly simple idea to reason about (context & locality) but it's even potentially automatable as a metric.
http://literateprogramming.com
(at least the programmer would explain the choices made)
and I've found that the need to explain the code in a literate mode has resulted in my better understanding the problem and possible approaches, resulting in a successful coding solution.
What you are actually looking for is code that operates in the correct problem domain. That requires abstraction which is done correctly.
My guess is that you have worked on codebases with crappy abstractions.
I would have agreed with you when I was 5 years in but now I like small functions and lots of the. Each function have one narrow easy to test task.
I may be generalising but I find older (40+) programmers more likely to write good code - they have all the war zone experience and battle scars
For example, in Common Lisp I am a big fan of
https://github.com/informatimago
and also Allegro Common Lisp’s open source code
I must admit, the _art_ of writing readable code really appeals to me and is one of my finest joys in programming
I think it's also because they don't deal with complexity as well as younger programmers so they strive for simplicity (which is good).
> I must admit, the _art_ of writing readable code really appeals to me and is one of my finest joys in programming
Totally agree. Unfortunately, this isn't rewarded as much in a professional settings were you're expected to write features. Refactoring is much less visible, unless you're refactoring a piece of code which everyone was struggling with.
Yeah, I'm going to have to ask for details here. When I was a youngster, it was the older folks who were my guides through the forests of complexity, especially when it came to interconnections between things, unintended side effects, and far-ranging consequences of design choices. When I got to that ripe old state of being 40+, that was my role as well.
One of the best compliments I've ever received was when a team member told me I created the most beautiful looking code she's ever seen. This was from both a pleasing to the eye and ability to grok perspective. The latter was mostly due to naming things in ways that made their purpose obvious and avoiding the temptation to do as much work in a single line of code as possible.
[...]
>For example, in Common Lisp
Technology choice can correlate with age. In Lisp's case I would expect that it's long past "cool", i.e. that it's attracting fewer people than it used to, and so I would expect it to skew older.
Just like perl and tcl and awk.
Also you would have to take survivor bias into account - if you only see good lisp projects, maybe that's because the bad lisp projects died out? Maybe the bad old lisp programmers left?
But I agree with your first point, the _older_ lisp programmers may very well be the ones that kept at it and honed their skills while the poorer programmers jumped ship before then
Clean code is easier to reason about.
The easier to reason about the code the cleaner it is and the less context one need to understand the problem the better. State is the real enemy - field variables, mutable objects, lots of incoming parameters, context holding objects all makes it incredibly hard to understand the edge cases and general mechanism of the code.
Thankfully goto statements, global variables and singletons already has stigma attached to them...
As a web and mobile app developer, I find that globals can be useful. Developers will go through extraordinary lengths to avoid making things global, but the truth is that for end-user applications a great deal of the relevant product requirements are essentially singletons. There’s only one active profile and one user and one catalog and one persistence layer. There’s only one DOM and one window. The result is that code which endeavors to make state which is truly shared and global actually shared and global will often be much simpler and less bug-prone, provided that reasonable abstractions are chosen for any mutation and event APIs related to global state.
Globals are appropriate when "there can be only one"; however, I tend to find that's rarely the case. Much more common is "there can be only one at a time"; in which case, dynamic variables are far more useful than global variables. The most obvious example is in tests.
But --- hopefully --- there are more than one unit/component tests. ;-)
But I agree, there are contexts in which global variables may make sense, e.g. sometimes in an embedded system, sometimes in a short throw away program. As usual, the community of software developers discusses guidelines without clarifying the context or the assumptions they are working with. Since there are almost no universal guidelines for programming those discussion go on forever.
Some people write code that reads like that. Readability counts.
Having been on call for other people's code I can tell you that the idiom "Always code as if the person who ends up maintaining your code is a violent psychopath who knows where you live" is a very good idea and you should stick to it.
If I get woken up at 2am due to an outage you caused and open your file to find you've written War and Peace code instead of The Hungry Caterpillar code I will be very very very upset with you.
This article seems to make the case that it's all about taste or lack of familiarity. I'm not convinced.