How to reduce the cognitive load of your code
chrismm.com
chrismm.com
- Thinking of paragraphs as functions with one purpose
- keeping sentences short to reduce load on working memory and increase comprehension
- create visual breaks to help the reader by grouping common stuff together as mini-functions
- reduce intimidation factor of reading by removing convoluted stuff
- remove cognitive noise (dead code, unnecessary comments, variables declared out of context or too soon etc.)
- keeping terminology consistent across the code (domain language)
- not using double negatives e.g a = !notLocked
- using automated systems to simplify expressions (e.g weird boolean conditions http://www.wolframalpha.com/input/?i=a+%26%26+b+||+c )
etc.
Unlike English, it is less difficult to have a program reorganize your code to make it more readable based on a set of cognitive principles.
But, beyond the readability and understanding of a function, we should also learn from other engineering fields. For example, system thinking helps tremendously in organizing code if the cost of refactoring code wasn't so risky in dynamic languages.
Anybody who has written a book knows the feeling of slowly going crazy because you lose touch with how readers will take it. People writing long works in English have a host of tricks to overcome this. First, they'll have editors. Generally, they have more than one. They'll have sample readers, often quite a number of them. And then they iterate, going over something repeatedly to make it more readable. Writers I know will spend 2-3x the time in the revision process than they did in writing the first draft.
For me, the best way to get that same feedback is pair programming. I was recently looking back at a code base produced by a small team I was on a few years back; the 4 of us did it with pair programming and frequent pair rotation. It is a great code base, one I'm entirely proud of. Clear, readable, well factored, intellectually coherent, and with amazing unit testing coverage. I think that's because every line, every change had two pairs of eyes on it, which meant that we were constantly evaluating readability, constantly testing and reducing cognitive load.
Stop drawing parallels you aren't helping anybody! :D
The beauty of human language is that it allows us to express our individuality as humans. There are an infinite number of ways to write the same thing, and each author can have a unique style based on how they choose their words, structure their sentences, etc. This is fantastic, and is truly one of the pillars of a free society. Consider "Newspeak" in Orwell's 1984 - its entire purpose was to eliminate free thought by eliminating choices in writing and speech.
Now in computer programming, I think unique styles are to be strictly avoided. A Newspeak-ish programming language that minimized the ability to write "creatively" would actually be pretty nice.
On the other end, on static language you will likely end up finding out that your initial architecture did not take in consideration this use case or this API, and you may spend a considerable amount of time fixing what would be trivial to do a dynamic language.
One cognitive problem I've not yet found a good solution to:
So I have a high level routine to, say, sort the items in a display. Said sorting has some cascading effects, so I'll have a high level function that calls some mid level functions that call some lower level functions.
But at some point I get a name clash between a high level function and a low level one. Like...I'm calling the high level sortDisplayedItems, (or sort_displayed_items) and there's a low level function that ACTUALLY sorts them (as opposed to handling the other bits). Sure, I can give it a distinct name (e.g. sort_items), but when skimming the code, it's unclear which is which. If I have to mentally parse the code to understand what is happening, I've failed at readability/maintainability.
How have others handled this?
You're absolutely right: naming things is hard. When I was writing Scheme long ago [in university], our convention was to name helper functions with `-helper` in the name:
sort-foos
sort-foos-helper
sort-foos-merge-helper
Now that I write mainly in Python, the convention seems to be to use underscores to indicate things that you should ignore. Usually that's done at a method level, where your consumers are calling the object's public API (sort), so the object's namespace helps reduce such name collisions. The addition of triple-quote docstrings for method comments helps a lot at clarifying purpose as well: class FooSorter(object):
def sort(self, foos):
"""Sort an iterable of Foo objects"""
# ... sanity-check inputs ...
self._sort(foos) # do the sort operation
def _sort(self, foos):
"""
Sort things we already know are Foo objects.
This is a helper for sort(), which already
sanity-checked our inputs.
"""
# do the actual sorting validateSortDisplayedItems
{
validation Logic
...
sortDisplayedItems(); //Actually sorts items.
}
This can be harder to maintain, but really long names end up a useful code smell.I mean, your above code LIES. If I call validateSortDisplayedItems, I don't validate, I validate AND sort. Plus, what do you do if you have "validateItems" and "sortItems", and then one function that calls them each in turn? call it "validateAndSortItems"? Yuck.
validateSortDisplayedItems
{
if(!DisplayedItemsValid())
{
CorrectDisplayedItems();
if(!DisplayedItemsValid())
{
DisplayValidationError();
return;
}
}
sortDisplayedItems(); //Actually sorts items.
}
AKA validate means try and make valid, not verify that data isValid. So, you can't just do if(isValid) sort; the bonus is unrecoverable errors end up at leaf nodes vs. the happy path.At the high level, your function might be sortClicked, which can then respond to that by calling a wide range of functions. (userCanSort,SortData,UpdateDisplay)
PS: I find the validate > correct loop is generally the important and error prone part of code, so I give it priority. The happy path where everything works is more or less an addendum.
All kinds of functions everywhere validate their arguments; that doesn't deserve to be in their name.
Granted, if it's pure pass-through then reusing the name is a non issue. Foo(x){foo(x);} is not confusing. Foo(x){junk; foo(x);} is.
One option is Foo(x){validateForFoo; ValidedFoo(x);}
Alternatively for private functions JunkFoo(x){junk(x); foo(x);} And replacing JunkFoo with a meaningful name if possible.
Again though, public function Foo(x) should really call internal function Bar(x) otherwise it's really easy to couple things on both sides of an interface.
If someone wants to know how to name the low-level helper which does the real work and the interface over it, we just have to take it for granted that the separation is necessary and that it makes sense for either function to have the name "displaySortedItems" or whatever.
SortGuiElement(){junk; SortUserNames(); junk;}
SortUserNames(){junk;}
vs. SortNames(){junk; sortNames() junk;}
sortNames(){junk;}
It might seem obvious and the code might be identical, but I see the second case very frequently in other peoples code.You want:
SortNames(){junk; SortNamesImpl(); junk;}
sortNames(){junk;}
or whatever: SortNamesGuts(), InternalSortNames(), DoSortNames(), LowLevelSortNames(). Anything but just flipping the case of a letter or two in the "SortNames" identifier.Amway IMO, SortNames() and Internal/Do/LowLevel/SortNames() are almost as bad because just looking at the names it's not obvious what's going on. Yes, visually they look different which helps and consistency can make things even more readable. But, even just SortNamesHappyPath() gives some idea of what's going on.
Contrast this with say, Python, where readability is highly valued and they have a one-way-to-do-certain things mentality, like sorts for example, means code is generally easier to grok, debug, and extend.
Thus, readability becomes especially important when you have to deal with someone else's big ball o' mud and you need to fix a high stress, high visibility, lines down issue at one in the morning because the original programmer is long gone, didn't leave any comments in their code, and decided to play "look at me, I'm so clever with how I use this operator", but forgot to account for a certain scenario that manufacturing decided to roll out the other day without telling anyone. Can I get an amen?
That said, for every time I've hated that someone got terse and clever, I've loved that I wrote something that _read_ well, particularly when returning to previous code I don't remember. Coding is like a language, you can use that expressiveness to be terse, to be rambling, or to be clear. Learning to do so is a skill that has to be learned (cost), but can be more expressive (benefit). Languages (or patterns) that restrict that reduce the cost, and reduce the benefit.
Consider, for example, that Python takes great pride in it's explicit readability, and yet it chose "lambda" as a keyword (making no one happy outside of higher math), that "def" was chosen instead of "define", and that one of the most powerful parts of the language (comprehensions) are highly prone to using secret knowledge. [] works differently than (), for example.
So, what you've described are all valid reasons that Perl has a bad rep...but I don't blame Perl for them anymore than I blame JS for the hideousness that is the DOM, or Python for the bad scripts that have been written in it. Regardless of the language, cleverness and/or terseness without regard to future people (including your future self) is just not a good practice.
though I will say that having one way to do it does NOT mean that
said one way is particularly readable/maintainble - looking at you,
Java
..
So, what you've described are all valid reasons that
Perl has a bad rep...but I don't blame Perl for them
anymore than I blame JS for the hideousness that is the
DOM, or Python for the bad scripts that have been written
in it.
Will you be willing to extend the same charity to the bad Java code you have seen? Or do you have specific complaints for your harsh comments singling out Java?Yes and no.
For the bad Java code I've seen, Java does not bear the responsibility. (And I've seen my fair share of bad Java code because I spent 5 years as a Java dev in a place that had a collection of bad coders and worse code). I've also (since) seen great coders with quality code (in Java), but I definitely retain some bias from my earlier experience.
BUT, when your language design promotes certain things, then yes, that language takes the hit. For example, some of the Perl options are just plain terrible (e.g. changing 0 indexing) and Perl takes the blame. Perl also moved away from those, recognizing the errors. Books like Higher Order Perl (and associated efforts) came out to promote best practices.
Java expressly sought verbosity, and thus takes the blame for verbosity at the expense of clarity.
Other practices are encouraged by the community. That's a gray area - it's not the language at fault, but...it's common. These can range from small and almost petty (Perl promoted underscores in variable names for REASONS, while Java promotes camelCaseBecauseIGuessSomePeopleLikeToSquintToParseVariables.) Not really the LANGUAGE fault, but definitely an expectation. (and one I've grown accustomed to in the name of working with others). Other problems are larger (Kingdom of Nouns) - still not the Language fault, but you can expect it if you're dealing with the code.
Many a time I've tried to trace through some Java code...weaving through "Impl" classes and factories, trying to find where some piece of logic is implemented...only to find myself in an empty class definition. That's a fault of both encouraged practices and poor practices.
When people complain of poor Perl code, it's usually code that was either written by someone that didn't do that as their main job, code from newbies, or code that has changed hands repeatedly with no one trying to make it maintainable. Most any code will suffer in those conditions. Java, thus far, has had the worst overall quality of code in "professional" collections that I've experienced. But then again, I've really only seen code in any quantity in 3 or 4 languages, so....anecdotal experience is anecdotal and subjective.
(I have to give props to Python here...while I have a few nits about a few things, and I'm not well-versed enough to discuss code quality of anything I skim, I've seen a lot of code that was probably "bad" but nonetheless avoided some mistakes that are commonly made in other languages, particularly with new coders.)
- sort_displayed_items_guts
- sort_displayed_items_impl
- ll_sort_displayed_items (ll == low level)
- do_sort_displayed_items
public sort() {
preconds()
try {
doSort()
} catch (ex) {
handle(ex)
}
postconds()
}
private doSort() throws Exception {
// core impl
}And unlike bad abstractions -- which you can "learn" about and mentally model -- the friction of clutter is a constant damping to your productivity.
Building features in a cluttered codebase is like being asked to install plumbing in a hoarder's basement.
Understanding a cluttered codebase is like reading an advanced math textbook by candlelight that was written with broken typewriters by a group of modernist poets.
De-cluttering is the first thing you should do, and you should stick to it above all else. Before you add tooling, before you refactor, before you abstract and framework-ize, get rid of the clutter. That means format your code, and use a linter whose rules are based on widely adopted community conventions.
Hah, yes! That's a great analogy. It doesn't mean you can't get the job done, it just slows you down every time you move around the code base.
see also: https://en.wikipedia.org/wiki/Lint_(software) (this article refers to one particular (original?) lint tool, but there are numerous others for various other languages)
But then, after a certain job, I realized that this is probably not the case.
Now I belive that most bad code out thare is made by overworked and tired programmers in a rush to deliver something that works.
developers have personalities with differing values. That super productive developer may just not value the flexibility of the code due to his personality coupled with that developers personal experience.
That developer will absolutely shine in some environments and be abysmal in others since there are absolutely environments where flexibility isn't an overriding concern.
As you get more experience you learn to be more explicit about your decisions, but you'll still have your preferences and your defaults and it takes a certain threshold for you to move away from said preferences and defaults.
Maybe productivity should not be regarded isolated from other metrics and/or certain style-considerations.
It seems to make sense:
N * (debt factor) = debt
10N * (debt factor) = 10 * debt
I say 'framework code' but it could be called structural code, or glue code or something, too.
So do I spend more time refactoring and bringing this application to a decent state, or do I just keep on hacking more crap on to it? No one except for me seems to give a shit as long as it seems to be working.
Problem is: working memory is short term. After a few days, you revisit that code and you realize how much it costs to load it all in your head again.
Basically, we need code reviews of code we wrote last week as opposed to code we wrote yesterday.
The insidious things about this are twofold. First, it makes all new team members look like idiots. These other guys are getting work done, what's your problem? Our problem is we can't figure out wtf is going on without breaking things to do it.
Second, everyone doing work to fix the problem is a threat, because they are moving code that is memorized.
I honestly don't know how to fix this problem, I used to just walk away but I feel like I should be able to do better than that. Boiling that frog is a long term commitment.
Thankfully the Old Ways are dying, but there are still pretty huge pockets of holdouts. Sometimes it can be hard to tell from the interview process where your prospective employer is at on the continuum.
Both [of the last two, big enterprise] places asked me about testing, and I wrongly assumed that meant they knew and cared. At the former they neither knew nor cared (in fact their dependency graph was so very broken that I couldn't write tests even though I wanted to, and I wasn't going to be able to unwind 250k lines of Big Ball of Mud in the time I had. Hated it). At the latter they care a bit, but are only beginning to understand what kind of trouble they've gotten themselves into.
This isn't the issue. I, and I suspect many others, simply don't have time to thoroughly test code.
It's the age-old issue of "Business doesn't rely on good code. It relies on delivering the product of that code." Up until the ship is both on fire and sinking - nobody cares about sailing a good ship. They'll settle for the shoddy lake boat and ride it out as long as it will last.
I think that's one area where executives and salespeople have acumen that us poor plebs lack in spades. I'm not saying it's a good thing, I just think we get left holding the bag.
Some time before, I set out to outline a disciplined approach for writing code that resembles the practices in the article and commit to following it, regardless of whatever strain of laziness I'm afflicted with in the moment telling me to just break the rules and let it slide.
Contributing to Mozilla in a time before GitHub was a big part of my coming-of-age story, and I credit exposure to the Mozilla code review process and a lot of the other good practices as the number one reason for this—as well as the source of my annoyance when I try digging into some project that I'd like to contribute to, only to find that the maintainers are basically flying the thing by the seat of their pants.
Reminds me of the anecdote about Bill Atkinson: His manager required a "productivity" form that asked "how many lines of code did you write this week?" Having just refactored QuickDraw, with vast gains in speed and simplicity, answered "-2000."
http://www.folklore.org/StoryView.py?story=Negative_2000_Lin...
Like e.g. if you are parsing a file, and it is complicated you can save a lot of time and give a lot more clarity if e.g. you use some form of state machine or look into a little bit of theory of textual transformations and representations.
I mean I have seen people managing a forrest of state variables and who have never heard about regular expressions or any kind of methodology on how to deal with text formats.
You end up with a crap load of messy unmaintainable code.
But of course it is also a problem when people turn every operation into a ceremony involving factories, command objects and any design pattern you can think of. In the Java world that seems like a common problem. I think my own philosophy aligns more with how Go libraries are written. They are pretty simple and straight forward. They use smart approaches were needed but don't insist on adding indirections, encapsulation etc everywhere.
The disciplined developer will produce better code.
I would agree. But I think putting up with those conditions is a different sort of laziness, and agreeing to produce garbage is a different sort of incompetence. Looking back on the times I've done that myself, I deeply regret it.
Programmers have a lot more power to shape process than they realize. There is a giant deficit of good programmers right now. Few managers understand what we're doing anyhow, so if we say, "Nothing gets marked as done until the code is up to professional standards, which includes sufficient unit tests and a well factored design", we can frequently make it stick.
Sure, they will still push you to go faster, but they will always push you to go faster. I think there are better ways to respond to that. E.g., by redirecting their "go faster" energy to breaking units of work down into smaller lumps so they can do better scope control. Or by having them get feedback on new systems early and often, so their decisions about what to make next are informed decisions, not just executive fantasies expressed in bloated 300-page Word documents.
And a given check only has power over you to the extent that you let it. Most programmers will have little trouble finding a job if they have good professional networks or are otherwise willing to work at having options. If you are confident that a just-as-good job is easily available, you can be much braver in the one you have.
Or don't even say it, just do it and don't tell the manager, while incorporating into your estimates.
Overall time to deliver working software with acceptable performance and bug count will still be faster, anyway.
But, I think this is just a symptom of a lack of capacity and resource management.
Contrary to popular belief, it was actually possible to do iterative and incremental 'agile' development prior to 1999. Most engineers back then used the term 'waterfall' to describe a well-known anti-pattern.
Not all programmers are "good" programmers.
Much of it is decent code that was then altered several times, without anyone taking a fresh look to clean it up.
But I'd guess the biggest factor is that most programmers can't tell good code from from bad very well, when writing it.
"null != variable" will confuse people is downright silly. People confused by this won't have an inkling of what any non-hello-world program does.
The rest has some validity, but it focuses on syntax and programming in the very small. It might take a bit of effort, but I can make sense of a tangled function (that's not an excuse to code sloppily though).
The real challenges are architectural: understanding how an application is structured is often a daunting task, and yet there are some very simple things one can do to combat this, starting with grouping things thoughtfully, generous pointers in the documentation, and a few paragraphs of architecture overview in the readme file.
That's why you have coding standards on your team. People agree on the idioms they use. Some people need to be told that it isn't an idiom if only one person likes it, though ;-)
Only because this is how people are traditionally taught to think. This can be untaught easily enough.
If I say (A != B), people may think of this as being equivalent to saying that there's a property we care about- being not equal to B- and we are checking whether A has that property.
If I say (Joe != null), then we are saying that there is a property "being not null" and we are checking whether Joe has that property. Feels pretty natural.
If I say (null != Joe), then we are saying that there is a property "being not Joe" and we are checking whether null has that property. This is pretty amusing way to think. In the lines after one says something like
if (null == myArgument) throw new ArgumentNullException("myArgument")
one may proceed with assurance that some properties of null must hold :-)The distinction between null and not-null is so generally useful that I could imagine making a partially-applied function like NotNull, and applying that to Joe (and many other objects). Would NotJoe get much use other than applying it to null? Probably not.
I don't think it's necessarily confusing, it's just that usually we tend to think of the elements we're working with (variables, objects, functions, etc...) as taking on values, and so linguistically, we ask "is my thing null?" Not, "is nullness something that applies to my thing?" Hence "thing operation value" is arguably cognitively cheaper than "value operation thing."
So you could argue that it's one of those things that increases cognitive load if you're not used to it. That last bit is important because almost anything no matter how convoluted can become acceptable once it becomes a convention. In java I used to wonder why people did this:
if ("".equals(str)) {...
rather than if (str.equals("")) {...
until I got more familiar with NPEs. Now the former seems quite natural.I'd rather write it that way, and then later change it in code review to `(myvar == constant)`, than risk writing it as `(myvar = constant)` and have it sneak through. Granted unit tests might catch this, but perhaps the error is in the unit test. ;)
This is why in modern Java, or in Scala, we have Option types (and in old school java, Nullable annotations), which makes trickery like that unnecessary.
A codebase is like a home. We all have to live in it. Architecture and abstractions are like the furniture and appliances. Code formatting and syntax are the nicknacks, paper bills, decorations, cat toys, board games etc.
Does the toaster really belong on top of that highly flammable mattress? This is a big, deep question about the overall functioning of the home. But throw a pile of ratty blankets on top, put some boxes of moldy half-eaten pizza around the floor, and maybe wind that kite string over the doorknob and around the teetering bookcase. Can you still see the toaster? How do you plan to get in there to move it to a proper place?
Small issues can block, conceal, misrepresent, and make big issues even more dangerous. Which why I say: fix all the small things first and stop new clutter from appearing (lint your code), then talk about rearranging the furniture.
I feel this is one of those things someone new might get confused about the first time they read it. After that, no. IOW, understanding this is part of learning and once it is learned, it is OK.
`if (null == variable) or if (false == variable)` reads well that you are checking for null/false.
`if (variable != null) or if (variable != false)` reads well that you want non-null/false conditions.
It's a "reading micro-optimization" in the same way that not littering your code with comments is a "reading micro-optimization".
It makes you pause and interrupts your flow of thought as you contemplate something not directly related to what you were modeling in your head 2 seconds ago.
Littering your code with comments IS a huge readability problem. You are forced to filter them out to find the real code, since they are mostly garbage, but you can never be sure you aren't skipping some vital bit of information. Likewise, boilerplate and needless ceremony obscure the intent of the code, so they are obstacles to be reduced in good, clean code.
Contrast this with writing Yoda conditions. They are no big deal. They interrupt your flow of thought exactly once: the first time you encounter them. If your train of thought is interrupted every time you encounter them, you are way too novice a programmer. So I think this is a really minor issue in the sea of software complexity; so minor, in fact, that it seems bizarre to me to mention it.
I think whether to use Yoda conditions, much like the placement of braces or the number of whitespaces for indentation, are the stuff of flamewars and endless argument because they are the kind of things we programmers love to obsess about, but they are really not very important.
You're unimaginative.
This conversation is finished.
either ban me or leave me alone until I do something that's actually ban worthy please.
And to be clear, so there is no confusion here.
Your guidelines state the following:
> When disagreeing, please reply to the argument instead of calling names. E.g. "That is idiotic; 1 + 1 is 2, not 3" can be shortened to "1 + 1 is 2, not 3."
According to your guidelines, his comment:
> If your train of thought is interrupted every time you encounter them, you are way too novice a programmer.
Should have been something such as:
> If your train of thought is interrupted every time you encounter them you probably don't see yoda conditionals very often.
Please, do me favor, go warn him as well.
consistency and fairness.
Or just ban me since we both know that's what you want to do.
My use of "you" was the same as yours when you said "it makes you pause and interrupts your train of thought"; I didn't take it to mean you were referring to me specifically! Unfortunately written language is prone to this kind of misunderstandings :(
I ended the conversation for exactly the reason I cited, a distinct lack of imagination. Anyone who immediately reaches for "you must be a novice programmer" isn't really someone I'm interested in conversing with.
Especially since anyone who stopped to think about it for more than two seconds would realize a "novice programmer" would find the yoda conditionals easier specifically because they're still learning to read code.
No one in their right mind would claim an adult reader would be better at having random bits of text in their novels read from right to left in the middle of their left to right text. Most reasonable people are going to agree that it's easier for a young person just learning to read to pick up on that.
Yet there you were, making exactly that claim for programming.
It indicated a reactionary comment with a complete lack of critical thinking on your part and I just have no interest in spending my time speaking with a person whose thought processes work in that manner.
And if that offends you, then so be it. I personally do think you were simply being passive aggressive and that's why you didn't stop to think about what you were actually saying.
Which is your right to do, and I would defend you for it. But it doesn't mean I'm willing to continue engaging you.
Inconsistencies in HN moderation are random side-effects of the impossibility of reading all the threads. You're welcome to bring them to our attention, but please don't take them personally.
Unless it's one of mine, of course. Then you seem fairly consistent.
ban me or leave me alone.
The last two applications I have inherited are horrendously over complex and they are performing relatively simple tasks. Far too many joins done at the application level. The front page of the current application I am working on makes over 1000 calls to the database. I don't think it needs more than one for the main part. (And if the logic is really so complex that it can't be done in the database, that's a sign that the database schema probably needs updated.)
In fact one of the most counterproductive things I have seen at work is people arguing at length about tiny details in the code standard and wasting hours work and goodwill.
My take on this is, you give people advice on why a code standard is good to follow. But if they insist on tiny little quirk, let it be. It is not worth fighting over. The whole development community is full of people with all sort of little quirks.
At the end of the day we got to get stuff done, and you have to weigh the benefits of stepping on somebody's toes against what you gain from it.
Kinda makes my point.
I suppose it could be explained by the fact that it's much easier to say something about the intuition of "null != variable" than to share some architectural insight.
This is the main reason to avoid Spring.
Having so much critical code in XML and properties configuration files, annotations, and magical interfaces that somehow sprout implementations with no source code, make it impossible to debug using the normal techniques of Java development.
"Show me all of the implementations of this interface method"
"Show me where this object is instantiated"
"Show me the code being invoked when I call this method"
In my day to day work, Spring's only purpose often seems to be hiding the code that will actually execute when my program runs.
But there are definitely still times when Spring manages to elude even IntelliJ. How anyone thought "magically appearing interface implementations with no code" was a good idea still baffles me.
Same with a lot of the DI frameworks that were so popular for a while. New guy doesn't wire something up correctly, checks it in and everyone spends the next 3 hours debugging a rabbit hole. (OK, slight exaggeration.)
Yes, yes, readability and problem separation and stuff is important, but so is efficiency and usability and security and deadline and everything. Bottom line: balance.
And it's the hardest thing to achieve.
Literally anything taken too far is a bad thing. And just about everything in life is a matter of finding the right balance.
I really like aligning multiline blocks, adding whitespace and useless braces here an there. e.g: Having just a single space between function name and arguments makes it look less like a call. Yet almost all lint presets/defaults forbid this. Typography is all about the whitespace between letters forming easily recognizable shapes.
But our current programming systems, we are still very much in the dark ages of code typography.
This is brilliant. A neat way to bootstrap getting this sort of thing implemented in most code editors would be to write a plugin that can do this for existing code and make it good enough to turn heads. It would extract documentation blocks to be presented as prose in a vertically split pane to the right and present the file itself with those blocks hidden—as if automatic folding were turned on. It could even apply some fairly simple heuristics to automatically link to other relevant comments. The goal should be an experience indistinguishable from an embedded iframe showing human-generated API docs from the Web.
Another thing I'd like to kill is the file tree that you see in most VCS Web frontends that tell you the last commit message that touched the file/directory, rather than about the structure of the code.
Netscape's old Bonsai tool tried to do something like this. (When Netscape open sourced Mozilla, they also opened up a lot of their internal tools. This is where Bugzilla came from, but there were others, too.) When you were looking at a directory listing, if the file contained what looked like a short description of its purpose in the comments near the top of the file, Bonsai would grab description and present alongside the file name.
In my own projects today, I try to always include a file overview containing a short, single line description and then write a paragraph or two going into further detail, documenting the whys of the code, and generally explaining its overall role in the project/justifying its existence. I'm basically writing for a tool that doesn't exist but that I'd like to see get created and gain widespread acceptance.
Unfortunately, literate coffeescript is not really mature enough to be used in a large project, IMHO. However, I hope that more people think about separating human language commentary from computer code. I think the idea behind literate programming is a good one and I hope that it gains some traction some day.
And I find the idea of a tool that automatically rewrites machine readable code into a natural language to be of dubious value beyond use cases where someone is first picking up the language. Similar to those tools that exist to generate comments in the form "Set global position" based on a method named setGlobalPosition. It just creates redundancy, and if you're committing the output to your source tree, then it's redundancy in the form of clutter, too.
What I'm thinking of, as I said, should shoot for parity with having a half-screen browser window open to the right containing the relevant docs. Only in this instance the docs are "live", and the lookup process is context-sensitive, requiring very little manual effort to perform it. I know that a basic attempt at something bearing minimal similarity is available in most editors that try to implement Intellisense, but generally I find the helpfulness of the small popup in most implementations to be limited to helping you get the method signature right, and not much else.
One of the nice side effect of the design I'm talking about would be a system that encourages keeping the docs up-to-date and useful as much as it encourage their consumption.
Literate coffeescript does not have the tools for extracting and rearranging text, so it's really just a way of embedding markdown text into your coffeescript code. However, it's useful because you can embed html hyperlinks which can do things like enable you to click to get to the tests, etc.
Here is a small example of something I wrote in literate coffeescript: https://github.com/ygt-mikekchar/react-maybe-matchers/blob/m...
Now imagine that you have the English text on the left hand side and the source code on the right hand side. Ideally you would have tools that would allow you to make the hyperlinks (possibly automatically) and keep the documentation in sync. Such tools do not yet exist at the moment, unfortunately.
Edit: I should admit to being embarrassed about my fluent interface abuse in this code ;-)
> formats the output in a similar way - you have the english text in a pane on the left and the code in a pane on the right
... to be a description of a system that didn't really sound like what I had in mind, and didn't really sound like literate programming, either. After reading this comment and a reread of your original one, I understand I was wrong in my interpretation of what you meant. Sorry about that.
if ((mouse.x == 0) &&
(mouse.y == 0)) {
scores more points than: if ((mouse.x == 0) && (mouse.y == 0)) {Take a look at the source for tcsh, this is exactly how it is done there. Unless you're interested in supporting a long dead platform - it makes things unnecessarily complex. I love tcsh, but so many ifdefs in so many static functions...
For example:
sqrt((x * x) +
(y * y) +
(z * z));
You can run your eyes up and down each column to verify it's squaring x, y and z.That reflects the structure and symmetry of the expression better than:
sqrt((x * x) + (y * y) + (y * z));
Did you spot the error? sqrt((x * x) +
(y * y) +
(y * z));
How about now?Here's some code that has a lot of examples of that style, a JavaScript implementation of a weird hybrid Margolis cellular automata neighborhood, which has a lot of two-dimensional patterns:
https://github.com/SimHacker/CAM6/blob/master/javascript/CAM...
It might not be the default linter styles, but set up your linter for your project and give your mind more important things to focus on.
annotateAsts = import ./annotateAsts.nix { inherit stdenv annotatedb; };
runTypes = import ./runTypes.nix { inherit stdenv annotatedb jq; };
dumpAndAnnotate = import ./dumpAndAnnotate.nix { inherit downloadAndDump; };
It would be nice to have a tool rate various equivalent arrangements and warn if it finds one with a significantly better score, e.g. showing me the above if I'd given it something more confusing like: annotateAsts = import ./annotateAsts.nix { inherit stdenv annotatedb; };
runTypes = import ./runTypes.nix { inherit stdenv annotatedb jq; };
dumpAndAnnotate = import ./dumpAndAnnotate.nix { inherit downloadAndDump; };
Of course, as well as formatting it would be nice for equivalent representations of the same expression to be compared, e.g. using an SMT solver or genetic programming. For example, in Nix the variable names after "inherit" can be in any order, so it's easy to find permutations which highlight common elements (like "stdenv" and "annotatedb" above); if I'd written these in a different order (e.g. "inherit jq annotatedb stdenv;" on line two), it would be nice to be shown rearrangements which score more highly.It's not just linters either. I can't even find an indenter which handles 2D alignment. For example, indenting something like (random bash code I have open at the moment):
jq -n --argfile asts <(echo "$ASTS") \
--argfile cmd <(echo "$CMD" | jq -s -R '.') \
--argfile result <(echo "$RESULT" | jq -s -R '.') \
--argfile scopecmd <(echo "$SCOPECMD" | jq -s -R '.') \
--argfile scoperesult <(echo "$SCOPERESULT" | jq -s -R '.') \
'{asts: $asts, cmd: $cmd, result: $result, scopecmd: $scopecmd, scoperesult: $scoperesult}'
Emacs wants to put the second '--argfile' directly beneath '-n', which is clearly confusing compared to the above. If linters solve all typography issues, are there any which can be queried for the local-optimal indentation on a line-by-line basis? loggedIn = foo
isAdmin = bar
It doesn't look like much, but when you are glancing over code it really helps to quickly read it.I'm of the opinion this should be extended to "does not play well without an IDE". Because even in projects that said "everyone, use Eclipse(/IntelliJ/whatever)", and tried to share project files, there was constant pain in ensuring that everyone had the same development environment ("Oh, yeah, I made a local change to my project file, but then I also made a change that needed to be shared, and whoops, I broke everyone" and the like). Not to mention trying to work with code deployed onto boxes without an IDE. I've gotten to where "if I can't make sense of it and be productive in it with just vim, grep, and find, your code is too complex".
Requiring an IDE is, to my mind, a symptom of a far-too-complex environment which will lead to breakage.
And soon enough, my next job the code was impossible to understand without autocomplete and a debugger. I am still sad about that to this day.
I think that if you do things correctly, your concern is mostly mitigated. My last job was Java/Spring. As someone noted above, you really don't want to be working with Java/Spring without IntelliJ.
But we never had issues with anyone breaking builds due to screwing up project files, because those weren't shared and the project would run/build/test immediately after import into the IDE without any additional tweaks.
I think it's sort of odd that many folks complain about IDEs, but then use vim with a million plugins. Isn't vim + plugins = IDE? Or, if the overall point is "java sucks, don't use java" then fine, but if you have to use it, I couldn't even imagine using plain old vim without any language support...
Whereas I've gone to projects that were written without all of this, and even without getting the code to even compile locally (because of dependency hell that I didn't want to suffer through if I didn't have to), was able to diagnose and fix issues, because the code was written to be understood with only a text editor.
By not sharing IDE project files, not having any IDE-specific machinery in the project, but still expecting CI to be able to run it, being able to theoretically hack on the project and get results (build, test) even with just bash and grep and nano, you get the best of both worlds.
Modern IDEs are better at this than they ever used to be. I was really anti-IDE for a long time, but JetBrains' products made it make sense finally.
It's funny that you mention the Java ecosystem as one of the worst offenders, since the nature of the language itself and the culture around best practices in the early days should have put it in a particularly favorable position here. This is covered partly in another great post titled "Java for Everything": http://www.teamten.com/lawrence/writings/java-for-everything...
Unfortunately, I too find most Java projects are developed in a way that make them completely unapproachable without the assistance of an IDE. I'm secretly hoping that Microsoft announces first-class support for Android app development in VSCode, given that the team's overall gestalt is on producing a nimble, code-oriented editor. They may not otherwise have any incentive to make that kind of thing happen, except that I think it would be a real boon for VSCode adoption. Anxiously waiting to see if there will be any interesting announcements to come out of their developer conference this weekend, although I'm not counting on it.
Refactoring code without side effects is so incredibly liberating: I barely have to compile many changes, much less carefully test them, because they're so obviously correct, because I can SEE the flow. I don't have to speculate on what the rest of the system thinks is going on.
This is the main reason that Haskell is such a force multiplier for me. All the small minutia is well defined and it frees you to think about the big picture. Sure, it's a big learning curve, but so worth it :-)
The readability wins I've seen come from using common tools; I'm going to be much more productive on a codebase written in a library ecosystem I understand. More generally, a powerful standard library with consistent calling conventions can help a language be useful. It's not productive for me to spend 10 minutes reading a function only to realize 'oh, this human wrote their own string split'.
Tests help too because they make clearer the dependent typing of a function's call signature; stuff like 'don't pass both of these variables' or 'a must always be > b' are never clear (and I've never used a design by contract langauge). Tests are sometimes better than documentation because you can sometimes get alerted if they're wrong.
Most important rule for readable code: hire programmers who know how to read. Reading a large project is a skill and lses people can read than write code. Every project is going to have a quirk of its evolution that's hard to understand without a close reading. (Every large C project contains a buggy partial implementation of LISP). Hire programmers who can survive that.
It feels like maybe people are almost ready to listen now. I don't know how this relates to the number of libraries we use now, or the StackOverflow culture, but I have my suspicions that they are related. Maybe we should start talking about Library Fatigue...
The one time I let comments slide is when they document unexpected and unavoidable behavior that I can't fix in code. For example working with an inconsistent data set and having to add some convoluted code because what is a string in 95% of the rows might parse as a tuple in the remaining lines, so I need to account for that. This is so that I, or the next guy, knows why the code was written as such and won't go off trying to fix it.
He eventually gets to "hungarian notation" (those redundant prefixes you don't like) towards the end of the piece. tl;dr: misinterpretation of a good idea, cargo-culted down through the ages.
All strings that come from [user input] must be stored in variables (or database columns) with a name starting with the prefix "us" (for Unsafe String). All strings that have been HTML encoded or which came from a known-safe location must be stored in variables with a name starting with the prefix "s" (for Safe string).
us = Request("name")
...pages later...
usName = us
...pages later...
recordset("usName") = usName
...days later...
sName = Encode(recordset("usName"))
...pages or even months later...
Write sName
The thing I want you to notice about the new convention is that now, if you make a mistake with an unsafe string, you can always see it on some single line of code, as long as the coding convention is adhered to: s = Request("name")
is a priori wrong, because you see the result of Request being assigned to a variable whose name begins with s, which is against the rules.[1] http://www.joelonsoftware.com/articles/Wrong.html
(edit: formatting, citation)
The article goes on to advocate prefixing the methods that return safe and unsafe strings with 's' or 'us' I would hate that! I can easily see you ending up with string methods like safePadLeft and unsafePadLeft that do exactly the same thing but are named for where in the code they are used...
Coding is hard, and as other comments have stated all these blog posts and books need to be condensed into 'balance' and that balance point is different for each coder, project and team.
I don't follow. Since the argument type to either method would be the same (a string), why are two methods necessary? Or are you inferring that we define subtypes of String: SafeString and UnsafeString? And wouldn't subtypes inherit behavior from their parent?
Advocating for prefixes like this is just going to result in badly obfuscated code where we get the exact opposite of, "JavaBeanFactoryConstructorXML_dom_to_json".
You're not supposed to write
iLength = iCount * iSize
but rather (something like)
meterLength = unitCount * meterSize
newtype UnsafeString = Unsafe String
(UnsafeString is the name of the type. Unsafe is the constructor you use to create an UnsafeString from the String).
With this and other techniques you can lift these characteristics of you data into the type system, which is IME much more useful than using hungarian notation.
Marking something unsafe is a best case scenario. Separating lines from rows would require a huge amount of boiterplate to preserve the numeric operations, and one still can not create a library that will check something like:
let
l = 5 :: Meter
w = 4 :: Newton
in l * w :: JouleNot sure what you mean by separating lines from rows. In the example you give of computing with units, I've seen examples of such systems in Scala (e.g. http://www.squants.com/) and F# comes with something built-in for this (IIUC). You might be interested in checking them out; it's quite neat.
It's the example Joe Spolsky used on the blog post about prefixes that everybody keeps pointing to. He showed code that worked on rows and columns (not lines, sorry) on a table, using prefixes to avoid mixing them.
Haskell types do support this use, but you'll have to declare both types as Integral and Numeric, with the resulting boiterplate. I'm not sure exactly how to solve this, maybe a way to "lift" typeclasses out of a newtype declaration would be general enough. It is probably possible to do with template Haskell, but that is already outside of the language.
That Scala library is interesting, although it looks like lots of the work is done at run time. Also, it looks like Haskell is able to do it [1]. (It's a recurring theme of mine, to complain that something can not be done on Haskell, just to discover somebody that did it.)
I assume that global variables would be horrible in such an environment, but typing out the prefix and maybe one letter of the actual name would cut the completion list down to a manageable size.
Perhaps apps hungarian started because it helped avoid common errors in complex GUI logic, then similar practices were carried over to the systems side where they were useful for mostly unrelated reasons, eventually reaching the world in general through the header files produced by the systems folks but lacking any of the context or history needed to go beyond "microsoft does it that way, there must be a good reason". Hopefully I'm just being overly pessimistic and extrapolating from massively insufficient information, though...
1) Reduce the number of variables that need to be kept track of in order to make a function easier to understand.
2) Avoid metaprogramming unless there is a clear need for it (doing something over and over).
3) DRY isn't always good. Sometimes being more verbose is easier to read than being clever so you don't have to type as much. Concentrate on readability first.
4) When there are a lot of interacting components consider the DCI pattern. It will save the developer from bouncing from file to file, module to module, just to understand the flow of an algorithm. Each time a developer needs to look up a different file more cognitive load is introduced. An algorithm should be easy to follow in a single file with code in sequential order. The opposite of this is message passing and having components pass messages to other components.
5) Syntax matters. Unfortunately once you choose your stack there isn't much you can do about this. Some syntaxes are much noisier than others. Every little bit of noise adds more cognitive load. Don't believe me, try do long division with Roman numerals.
6) Compress complex concepts into shorter ones. This builds on the Sapir-Whorf hypothesis. This might be a moving a part of an algorithm into its own function or storing an intermediate state in its own variable (as opposed to function composition 4 levels deep). `map` is much simpler to understand than `for (var i=0; i<arr.length; i++) { doSomethingWith(arr[i]) }`.
7) Spend time getting to know your editor. When you can reduce the amount of muscle movement required to perform an action it generally reduces the cognitive load as well. Not to mention making you more productive.
People need to see how the code flows and how things are logically connected. That is impossible to do with reams of methods operating exclusively on member variables.
I guess this is just another way of stating the benefits of functional programming. People do modularization at the function/method level so badly so often that I often think talking about modularization at class level, is beginning at the wrong end.
So many struggle with modularization at small scale.
In my experience, programmers generally have a sense of what parts of their code need to be cleaned up. If they are a newer developer their ideas might be a little screwy. They might not see the root problems, but they know which closets are producing ghouls.
The issue I see is that developers don't think refactoring is a good a use of their time, or someone up the food chain doesn't think it's a good use of their time, or there's just so much institutional inertia making people question whether it's a good use of their time.
If you want your code to be better though, don't read about architectural problems, just take some time to fix the stuff you know is bad in whatever way you know how.
That's better than any article or class and it pays you money and it makes your code more fun.
Ideally, sequence diagrams (or a similar type) can help shed light on just how the project comes together. Explaining how a project works from a high level can leave unintended knowledge gaps. The great thing about these diagrams is that they can largely be automated with very few inaccuracies.
Diagrams are definitely now a silver bullet. I have found that they can easy grow unwieldy when attempting to incorporate too much information into a single diagram. Much like separation of concerns or the single responsibility principle, I believe that diagrams work best when they're focusing on a single piece of major functionality.
MyObjecter.GetObjectRelatativeThingy(ThingyHolder.WhatsMything(IndexKeeper.Get(), OtherWhatsIt.BuildOther( MoonCheese.Intensify(true)))
It's probably a result of the letterbox effect of widescreen monitors and IDEs with horizontal toolbars and status bars.Is there anything wrong with expanding the idea of 'one purpose per function' into 'one purpose per line' and split it all up into multiple lines with temp vars inbetween.
this gives you benefits of:
- short lines and simple ideas
- easier to set break points
- allows for adding logging between calls.
- allows for setting debug watches on the temp vars
- the compiler tidies any temp vars away so there shouldn't be any extra copying.
my fave is clean code... though uncle bob seems divisive these days for some reason.
http://www.amazon.co.uk/Clean-Code-Handbook-Software-Craftsm...
This technique is named 'carbo-loading' (because of all the spaghetti).
https://github.com/johnpapa/angular-styleguide/blob/master/a...
Unfortunately, since the simple examples use Folders-by-type, and many people are incapable (or unwilling?) to think for themselves - they just continue that way of doing things, even in giant projects. Hell, EmberJS embodies the Folders-by-type into the framework.
I can read code by the page, hitting next page at about a 1Hz rate. If the code is not overheated. That means, avoid lots of syntactic bloat, keep it concise, keep it modular, with low branching. Just about what the OP says.
Until you have some code reviewer that thinks otherwise because of some "stupid reason" and you can't get around their hard heads.
One example, breaking a 81 char line because it goes over the limit and getting two shorter lines that are awful to read
So yeah I'll go for this when I'm working with reasonable people
(Not being snarky, I would like to do some iOS development for fun but want lines under 78 chars, so it don't go over 80 with diff.)
Everyone else on my team is in love with writing 120+ character lines, which makes me sad.
It's kind of how legible handwriting is desirable, but not really the most important factor, when it comes to the amount of brain required to solve a math problem.
There are 5 concepts in the article. Off-by-one error? :)
And given most languages in use today use = to mean assignment instead of equality, it's hardly "no reason".
Equality operator is commutative (sans operand side effects). It's not a quirk.
$ jshint test.js
test.js: line 3, col 10, Expected a conditional expression and instead saw an assignment.
1 errorSometimes the workarounds are necessary for performance reasons, or they're just good practice to avoid common pitfalls. But in a lot of languages they aren't, so unless you have a good reason for increasing the overhead, it's better to lean towards readability.
Edit: Also, for this particular example, a good linter is all you need to warn you about an accidental assignment. If you don't have a linter, sure, then put null first, but it's important to recognize that this sacrifices a small amount of future scannability for the more immediate avoidance of a bug.
The blog post over states its own importance as after learning these clean code would be easy. Like fade diets, magic paradigms etc... another example of someone believing or trying to convince others that there is a "magic path". "Just remember these four concepts", "get a six pack with just 20 minutes a day"...
There are a few good tips in here but really nothing new. A nice reminder that there are some simple steps to help improve your own code.
A large devil of clean code is not just in nitty gritty "have one line per logical action" but rather the deconstruction of a problem into easily followable steps.
In that way this is definitely not an exhaustive description.