How do we read code?
blog.theincredibleholk.org
blog.theincredibleholk.org
For anyone curious, I wrote up some of what I learned while exploring the research a few years ago: http://www.clarityincode.com/readability/ (I apologise for the less than stellar formatting; I haven’t updated that page for some time. I also apologise to the actual researchers I cited, if I’ve dumbed down their work too much in aiming for a non-expert audience.)
For the eye tracking reported here, I wonder whether the early emphasis on the top part was a combination of trying to figure out the data flow in the between() function and then its significance to the wider program.
I think it would be interesting to compare the results with a similar eye track of a program written with more emphasis on data flow rather than control flow, e.g.,
def between(numbers, low, high):
return [n for n in numbers if low < n < high]
def common(list1, list2):
return [i for i in list1 if i in list2]
x = [2, 8, 7, 9, -5, 0, 2]
x_btwn = between(x, 2, 10)
print x_btwn
y = [1, -3, 10, 0, 8, 9, 1]
y_btwn = between(y, -2, 9)
print y_btwn
xy_common = common(x, y)
print xy_common
It might also be interesting to compare the results with a functional programming language that expresses those ideas more concisely and/or with tools like between() and common() as part of the standard library that programmers would probably be familiar with.Final thought: How much does the absence of a clearly marked starting point (like a main() function in C) affect how a reader approaches unknown code in Python? If this had been a C program, would the reader have aimed straight for main() and then worked down from there to functions like between() and common()?
Otherwise the results of that research would be biased towards comprehension of meaningless chunks of code, rather that of a real thing.
Then you can move on to what we're really interested in: how do people build an understanding of complex applications with many interacting parts? But you can't go there first; if you do, you won't be able to interpret your results because there will be just as many fundamental questions as experimental variables.
Ideally, we’d compare like-against-like using industrial scale applications, implemented by professional practitioners, controlling for everything except what we’re trying to investigate. Finding opportunities to do that in practice is rather harder, because obviously most real software development projects don’t get implemented twice by identical teams making exactly one significant change in their approach.
What I thought was interesting here was that even though it is only a toy program, there were still some patterns to how the reader explored it that might suggest more general trends. As long as we understand the limitations of the experiment and don’t overgeneralise any conclusions, isn’t some data still better than random conjecture?
It is just, that this synthetic code is broken. It doesn't read. There is no flow in it. It just looks from the first glance, like a non-interesting, unimportant piece, that doesn't do anything, so there is no point in reading it.
Does that sound convincing enough, that there is a big difference between reading synthetic examples and actual code?
Of most experimental research with humans in general, for that matter. It's a big topic of controversy in psychology, because a lot of quantitative psych results are from synthetic tasks in lab settings, which leads to argument over the extent to which those are accurate proxies for real-world behavior.
It’s a specific instance of a more general problem: it seems clear that programmers often work differently depending on their familiarity both with programming generally and with the specific domain they’re working in, but we’ve only scratched the surface in identifying exactly how they work differently, and therefore what practical steps we might take to make things better for programmers in different situations.
Hey, you can even listen to it ;)
It would also be interesting to compare with code with decent syntax highlighting. For some reason I don't quite understand, most syntax highlighting out there highlights the wrong thing: keywords. It needs to highlight the site-of-definition of functions, variables, etc. For example: http://david.rothlis.net/code_presentation/distracting_synta...
Maybe what I don't understand is all the effort people put into making pretty themes for their editors --there are millions of these themes out there-- and almost all of them insist on highlighting (drawing attention to) the language's keywords of all things. The least important part of any program.
1. Establishing familiar patterns is good, and being able to identify deviations from patterns (intentional or otherwise) is probably far more important than we usually assume.
2. You’re going to need to figure out the data flow for some code you’re reading anyway, and a lot of the time the control flow is just noise/accidental complexity, so why not put the data flow front and centre if possible?
Applying these to the example code we’ve been discussing, I might prefer to write it something like this:
x = [2, 8, 7, 9, -5, 0, 2]
y = [1, -3, 10, 0, 8, 9, 1]
print filter ( 2 < _ < 10) x
print filter (-2 < _ < 9) y
print intersection x y
with the proviso that the previous common() function doesn't quite do what an intersection() usually would, which in itself might prompt us to think about whether the previous common() function that is asymmetric in its parameters is actually what we wanted. (Maybe it was, or maybe it was an unintentional error but no-one noticed because the results are the same with this set of sample data.)I have a theory that part of the attraction/intrigue of functional programming for a lot of working programmers comes from the way it inherently emphasizes these ideas, particularly making data flow almost the default form of presentation.
I also have a theory that maybe there's another mental model to be found/learned/created, somewhere above control flow and data flow but below the general domain model, capturing the flow of observable interactions with the outside world: side effects, interactions/dependencies between concurrent operations, and so on. I suspect that part of the reason we find functional programming more difficult for some things (see the “awkward squad” paper) is that when you remove the implicit control flow, you also remove the implicit ordering of events/effects that you get with imperative programming, and there is a cost to that if you don’t replace it (or at least the parts of it that matter for that higher-level view of what your code is doing) with something else.
One of the things that stood out to me in watching the video was how much his mind seemed emphatically not to work like a computer at all. His process gave the appearance of a network self-training for a little while, then simultaneously training and producing output. The more times an area of the program was visited, the better the training, and consequently the longer it could be retained. Notice how the results of calculations are “picked up” from the source and “dropped” almost immediately into the output, as though they’re heavy and difficult to hold on to!
For example, the ternary operator is consider by many to be an elegant solution to simple if statements but from a comprehension (and therefore bug-finding) standpoint, is it superior? Also, how does this effect change as the size of the codebase grows from a single page (as depicted in the video) to a more complex class file.
value = test ? (some_value * multiple) : false_value
# vs
if (test) {
value = some_value * multiple
} else {
value = false_value
}My personal opinion is that this is an excellent use of the ternary operator. When writing software, you want your code to be as simple and short as possible provided it is readable[1]. I find both of them to be equally readable, therefore I prefer the shorter one by a lot. We are talking about 5 lines of code vs 1, which can really add up if this appears commonly all over your codebase. One limited asset every engineer has is monitor real estate, and the more code you can fit on your screen, the better (provided all of that code is the readable variety).
I also find in this particular example, the ternary could get a slight nod for readability as well. This is because with the ternary operator, you start with:
value =
Ok, value is getting assigned to something... then you read the ternary expression. If you are grepping through this code and looking to see what value might be set to, you get to this line and you know you have found it, then you parse the rest of it to figure out what it is being set to.
With the other example, your grepping will lead you to a "value =" that is within a conditional, so now you have to look up and down and explore a little more to see what it might be set to. This is because the ternary operator can only do one small thing, whereas the if statement can do lots of things. Since you are only doing the simple thing, using the simplest possible operator to do that in some sense helps future people reading the code [2].
[1] Which might seem like a spectrum, but I think of it as much more binary. Code is either readable or it isn't. This might seem strange, but my fellow engineers and I spend quite a bit of time grading code submissions to our programming challenges, and of course the "readability" of the code is a key thing we grade for. I think we probably match up our independent up or down votes on readability at least 90% of the time.
[2] If you want to frustrate experienced engineers, have them look through a bunch of code that hasn't been factored down to its simplest form. This is extremely common among inexperienced engineers who refactor code and don't delete a bunch of cruft that now exists due to code or logic changes because "who cares? the code works!", and you get stuff like this:
if ($has_phone_number) {
return TRUE;
} else if ($has_phone_number || $has_email_address) {
return TRUE;
} else if (!$has_phone_number) {
return FALSE;
} else {
return FALSE;
}
return FALSE;
ahhhh!1 file changed, 1 insertions(+), 10 deletions(-)
value = if (test) {some_value * multiple} else {false_value}More to that, in my opinion, "code flow" in functional languages is completely broken.
value = if (test) {some_value * multiple}
else if (test2) {second_value}
else {false_value}
which falls in a straightforward way out of the if/else syntax, which the language authors have put a lot more thought into than the rarely used ?: syntax. value = test
? (some_value * multiple)
: false_value; value = test ? some_value * multiple
: test2 ? second_value
: false_value;
Which, you know, isn't that bad. But it looks really weird since people don't expect you to use ternary that way, while multi-part if statements are fairly normal.2) it's not punctuation. a ternary expression is a 3-input mapping to begin with. punctuation only makes it noisier.
Eh, if you use whitespace effectively it is perfectly readable. Moreso than the alternative, IMO.
I always like to think the text editor is my canvas and the characters are my brush. My job is to make something functional, of course, but also something that looks beautiful.
R ← data BETWEEN limits
R ← ((data>limits[1])∧(data<limits[2]))/data
R ← a COMMON b
R ← (∨/a ∘.= b)/a
x ← 2 8 7 9 ¯5 0 2
y ← 1 ¯3 10 0 8 9 1
x_btwn ← x BETWEEN 2 10
y_btwn ← y BETWEEN ¯2 9
xy_common ← x_btwn COMMON y_btwn
If you know APL the above pretty much reads like the palm of your hand.Since APL is unknown to most, here's a quick explanation.
R ← data BETWEEN limits
Dyadic function declaration. Takes two arguments. R ← ((data>limits[1])∧(data<limits[2]))/data
Let's break this up: (data>limits[1])
Takes the "data" vector and compares it to the first element in "limits", which happens to be the "low" limit. You get a binary vector as the result with a "1" anywhere the comparison is true and "0" otherwise.If "limits" is 2 10:
0 1 1 1 0 0 0 0
Now: (data<limits[2])
Does the same thing with the upper limit: 1 1 1 1 1 1 1 1
Then: 0 1 1 1 0 0 0 0 ∧ 1 1 1 1 1 1 1 1
Performs a logical AND of the two binary vectors, resulting in a new vector: 0 1 1 1 0 0 0 0
Finally: R ← 0 1 1 1 0 0 0 0/data
Selects elements from the "data" vector based on the values in the binary vector and returns the result vector: 8 7 9
Anyhow, that's why I think that language is important. If I wrote this in Forth the pattern would be very different and the thought process required to understand it more convoluted. Probably true as well for assembly.So, reading a code with familiar shape is one thing (that why Lisp has such emphasis on the form of an expression, and Python put that into extreme), while reading long.chains.of.unfamiliar.methods.is.another.)
Then comes recognition of a familiar zones (areas) of an expression, expecting particular kind of sub-expressions here and there. Then match what you have seen with known whole things.
Lets say it is a recursive process of reduction to something already known by examining shapes, forms, and details.
So, shape matters. Small procedures, around ten lines matters. Naming matters, and, especially, using one-letter, non-confusing (no meaning) names for just a placeholders matters.
Let say that this solved in Lambda Calculus (by a naming strategy), and then in Lisp (by shapping strategy) by accident.)
Once I've found the place the problem might be, I start hand-executing (or eye-executing?) for problems.
Does it really matter much how we read new code? We read many tims faster than we write; re-reading is the norm.
It seems to me that in the first few passes on-sighting or sight reading code (as climbers and pianists call it), you're looking for easy to comprehend structures, and blocking off difficult to decipher, simultaneously, so maybe a bimodal distributions at work here
Seems a lot more like abstract interpretation to me ! Would be a lot more logical too :)
There is also alternative approach. If you are just recording and don't need interactivity - all you need, is a regular camera with zoom and a small tripod. You can record on an SD card and synchronize the stream stream time in post-processing. Just flash the monitor screen couple of times at the start and the end of the recording, and play tracking alignment sequence.
Here's an attempt to do just that: http://www.youtube.com/watch?v=zJqKRL9D2qY
A more accurate way of thinking about this is: how much the computer works like our mind. Not surprising considering that people build things based on how we understand everything else. Whether it was done consciously or subconsciously.
Perhaps it's not relevant to the experiment, but it seems worthy of mention at least. Following along, I was second-guessing myself because my answer didn't match...
... that modern languages, such as Scala, offer advantages as human communication mediums. I describe an experiment, using an eye-tracking device, that measures the performance of code comprehension.
Why winners? What a bizarre variable name, to me anyway. Does anyone else use that? I always go with `matches`, `retVals` or `returnValues` depending on the language/IDE I'm using.
Several people have commented on this, however, so I may just change the name to "matches" or something. Thanks for the suggestion :)
(Spoiler alert: Perl doesn't fair much better than the random language!)
Some quick reading on the phenomenon: http://en.wikipedia.org/wiki/Saccade
8 7 9
1 0 8 1
8