Functions Should Be Short And Sweet, But Why?
sam-koblenski.blogspot.ru
sam-koblenski.blogspot.ru
If you're writing big functions by habit then you're just laying more code after existing code to build a program. If you split your code in functions and, instead, start talking in the language you've just invented then you're composing a program. And I think compilers should do the building while I refrain myself to composing.
Of course, one-line functions don't have inherent value but sometimes they're needed. For example, to present a higher-level concept that just effectively happens to be a simple assignment or an inc.
Short functions themselves don't have an inherent value either: if you manage to write your function with less tokens, possibly due to composing it out of other functions, then that's all good but isn't the real value of doing it. Shortness for the sake of shortness only matters until you're down to a screenful or so.
The real value of defining and writing functions is thinking of the proper words for your language. As soon as you're making your own language, then you're just using functions as vehicles to represent words in that language. Some words might take 30 lines to code, some words might take 3 lines. It doesn't matter as you've got your language together in a way that is concise and beautiful which is the real point.
This is something Abelson & Sussman were trying to teach us in the 80-s (stacking abstractions-as-languages on top of each other is one of the main themes of SICP). Unfortunately this insight got somehow lost in between people parroting design patterns and repeating things without understanding them (programming pop-culture?). I've personally seen only a handful of programmers who know of this concept, and even less of folks who understand it.
I'm not unduly obsessed with classes or design patterns in object-oriented languages, but to me the key benefit of functions, classes and design patterns is about defining a shared language of named abstractions.
These abstractions take care of details, so you can think at a higher level. The names are a common vocabulary shared between the designer/implementer and user/client of the code; and as yason says, if you choose your abstractions at an appropriate level, you can "talk about" (or describe/define) the problem+solution in terms that are best suited to the problem+solution.
And to note one more benefit of short functions: When you're talking about functions (operations) and classes (encapsulated state), you rely upon these abstractions to behave in a certain way. Shorter functions and simpler classes generally have fewer moving parts, thus making it easier to verify the correctness of the implementation of the abstraction concept.
Though the naming seems to give documentation benefits, naming variables and functions is one of the hardest things we do as coders. The practice of properly factoring your system comes more out thinking in terms of values rather than steps. When you think "what value am I computing?", you not only write short sweet functions naturally, but also build more robust code in the process.
And yes functions are values as well.
I'm not saying this happens all the time. But I'm more concerned about it, because I see it more often than the inverse, and it's more damaging to the codebase. At least long spaghetti methods can be refactored into something nice. This kind of mess has to be refactored into a single area first, which is much harder.
indeed, I have seen functions that contain only one line of code. Even if that is intended for reuse, it makes no sense, because the function call is one line too. Readability goes down the drain this way.
If a reader tries to understand a logic and thinking has to be interrupted by searches and trying to find the right code window, split-ups don't do good things.
Polymorphism, if applied wrong, is a notorious tool to hide the logic from the readers too. If the code decides at runtime, which implementation is used, comments become really important. Especially when this decision is made completely elsewhere in the code. Think strategy pattern, for example.
The underlying problem is IMHO developers blindly following rules. This can be seen a lot. If you ask them 'why' and you get a textbook answer, be careful.
def logUnexpectedError(t: Throwable) = log.error("Oops: {}", t)
After I get it together and start keeping proper metrics: def logUnexpectedError(t: Throwable) = {
log.error("Oops: {}", t)
statistics.errorCounterByClass(t.getClass()).inc()
}Summary: Performance almost certainly isn't an issue (compilers are smart). Abstraction leads to better design, using a well-named function can improve readability, making it easier to understand, easier to prove correct, and easier to modify. When you keep code inline it increases the inertia of the given design and implementation, extracting even a single line into a function (when warranted) can make it far easier to drastically change implementation. And if you want that single line of code to become several lines, or to switch to calling out to a web service, implement caching, add error checking or what-have-you that's made much easier when it's encapsulated inside a function than when it's sitting inside a block.
Componentization is always good, even when the components are sometimes very small. Like all design choices extracting code into its own function requires judgment based on experience but when done right it can have many benefits, even when it's a single line, even when it's a fraction of a single line.
Blindly following a rule such as avoiding making functions too small is just as big an error as avoiding making functions too large.
odds = filter odd
While that might not make sense if you only call odds once, you might end up calling it a dozen times! You may also want to limit the scope of this function definition by using a let or where clause. This can aid readability if it allows you to avoid repetition.Probably the classic case of Imperative-Guy-Struggling-To-Get-FP.
While I'm admittedly not best placed to comment on the matter, it seems to me that what Haskell offers is not so much greater power than a Lisp, as greater concision, as well as a language structure which doesn't allow for the sort of "escapes" into imperative behavior that a Lisp programmer can employ.
In fact, now I think about it, I'd say that Lisp and Haskell are in this way roughly analogous to Perl and Java. I'm sure this claim will infuriate partisans of all four languages, so let me hasten to point out that the only sense in which I mean it is that Lisp and Perl allow the programmer considerable stylistic freedom ("TIMTOWTDI"), while Haskell and Java go to great lengths to enforce that style which their respective implementers consider the One True Way.
Both these traits have benefits, of course, or one would've probably outcompeted the other by now; of particular note here, though, is that the One True Way style requires a lot more front-loading on the part of the developer than the TIMTOWTDI style does -- in the latter, you can more or less "fake it 'til you make it", while in the former, trying to "fake it" results in a nasty dressing down from the compiler, and you can't accomplish anything until you've become at least a neophyte, preferably a proper acolyte, of the One True Way.
All of which is a very long-winded way of saying that if what you're interested in is not so much Haskell specifically as the functional style in general, you might be well advised to start with a Common Lisp or a Scheme, where you can study the functional style without needing first to understand any large and rather dry bodies of theory. Granted, you'll probably screw it up a lot right at the start, at least if my own experience is any indication. But Haskell's a big elephant to swallow all at once, and starting in a Lisp or a Scheme will let you nibble around the edges at your own pace, rather than being required to choke down nine-tenths of the entire carcass just to get to Hello World.
@pl transform k z = reverse (foldr (\x y -> y ++ [fst(head(filter (\a -> snd a == k-1) (zip x [0..])))]) [] z)
Transformed into: transform = (reverse .) . flip foldr [] . (flip (++) .) . flip flip [] . (((:) . fst . head) .) . (. flip zip [0..]) . filter . (. snd) . (==) . subtract 1
Yeah, neither is readable. This function ought to be broken up into cleaner, more declarative parts. In fact, the whole function can be reduced to this: transform k = map (!! k)
Much more readable, wouldn't you agree? Now all you need to know is what map and !! do.IMHO it can make a lot of sense. In many cases you can clearly specify the intention of that one line by putting it in a function with an appropriate name. Just look at the example from the OP:
void HandleNop() { SPIDataPut(SPI0_BASE, HOST_REPLY_POSACK); }
When you want to handle a NOP, you're not interested in how that NOP is handled, you just want it handled...:)
Thats the beauty of a forum, if used in a productive way :o).
I wonder if it's a limitation of our IDEs using plain text and files. Maybe the next generation of programming languages will be nodes or visual blocks that chain together or something? Who knows haha.
The problem is, you still need to have control at a "line of code" level of granularity (though not quite a line of assembler mnemonic) if you want to have general purpose programming power. You could probably, with some work, make a graph tool for something like ifttt[0] and do powerful things relatively simply. It's not quite programming though.
I've used state machine UML code generation systems in which you "program" by moving blocks around (something akin to a flow chart). Even when the system's purpose is something extremely limited, e.g. handling a phone's menu transitions, there's still a mind-bending amount of complexity to handle.
I believe ImageJ[1], a freely downloadable image analysis tool, is among the tools with this sort of UML functionality built in if you want to play with it.
I'm not trying to claim some complete, graphically represented program will have as many operations (or boxes) as the equivalent C code. It won't, however, require less cognitive overhead than crafting the program in the language of your choice without making most of your choices for you (and in the process becoming much less powerful, not general purpose at all).
edit When a computer is smart enough to derive intent, then a picture will be worth a thousand words. Until then...
Also the fact it boils down to code is not true at all. It all comes down to an AST. Text based representation of computer languages is rather archaic when you think about it.
I mean, look at IDEs. So many handle code transformations for the user because of how tedious the work is. Not only that, but in order to perform these translations it is transforming the code into an intermediate form it can work in, performing the specified transformations, and then transforming it back to the textual output. It is crazy.
Well, you've gone over my head, but given me plenty of reading material. Thanks.
> Instead they like to emulate statement based languages with cute but not very abstract or reusable boxes.
Indeed, though they're rarely cute.
> Text based representation of computer languages is rather archaic when you think about it.
As a workaday sort of guy, it's not really. The code I'm writing at the moment is built on so many abstractions of the underlying storage, the atoms of functionality on the processor and other realities, that I am able to fairly naturally "think in" the language. My run-time makes decisions about sorting, searching, filtering, iterating and so on for me. The typing and context assumptions it makes, the terse nature of it (without golfing) - the attempt to divine my intent without so many words spilled, means I never feel like I'm doing busy work - the boiler-plating, the exception trapping and so on. Someone else's implementation of any given task I need to perform is probably a single command away from being available to me.
I may have twisted my imagination to its will, but we get on well now. I don't find it tedious and don't find I need an IDE or related tools to count the pennies.
Do you have a vision for a "next gen" (not quite Star Trek Next Generation - "Computer, extrapolate" level) programming paradigm? Are we talking an overhaul of computer architecture or will there always be some clever person writing our C for us? Abstractions all the way down...
Also, an AST for some given procedural code, when represented graphically, might look (superficially) like a Jackson Structured Programming diagram, which would have been created by hand... Maybe text is the wrong tool :)
I had confused it with LabView...
With some work and some luck:
- all functions are short enough but all of them do something more than just passing things around.
- there is never any ambiguity on where a function's implementation is to be found. (In Python, I prefer "import x.y.z" then "x.y.z.myfunc()" than using "from x.y.z import myfunc", for this reason)
- like-minded functions are grouped together.
- functions are nice enough (read: testable and tested enough) so that there is rarely a need to check its code to use it.
When I tried to solve some simple Java exercises with classic Literate Programming (noweb[1]) -- I found that it helped enormously with simple top-down design -- but I ended up with pieces of code that were a lot smaller than the natural java abstractions -- it felt natural to factor out parts of loops etc... mostly this was me overusing the power that comes with literate programming -- but if done reasonably it could also be rather effective.
It allows one to start essentially with pseudo-code -- but that pseudo code often ends up dealing with stuff like arrays, that with (classic, pre-iterator) java implies explicit counter-variables (like "i") to keep track of an array pointer. So you suddenly get a problem reusing blocks of code, for an inner loop that indexes by "j".
At the same time it made for wonderful compact and readable code (where "code" means the Literate Programming document, not the actual java noweb emitted).
http://www.johndcook.com/blog/2012/01/09/holographic-source-...
It's as much a nightmare to work with as overly coupled code is.
Bart Bakker references "Structure and Interpretation of Computer Programs" as an early source and I'm somewhat positive that I read about an similar approach in "How to design programs 2nd ed", called "programming with a wish list".
Edit: I found the mention in HtDP but as I understand it it's not exactly the same as wishful programming: http://www.ccs.neu.edu/home/matthias/HtDP2e/part_one.html
under "3.4 From Functions to Programs"
The main reason for going small at first, I think, is that it creates more names. Names especially help with very simple calculations that appear repeatedly and turn into visual noise(true of a lot of UI code). The naming process can point you towards an instant design change as you discover what you are "actually" doing in the code. A big function acts as a bigger black box, thus it can hide these issues. Sometimes you want the black box, but it's the exceptional case.
In my case the why (as well as many of the arguments here) is because I'm not a "superstar/rockstar ninja coder" and by keeping functions short, clear and with a descriptive name in 6 months when I open the file to fix the change request I won't have to spend 2 hours building a mental model of the system to fix the bug.
e.g.:
var showPreview = function() { ... };
var calculate = function() { ... };
$('a.preview').click(showPreview);
$('a.calculate').click(calculate);
Even nicer is to split this out into a separate CoffeeScript file. Then just call `new SuperAwesomeCalculator()` on the page.I think what made me do this was Erlang. After a point it gets annoying to type NewNewNewNewNewParams, so it's easier just to split it out into a separate function :)
So yeah, small functions good, but you can take it to extremes.
Compounding that issue, I've run into cases where people pulled out code into functions, realized that in two places the behavior was ~slightly~ different, and they addressed it by adding a variable, 'shouldThing' or something, that serves just as a boolean flag as to which bit(s) of logic you should fire. So now rather than just have to focus on the code that needs changing, I've got to go to an entirely different section/file, and wade through code that doesn't even relate to what I'm trying to fix.
You can have a fairly long function that's readable and maintainable; those last two are the important part, and, to a point, they hinge more on how much work is being done than how much code it takes.
Of course, if you're just plain writing more code than is needed, that's not a problem with the function itself.
Would you rather have mostly short functions with a couple long ones or the opposite?
Of course, I've seen many cases where people take this to the opposite extreme and start breaking up functionality that logically belongs together and can't be reused in other contexts. It makes it pretty apparent that writing decent readable code is hard (for many people anyway).
I know you said you're in the C++ world, so I don't think what I'm doing fits yours, but I've responded to the same behavior in the Java world in a way that tends to confuse and mystify my co-workers. In Java I've started to declare classes within methods. These classes can compute complex processes, but are nicely tucked away in the only logical place they make sense, that method.
My reason for this is that the method really belongs to the containing class. I can't break it down or abstract it out. I could use composition, but the new method-class isn't usable by anything else. The sub steps in the method aren't usable by any other method within the housing class. I could put them there, but they would just muddy the outer class. So now the code looks something like this.
public class Outer { void doSomethingComplex(final params) { class ComplexThing { private void step1() {...} private void step2() {...} public void goToTown() {...} }
new ComplexThing().goToTown();
}
}ComplexThing gets access to the outer params. This cuts down on constructor complexity and cements the intricate bond between the two entities, method and inner class. The logic is broken down nicely. Finally, nothing is needlessly shared.
This might make the outer class a bit larger, but I find this structure pleasant.
It's also a useful pattern for writing tests. Each test method gets a class (could have inheritance) that has a setup, test, tear down. All of that is encapsulated in one place.
I drew a table on a large sheet of paper, and recopied the code as a state machine. When I was done, it was clear where the problem lay: the state machine was woefully incomplete.
I rewrote the thing AS a state machine with all state-event pairs handled. Voila! It worked. Had to go back and diddle with some events (better way to handle parity error during checksum calculation etc) but at least I could find exactly the place in the code to do that.
So more, smaller functions winds up being a net win (up to some limit, of course...)
To that end, I would think the more segmented the call trees the better. But this is often hard to see when looking at a single function/data structure. This also seems to imply that having duplication of a function is not as evil as it would seem, since it can free up many of the dependency lines you have to consider.
All of this is to say that I do not disagree with the premise; though I do think it is less clear cut.
You should read it..
PS. It's the book Code Clean ;)