Why clean code is more important than efficient code
m.techrepublic.com
m.techrepublic.com
Without reasonable efficiency, nobody would want to reuse your code. In fact, one of the hallmarks of reusable code is that it is efficient in a wide variety of circumstances; frequently the author of an API cannot predict what it will be used for, and thus must at the very least avoid pathological inefficiency when possible.
> I was not thinking about performance, about the efficiency of what I wrote, at all; I simply did not want to think about it at the time.
Bad software happens when you don't think about the efficiency of what you're writing at all. This perspective is completely detached from reality; in the real world, your code runs on a real machine and does real work. Sure, in some circumstances it's entirely appropriate to write inefficient code, but that's a decision that should be made after an assessment of the problem at hand. To intentionally ignore efficiency as a matter of policy is to ignore the facts of reality.
At the very least, you need to consider the performance impacts of the data structures you use. In most modern languages, it's not terribly difficult to use e.g. a hash in place of an array, and in the right circumstances it can make the difference between usable and unusable software. I can't even count how many tools I've used that, due to a poor choice of data structure, could not scale beyond the small data set that their authors tested them against, and were thus useless beyond a very limited scope.
I was working out a prototype recently and we're turning it into production: it's falling down hard in the 'efficient' category, although it's really easy to read.
If I had thought about efficiency more when prototyping, I wouldn't have had to do the rework that is now looming.
Efficiency matters. Speed matters. Let's not delude ourselves that they don't.
There are still reasons to pay attention to efficiency when writing code.
Did you even get to the second sentence of the linked article before launching your ridiculous rant? Obviously not.To suggest that nobody wants "reasonable" efficiency anymore is absurd. To do so when the article actually argues your point for you is painful.
To suggest that the people who favour readable code are proposing "pathological inefficiency" is insulting.
> absurd
> painful
> insulting
I tried to explain my position on the topic in such a way that we could have a conversation about it. Please, argue to the points I tried to make instead of slinging inflammatory remarks.
If you want to have a discussion, don't start by putting a ridiculous set of trousers on the straw-man you made.
My argument was not a straw man. The original author specifically said that he, during a specific part of his development process, "[...] was not thinking about performance, about the efficiency of what I [he] wrote, at all [...]". I understand that he circled back on the performance issue later, but in my view, it makes more sense to start with an integrated performance/readability approach than to achieve both in multiple passes.
Someone who has been writing software for a few decades will be starting out having already implicitly rejected bad algorithms, unmaintainable designs, unscalable data structures, and poor security practices. Years of development has given them an intuitive feel for the best starting point for an efficient, scalable, and secure design.
Not to mention that anybody with sufficient experience knows to figure out the overall design first before cracking open the code editor. As such, their focus should be on keeping their code elegant and maintainable because it will be built on top of a reasonably performant architecture. And when they discover areas which need greater efficiency, the ease of refactoring the well designed architecture is an inherent bonus.
Clean code requires less resources in terms of maintenance, bug fixes, and extensibility, because it's easier to understand and modify. Efficient code requires less resources in terms of server maintenance and bug fixes, because it actually requires less maintenance and is less likely to catch you off guard by taking your server down.
I've spent enough time diagonalizing 10,000 by 10,000 matrices to get frustrated with anyone who claims that efficiency in time and memory ain't important any more. Maybe not for your code...
So often times coding efficiently is living giving yourself a bigger server for 'free'.
As soon as their code is made more readable, not only does this make life easier for the next guy, but it also makes it a lot easier to optimise once it is clear what it is actually doing.
People who don't want to write readable code seem to subscribe to the belief that there is only one way to write a particular line of code, yet they could not be more wrong. For any non-trivial section of code in any non-trivial language there are dozens, if not millions of different ways of writing it. Given that, making it unreadable is either maliciousness or incompetence, and Hanlon's Razor tells me it's probably not malice.
Moreover, most of the time the poor sod who is the "next guy" to read/maintain the code is you. You are the next guy. You are the audience of most of your coding efforts. If you cannot read your code 6months down the track, then the only person you're shafting with your bad workmanship is future you.
You can have your cake and eat it to. Code can be both efficient and readable.
Numerical code is actually one of the worst example you can find, because first the API is very clear, and there is little chance to have a change of spec later, making it efficient is generally purely a matter of algorithms + implementation. It is obvious the OP is talking in a context of web apps and co, where this is not true at all.
Understanding a problem is more important than writing a solution that is clean in itself. However, we often don't understand a problem immediately. Writing code (and seeing the results) is one way to quickly gather data about the problem, and so help understand it.
This is where clean code can help: code that clearly communicates your current model of the problem is very helpful part of the process. After a break, it reminds you of the model; someone else inspecting the code can see the model; it is easier to see where the model is incorrect; it is easier to modify the code to correct the model; when the problem itself changes, it is easier to modify the code.
However, there is a trade-off between effort and clarity: it is a mistake to spend too much time clarifying something that might not be correct (or, if the problem will change soon, making it incorrect). Just write code that reflects your understanding of the problem. Regarding efficiency: if you know how to easily make it more efficient without compromising that (e.g. without making it convoluted), then of course do it.
Code is not a final product, but a process. The faster you can iterate, the better it is. (An exception is if you already understand the problem perfectly.)
I think the problem is that too many people disagree on what is 'too much time'. Not trying to nit-pick but obviously it is a mistake to spend 'too much time' doing anything. That's implied in the too much. That may come off as being pedantic, but I don't mean it to be so. I think defining 'too much time' is right at the heart of the problem. And it's a hard problem to be sure.
I'd also disagree that the possibility of something being incorrect is a reason to spend less time clarifying it...I'd say just the opposite.
By clarifying code you're making it easier to comprehend and thus change in the future. I'd say you should write code with change in mind.
Regarding 'incorrectness' itself...that's the only reason code is ever changed. Whether the code is fundamentally incorrect (the type I believe you were referring to) or temporally incorrect, the incorrectness of the code is always the reason for the change.
As for the exact trade-off: when you are iterating quickly, it is not effective to try to nail down these trade-off precisely - you need those cycles for actual coding!
For that matter, it's also difficult to define what is "clear" - or even what is "simple". Have you ever tried applying the "Single Responsibility Principle" thoroughly? It's amazing how many genuine "reasons to change" you can find - and amazing how less clear it can be. Another example is dogmatically writing "short methods". Clarity, like "simplicity" - and like the trade-off you note - is itself extremely hard to pin down precisely.
BTW: reasons to change other than correctness include: making it more efficient; making it clearer (e.g. refactoring); making it reusable (Brooks' says it's x3 the work). And, in practice, there's a continuum of correctness, not true/false, for "better" (more accurate, more reliable) results: e.g. more accurate search ranking; better collision detection; better webpage rendering; OS that crashes less often.
I put the word in quotes to try to express a much wider than usual connotation for the word (air quotes, if you will..). It's a shame that language is all we have to communicate on the internet.... :)
I also agree that we shouldn't precisely define tradeoffs before continuing on a single, specific, problem.
I just think that, in the macro sense, we can and should get better at defining the factors that would inform such a precise weighing of the tradeoffs for clarify, efficiency, expediency, etc... Doing so would give us better instincts for single instances where doing a precise comparison would be absurd.
But I think there's a fundamental barrier for software, because programs can differ from one another arbitrarily. This makes it hard to generalize - and generalization is the essence of theory, rules, systemization, modeling etc. I assume you'll accept that generalization is crucial, but you'll disagree that software can vary arbitrarily.
The reason I say this is because if you find regularity in software... you can automate it. Factor out a common method. Extract a "pattern" into a framework (or better, in functional programming). Derive one thing from another. It's like compression: as soon as you notice something that you can generalize, and construct a theory to model it... you can use that theory to eliminate the redundancy.
In a deeper sense, programming itself represents theories. And theories (e.g. of physics) are not generalizable: we don't have a theory of theories, that enables us to automatically generate the next correct theory. Oh, we can automatically generate theories, just not correct ones. For that, we need data to test them, to find out which one is correct. And so the unknown is the ultimate cause of the problem: we can't generalize what we don't know.
The best you can do is come up with rules of thumb, that help in specific domains - but they only work there. And if they are really accurate, someone will encapsulate them in a method, a framework, a library, a language or operating system, so that programmers don't need to deal with them, and that pattern disappears from their programs.
And to finish off with one of my favourite quotes, by Alan Turing: There need be no real danger of [programming] ever becoming a drudge, for any processes that are quite mechanical may be turned over to the machine itself.
NB: in practice, there is an enormous amount of repetition in programs (and parts that are common between different programs). This amounts to ongoing demand for libraries, frameworks, languages, operating systems, and extra features for them.
So, it's hard to obtain a precise weighing of the tradeoffs for clarify, efficiency, expediency when the playing field keeps changing. We can get better, but not much. It's different in fields like civil engineering, where materials and techniques don't change often. BTW Another fuel for change in software is Moore's Law - which enables different techniques to become favoured that weren't plausible before (like dynamic languages, VMs, etc).
sorry, such a long and self-indulgent response. Maybe it doesn't apply so strongly to trade-offs (though I think clarity is closely related to theory). At any rate, it doesn't stop us improving, which is what you suggest.
This reminds me of an approach I took on a recent project, where I wrote a variety of ultra-simple routines to look up very specific numbers and results. These routines would often go out and fetch a rather large file of data, and crunch that data down to the single result the caller desired. All in all, it was a highly inefficient way to do things. However, it was not a problem, because all I had to do was drop in a few caches here and there inside the routines, and voila, the whole thing became fast enough.
I found it a very liberating approach, because all of my code looked very "obvious" in its logic. In a few scattered places you see a couple lines of code to look something up in the cache first, but then it drops down into the ordinary and extremely clear logic.
The code also had a very nice "lazy", "on-demand", and decentralized feel to it, with no need for a bunch of gnarly and centralized initialization code.
And given migration of large clusters from one piece of code to another is expensive you want to think about performance ahead of time.
And in that scenario engineering time may be less expensive than tons of iron, the passing of said time is a killer. In most software project, time is a bigger issue than money.
After that, I'll worry about correctness and efficiency.
In the mainframe days, we used to have code reviews that forced programmers to structure things properly, or they'd hear about it in the reviews.
The most reasonable approach is to have clean code with a little sprinkling of efficient code where it matters. So if you've built a nice cocoa app you don't need to hand code in assembler each little part of it. Just that main image processing loop. Even in that case, document it well.
Now let me go a little further and say that clean code and efficient code are not at all mutually exclusive. If you find a programmer is producing efficient (dare I say even fast), but un readable code it may be time for them to find new employment.
It's not really hard, keep your Knuth/Dijkstra/Rivest handy, and refer to them when needed.
Sometimes it actually works ;)
Also I don't believe in the idea that the enhancements of hardware are to put up with lousy code.