The Narcissism of Small Code Differences (2008)
raganwald.com
raganwald.com
Those who fail at this will stop working as a programmer, move out in the woods, build a cabin and live happy as a farmer.
When you're young you will inevitably always have an attraction for styles that ring the most true to you. You have one single hammer you have learnt to use or even worse, have heard really good programmers at Hacker News prefer to use, and therefore you use this hammer for everything.
You are a poser.
Most programmers have been there. Posers can be extremely sharp and useful if they get to do what they're good at, but they are posers nevertheless and at the start of a humbling journey to agnosticism - getting shit done instead of bickering and posing.
There are, on the other hand, a lot of bad ways to do programming: FactoryFactory nonsense, software-as-spec systems, waterfall methodology. We've seen them several times in our careers and have a hair-trigger sensitivity to stupidity, because we've seen it cripple or kill projects.
Great programmers tend to be unforgiving in their condemnation of the bad ways of doing things, but hesitant in accepting one programming model-- even a very good one like Lisp-- as being the One True Way of doing things. As soon as you have a One True Way, some very smart people disagree with you-- and that's a good sign that you're at least partially wrong.
In other words, it's good to be passionate about using the right tools for the job. It's a problem when people think the same tool is right for every job.
Calling these unequivocally bad is counter to the argument you're presenting (which I happen to agree with).
I'm no fan of either, generally, but I also don't work on "move fast. break things" type projects.
I definitely do not work on "move fast break things" projects.
Trying to imitate them does not make you them. It makes you a poser.
The point is that we each have our own different style of doing things, and frankly, I'm a bit incensed that you're using this post to hold up your way of doing things as the "one true" way of doing things while calling everyone else a "poser". I think this comment is exactly what the author is trying to say is bad.
On monday, we said "hey, there's a new bug - when started in Server mode, this code still tries to open a connection upstream! It looks like a Client class is getting created when we don't need it! ... and hey, where did our code go?"
The Purist said "oh, well, that's a weird thing to want and isn't supported by the framework, but I can put a kludge in here...." The kludge is, unfortunately, too weird to describe in english - he played a game with lazy loading and well-when-this-happens-then-that-makes-this-happen. And the rest of us had better fights to fight so we just let it stay.
When I see terrible code and set out to fix it, I try to change it into what the author probably meant, rather than how I would have done it. And usually that's a lot faster, because I don't have to rethink and rewrite everything.
And sometimes the ugly approach really is better. I've introduced tight coupling where loose coupling was causing problems. The components were never going to be used separate from each other anyway, and this removed a whole pile of complexity that we didn't need. The simplest way is often best.
When I rewrite code, at least I make sure the result is simpler than whatever I started with. If not, I simply revert back to the the previous version.
But the agnostic developer doesn't even consider refactoring it, because it "just works".
90% of devs are true agnostics, never refactoring, never unit testing, barely above the fizzbuzz level. Raganwald's example isn't an agnostic at all: he's a faithful parishioner, dutifully worshiping at the Church of St. Agile of the Clean Code but wisely avoiding the fanatics and the church politics.
If the argument that form is irrelevant as long as everything works, well then I disagree entirely.
Just working doesn't mean it's good enough, and really the big reveal in the post is completely beside the point because all those rewrites did not work but that's not what the author is criticizing.
Code has to be maintainable, it needs to be readable, and that matters.
Also it's pretty much cliché at this point to trot out the straw-man turns everything into a verbose classes programmer. I've never seen anyone turn a simple one line function into 3 files of classes with tons of unnecessary methods ever in my career. In Java you do create classes for a single function sometimes, but that's because you have to.
I think he is criticizing the desire of people to force their ways of programming on others.
Code shouldn't be littered with half a dozen ways to do the same thing. It reduces maintainability and readability.
Use a consensus or a person in charge to pick a way to do something that can be done many ways, and then always do it that way, and yes force people to program that way for the sake of the project.
Some people, sometimes very skilled and extremely intelligent people, are just entirely process adverse, seeing it as needless as long as everything works. Yes working code is important, but being part of a team, and working with the team is pretty important. Maybe you don't need the process, maybe you can read code like Mozart writes music but others can't. But maybe think of your team sometimes who have to follow in your footsteps.
There are layers of trust.
At a fundamental level I trust that my compiler works. I trust that the the standard library is going to work as expected.
At a higher level, up in userland, I trust that libraries or modules that are used in a lot of places or that were written by respected members of the community. If I don't know them I can look at other things, like tests, demos, and documentation.
Only when there is no other information, when you don't know of the person or their work, when you see no tests and no documentation, will you need to actually go in and look at the code.
I take two things away from this.
One is that in organizations where this is happening there are fundamental problems with co-developer trust.
The other is that monolithic code bases where an organization owns the code rather than individual authors owning the code can lead to an atmosphere where no one and everyone feel that they have the say over specific implementation details.
We can design systems and workflows that can implement and support developers trusting one another. If a team is building a system, components could be authored and published under individual author ownership, and then stitched together by the team to build a product.
Of course this is fundamentally a transcendent problem. The issues aren't with the details of the system, the issues are outside of the system and with how the PEOPLE relate to one another.
The author isn't criticizing so much as observing. From TFA:
The post is not about the Agnostic or his code, it’s about the dynamic of programmers eager to rewrite code in their own image, and the hypothesis that our (I am equally guilty of this behaviour) motivation for doing so is to emphasize the small differences between ourselves and others.
I love how people missing the point actually reinforces the point. Well done.
2. I think you complicated your point by having the other programmers rewrite it incorrectly.
I realize the context of your article requires a coding example that is somewhat trivial to understand and pick apart by developers who have a "better way", but I also think it's a copout to claim that the programmers in your example are just refactoring code for the fun of it.
A lot of people are rightly fixating on the formatting "feature" of not padding 2-digit zip codes. If I was on that project I could easily see myself wanting to reuse the zip-code formatter and running into an issue where my use case requires the 2-digit zip codes to get padded.
On most projects a zip-code formatter shouldn't be such a big deal that it requires a meeting between programmers. If the original zip-code formatter had bizarro requirements then they should be made more explicit.
For some reason the article avoids the biggest mistake in the whole process. Why did the other programmers make changes without ensuring the test cases passed?
Sure, we could say there should be more tests, or people should run them. We can say there should be issues filed before you refactor things. We can try to introduce all sorts of processes for avoiding mistakes. And we'd be right.
But it's also nice to talk to a human, even if just casually. I don't want to say "shouldn't X without Y" because then "shouldn't make changes to other's code without talking to them" is much the same thing as "shouldn't make changes without running existing tests" and "shouldn't check changes in without submitting tests that the code fixes."
Shoulds and shouldn'ts should always be secondary to talking.
Talking to the person who wrote the code you want to change is always a good idea. It's not always possible (because the original coder has left the company or something), but if you have the opportunity, use it. You'll learn something.
I do this all the time. Sometimes the answer is: "You're probably right. Change it." and sometimes it's "It's like that because of good reasons X and Y, and we also use this in another project that needs it to do Z." This understanding is very valuable, and can save you a lot of time.
While one may or may not agree with Freud, the article is at least food for thought. Presenting software conflicts in this way asks us to acknowledge our instinctual psychological motives and question our claims of objectivity, technical soundness and pure logic. This kind of in-fighting serves only the individual and not the team.
Personally, I think there is great value in having a plurality of ideas and voices, but in a commercial setting, they must ultimately serve the team. This creates tension because individuals must be able to restrain their impulses.
For his next blog post, maybe Raganwald should write about the pitfalls of using allegories.
Granted, I could have written the comment a little better, i.e. more context as to why set semantics were an issue. My co-workers later told me they found my change and spent 3 hours or so trying to figure out if I was right or not, eventually concluding I was. Of course, I was just down the hall and neither I or them thought to go ask the other about why things had changed or if it was a valid change in the first place. A little communication could saved a large amount of time ... Ah well, live and learn.
In the article, author states that existing unit tests checked that two digit zip codes were invalid. The question remains how a refactor that fails unit tests was pushed to master.
Because rewriting code that looks ugly is one of the ways people use to begin to identify with the codebase and make it easier to understand for themselves. The example in the article is clearly made up to show how stupid it it to change things that "just work". Except in reality those things are usually ugly, large and complicated, and eventually need to be changed anyway. The way to deal with them is make people responsible for their own code, so that they add tests or at least some comments that explain why and how things are the way they are. Simply being fearful of change and blaming the last person who touched the codebase for all its problems is not a viable long-term strategy.
Surely you must see that if more than one person is working on the same codebase that this approach does not scale. If everyone on a 6 person team is constantly rewriting the code so that they each understand it better, and each of them has a different way of understanding code, well, that's going to be quite the clusterfuck.
If there really is a problem it will be fixed. If it is not considered a Problem it will be archived and documented as "NOT a problem".
And if the 'fix' breaks the system you will not be toasted because hey we talked about this and all(including the boss) agreed we should change this.
The true problem here lies with the poor choice of function name on the part of the original author and lack of commentary. The name "formatted_zip_code" automatically primes people to expect that you passed a (possibly mal-formed) zip code and the function should format it properly. Under that metric, the original code was indeed deficient and should have been changed. Whether it retained the original style is secondary, but there were clear functional deficiencies.
"Anything with fewer than three digits is supposed to be an invalid code" is not reflected in the code (why does it return digits and not trigger an error there?) and apparently there are no comments, so it should be no surprise that it needed to be fixed.
I've seen this kind of code time and again from coworkers who have the "code should be self-documenting" mentality, who refuse to put in comments explaining why they wrote a function a particular way. But what they think is self-documenting, very, very, very rarely is to others.
The whole tale is just "if it ain't broke, don't fix it."
I struggle between this and the accumulation of technical debt. Sure, it's not broken... but that doesn't mean it can't be improved.
def formatted_zip_code(digits)
case digits.size
when 4 then "0#{digits}"
when 3 then "00#{digits}"
# 2 or less digits is an invalid input
else digits
end
end-> if the argument doesn't conform to contract, an assert/exception should be thrown so that the caller can recognize the existence of an incorrectly used function.
Normally I'd agree with your argument about failing to meet a contract, but why would this formatter method get anything but a valid zip code.
Ask not, lest ye receive an answer.
Seriously speaking, you can't trust that callers will be sane unless you establish a priori with a higher level contract that all possible callers will be sane.
I'm not at all sure about my classification. Probably somewhere between Librarian and Purist, perhaps (but these things feel somewhat harder to apply when most of the code is in C).
Should I fight the urge to re-write code that I see as broken, since it's just me wanting to re-shape the code in my image? That's emotionally somewhat difficult (bad code hurts the eyes) but certainly doable, but it also seems a bit counter-productive since I really think there can be some value in sharing experience. It's also, of course (?) hard to really believe that it's all narcissism, to me bad code feels objectively bad. If that even makes sense.
I guess it boils down to "co-operation is hard" for me, right now. :)
(Note: I am a user of many ancient and venerable old FORTRAN libraries, which I have no intention of rewriting because "there is a better way." Yes the glue code is a pain, but you don't need to write it very often.)
[1] or passes the tests, if appropriate.
Understand why other people write code like that. If everybody is constantly "fixing" everybody else's code, you're never going to get anything done. Understand the differences and find some common ground. That's the only way to move forward.
This.
But maybe a cynical, world-weary Librarian like the discworld one, as touching someone else's crappy code without some larger plan in mind seems destined to cause you heartache, following the standard "person who last touched it is to blame" rule.
I've had to fire people because they could not stop "refactoring" code that wasn't broken. Their commits were half in the code the were supposed to be working on and half scattered, random changes like the changes from the article.
You're not "sharing experience"; you're wasting time, money and energy. You are insulting your co-workers intelligence. You are pissing off your boss.
Worse, your co-workers now have to merge your random edits into their work because they were actually assigned to work on the section of code you "fixed".
There is nothing worse than doing a pull and finding a bunch of small, conflicting and ultimately useless changes on the head, ruining your morning coffee.
I certainly hope I do whatever I do with a bit more finesse than what you imply. So far co-workers do not seem to be offended, and no boss has ever seemed pissed off to me.
From your remark, it sounds as if you consider "refactoring" to be almost a swear word, and something that should never happen. I just can't agree with that, even with code I've written myself I sometimes realize I did it wrong.
I practice commenting on other people's code a lot online, that seems to work pretty well, too. Of course it's even harder to notice if I'm offending people online.
Ha! Not really, in my career I think I am still "net negative" for lines of code; I once refactored a 100K line medical diagnostic program down to ~10K lines for instance. What is a swear word is people who use "refactoring" as an excuse to start dicking around with stuff that doesn't need to be fixed.
Refactoring certainly has it's place, but needs to have a purpose: new feature set that requires the rework, performance, what have you. Too many developers just pick something that is working and "refactor" it ... then you end up with the same thing you started with, just "better" and with different bugs.
My perception of "beauty" is the simplest code that could possibly work, with the lowest possible levels of abstraction. I like YAGNI and similar principles. That's a pretty vague formulation though, so while most can probably agree on a high level, the devil is in the details. What exactly is "simple code" for example? Short code? Easy to read? Flexible? All of the above (that'd require Carmack or Thorvalds)? It really is all in the eye of the developer.
I think the takeaway of this excellent article is that whatever your style is, do have respect for differing styles. Don't grade developers or code on what you intuitively feel. Don't assume that everybody else is an idiot and finally don't waste time on debating / enforcing the one true way of doing things.
>>> (datetime.datetime.now() - datetime.timedelta(days=2007)).__str__()
'2008-05-14 18:32:57.195554'
Looks like this old goldy originally turned up when hacker news was just a yearling...why would you call `__str__` explicitly instead of 1. using str() as you're supposed to or 2. just using print?
Although, amusingly, I can't actually find this rule in PEP-8.
This made me think about something(and perhaps it is a horrible analogy), but how would you play a game of chess if every few moves you had to let someone else make your move for you? Would the strategies that worked and allowed you to win by yourself in the past start to fail with this new constraint? Would you have to abandon certain types of plays or would you spend your time cursing the Gods because the other players couldn't see your strategy nor you theirs (even if you tried to discuss it)? Or perhaps you would realize that the problem is that chess is simply not a game that is meant to be played this way.
In the end, we had more success using clearly defined boundaries and trust rather than a collaborative process based on overlapping responsibilities and communication.
I'm not making a universal claim, I'm just telling my story.
I prefer code to comment in this case. ;)
01234,Hello
import it into Excel, edit something and export it as csv again. Unless you pay careful attention to what you're doing, you've just lost your leading zero.I think the point of the article isn't to compare these different styles, but to deliver a bigger message: Style isn't that important, just write code that works.
That does not really belong to OO. You would be unlikely to find an AbstractSingletonBeanFactoryStrategy in a Smalltalk project for example (even one under the control of a Purist.) You would find it in a JavaEE project, but even then it is not part the platform, it has more to do with convention that has built up around the platform (and the type of Purists it attracts.)
The example used as a demonstration is a bit like an episode of Three's Company. The premise revolves around a simple misunderstanding, and it's a bit hard to believe that the misunderstanding went on for so long. The original code is a prime candidate for documentation, and there were even unit tests. If you accept the misunderstanding at face value though, the lesson works well enough.
But still, if someone writes code like that and I'm not the one reviewing it, my instinct is to assume there's a good reason for it to look like that. If I'm reviewing it, you can bet I'll want them to explain why they do it like that, and more importantly: explain what's so wrong about the original code that it needed to be changed.
Indeed, imagine a world in which authors would publish in a magazine both their own short stories and a review of a short written by another author for the same magazine.
Unlike our authors in this Magazine From Hell we programmers have objective criteria on which we can base our critique of the work of others when we have to, namely a specifications, coding standards, etc.
As a perfectionist I'm more often than not engaged in rewriting code "the right way". Generally I've found it hard to accept that there are so many ways of doing things. Worse, some solutions perform better while others are easier to read. Even worse, there's usually a trade-off between those two.
To address this I've came up with two ideas. The first would be a language that would be ultra-restrictive, so that there could be only a limited number of ways of doing things. However my guess is that this has been tried and failed.
The other idea would involve some clever IDE and/or version control system, that would allow different versions of constructs to co-exist. In other words there would be different "views" on any piece of code. In fact, this is already the case with documentation which can be seen as a "natural language" view on the code.
This would solve the "The Narcissism of Small Code Differences"-dilemma, as every programmer could keep his favourite version. But what's more interesting: Based on the assumption that rewriting code for ideological reasons is common behaviour that can't just be stopped, it would be nice (and more efficient) if the rewritten code could at least serve some other purpose. As different versions of a function exist, they could be invoked based on some criteria (probabilistic or as a fallback) with the intent of increasing fault-tolerance.
Of course having a fallback is part of the motivation for version control systems, but these are not able to utilize different versions in a systematic way/at runtime, at least not as far as I know.
A few reasons I've seen where assumptions about how zip-codes work lead to not taking money:
Ireland - no zip-codes (outside Dublin) - dont do mandatory fields
UK - numbers and letters - dont assume only numbers in zip-code
Brazil - 9 digit codes - dont assume max of 6 or 8 characters
So then why if your code is outputting "57" to the input "57" while the librarian's cleaner and more reasonable solution gives you "00057", why is yours more correct? Shouldn't the function work something like
if digits.size<3 raise 'Invalid zip code'
digits.rjust(5,'0')
Its a craft yes, but really its also a logical profession. Strongly held opinions suck all the joy out of conjuring bits to do your bidding. Can't we all just get along?
I'd say just the opposite. If you care about programming, you care about doing it the right way. The people who don't have strong opinions are generally the people who don't care, and don't enjoy their jobs, and don't produce good results.
The solution is to have strong opinions, weakly held.
when you do so you introduce a risk of breaking it through lack of understanding.
more generally it introduces a potential point of failure when you refactor code. as with all things that introduce the potential for bugs and errors it should only be done if that chance seems less than the chance that it will fix a known issue of greater severity...
Usually, it's because the simplest solution is the best one.
This is what comments are for, people!
People not documenting the code's intention bring down the whole project. Sure, sometimes the code _may be_ right, but without documentation, you can never be sure... It inevitably results in a trust-no-one environment where you can't change a single line safely because it might be doing something else: the weird side effect might be used in some other module, or it's sanitizing data from library X that's no longer used, but nothing is documented anywhere, so you never know.
Anyway the others changed the algorithm, interpreting what they believed it should do. That's something unavoidable sometimes due to convoluted code and lack of documentation, but not in this situation.
I don't think any sensible person from any paradigm would have objected to the code as it was originally written. The substantive arguments between paradigms happen at larger scales of code organization.
So it comes across to me that the "agnostic" is trying to make other schools of thought look ridiculous. The trouble is he's done it by cheating.