Anti-Patterns Every Programmer Should Be Aware Of
sahandsaba.com
sahandsaba.com
And it's exactly this requirement of masterful manipulation of balance that makes me love this field so much and long to get better at it. Getting a handle on the artsy spirit of programming is really what separates the wheat from the chaff in terms of programmer skill. There is no formal programmer guild in real life, but if there was, and there was some sort of test a journey-man programmer must undergo in order to become a master programmer, it would be the test of being given a large project and then deciding correctly exactly how much technical debt to take on to be able to ship a product within a reasonable time-frame and yet have its internals not be complete unmanageable spaghetti.
It seems to me, that that saying works for only the small set of tasks that can be worked on in consistent overtime. Can someone clarify?
I'm usually a little meaner and tell people to pick 1.7 of the above.
It's the one attribute they can't see.
Good clients will understand and bad clients will filter themselves out quicker (which is a blessing in disguise even if they do sometimes give you money).
I prefer to let cold, hard reality do the teaching.
These are really just proverbs, and it is not unusual to find pairs of proverbs with opposite intent - e.g. http://www.wolaver.org/WordPlay/OppositeProverbs.htm
They have some use as starting-points for discussion, but identifying when one (or which of a complementary pair) is appropriate requires good judgement, and the people who get dogmatic over these things are missing the point.
I think the fact that these sorts of argument feature so much in the discussion of software design and development indicates that there is some way to go before it becomes a fully-mature technical / engineering discipline.
Or, taken to the limit, "Moderation in all things, even moderation."
(It may take a moment's thought to see how I got from here to there.)
I've seen this a lot, and is often made worse by attaching a load of unnecessary baggage/repurposing to classes.
For example, I've worked on PHP projects which, over time, have gained coding standards like:
- All classes should be in separate files, named to match the class name
- All classes should be built from a PHPSpec specification
- All classes should have their own test class
- No class can instantiate another, unless it's a "Manager" class which does only that. Everything else is dependency injected.
- Test classes can only instantiate their associated class; everything else must be mocked
And so on.
Now, each of these has its merits, but as each new practice was added, it increased the 'fear of adding classes', even if just subconsciously. Refactoring to break up classes became less and less common, since a simple code change would require:
- Whole new test classes (OK)
- New files (seems a bit like Java-envy, but OK)
- Injecting new instances wherever the old class was used (seems a bit overkill...)
- Mocking the behaviour of the first class in tests for the second, and vice versa (this is getting a bit silly...)
- Creating 'Manager' classes to instantiate the new classes (hmm...)
- Creating new test classes for these managers (erm...)
- Mocking the dependencies of the managers...
In the end, it was just far easier to just throw all new logic into existing classes, which grew and grew.
Other fun anti-pattern: make a Facade class hiding services. Then:
- add logic directly in your facade
- make the services using the facade depend themselves on the facade (circular dependencies FTW)
This an excellent way of ensuring your code will NOT be reused.
Another far more common antipattern that isn't mentioned here is premature abstraction. You see a lot of this in enterprise .NET codebases -- such as the widespread "best practice" to build ineffective and obstructive abstraction layers "just in case you might want to swap out Entity Framework for a web service."
So I consider it the effect of the"culture," "environment" and "organization" more than the particular language. It didn't help that the "patterns" were introduced exactly as the unstated promise to the managers to give them a "meta" approach. Like UML before etc. And that people started to think the more abstractions they force in the code, the better it becomes.
My favorite problem description is by Rico Mariani:
"The project has far too many layers of abstraction and all that nice readable code turns out to be worthless crap that never had any hope of meeting the goals much less being worth maintaining over time."
http://blogs.msdn.com/b/ricom/archive/2007/02/02/performance...
But in my opinion all big enough projects can be saved by redoing just some part of them.
And Java codebases
One of the most important steps in programming is:
How many elements could this code actually have to process, and if there could be many: is this datastructure/algorithm suitable for that?
This is almost never premature. This should NOT be called optimisation.
I would like an own word for that. Suggestions?
To answer your question how does thinkthencode sound?
Also, I would restate the rule as something like "don't spend too much time prematurely optimizing. If you're building a lookup, it's obviously a Good Thing to use a hash instead of a linked list, and in 99% of cases shouldn't take you any longer to code.
He does address this one in point 9:
Useless (Poltergeist) Classes Useless classes with no real responsibility of their own, often used to just invoke methods in another class or add an unneeded layer of abstraction.
Over-abstraction is extremely common in 'Enterprise' Java codebases too.
It's the same idea as the Knuth quote, but gives you guidance over what to optimize for (rather than just leaving you with the idea you shouldn't optimize for anything). Beautiful code will oftentimes include an elegant, efficient algorithm. Many times I've found that when I've had a performance bottleneck, the fix also made the code more correct (not just on a performance standard, but things like "Oh...yeah, we totally didn't need that there" or "tweaking that to be faster also meant we just fixed a possible, if unlikely, race condition"), or more beautiful ("oh, we just removed a duplication of effort...that when called this way was turning an O(N) operation into an O(N^2) operation"); had I just adhered more stringently to those two, I would have had the performance already.
This is still preferable to premature optimization.
What I really can't stand is developers who tell me that this piece of code that already runs in production with no complaints is 'slow'.
And I ask them "well, did you profile it?".
"Well, no, but..."
Some optimizations don't matter anymore. Back in the hoary old days pointers in C were faster than indexed arrays. Not anymore. But people still assume it's worth it. Twenty years ago, you paid a serious price for unaligned accesses. Not any more. Yet compilers still rigorously align data in memory. Which might actually make things slower.
As the follow on poster said, make the code clean and easy to reason about. People often like to think they are writing code like the big boys, where it'll get used by vast number of people. Much like a lot of hardware designers design stuff as if they are going to make millions of units. 99% of the time that's not true. Thus the NRE costs totally dominate.
Find out whether sacrificing clarity for speed is even categorically worth it in this particular function, and find out whether your alleged speed improvement is proportional to its readability costs.
"Not Invented Here".
The urge to rewrite things that you encounter is strong in software engineers, and you should always be suspicious when you find yourself thinking "I could do this so much better." Especially when it's true!
At some point, it's important for us as programmers to indulge that "I could do this better" instinct. If you're right you could make a real contribution to your employer or the community. If you're wrong then you'll still be a better engineer (and possibly a domain expert) for the effort.
If everyone embraced "not invented here" as a philosophy then the whole industry would stagnate.
And sometimes I look at components and I'm so disgusted by the code or the API that I merely use it as an example and re-write the code for my own purposes. Just because somebody put it on Github or built a package doesn't automatically make the code good.
Most of the "not invented here" I come across is from developers who never look for 3rd party solutions. It doesn't even occur to them. And it's usually obvious stuff like XML or JSON parsing! Which is why the message of "Not Invented here" is a good one to repeat over and over.
So, there is a problem where many developers would rather work on abstractions rather than concrete value-providing things. I suffer from this myself - if abstractions weren't interesting to us, we wouldn't be working in software.
However, over the years I've discovered that many software projects are gaining negative value from some of their libraries, and I'm not sure which is the larger problem.
http://yosefk.com/blog/redundancy-vs-dependencies-which-is-w... describes the essential problem.
There is one case when those wheels can be worth (partially) reinventing: if you only need them to parse some files which use XML or JSON in a very restricted way. Say you've got a bunch of XML files which only go one level deep, only use ASCII, don't use namespaces, etc. A full XML parser comes with a whole ton of bloat you don't need for that! Plus, if there's ever a problem with those files, the outsourced XML parser will tend to give the end-user absolutely useless error messages.
Also, whenever there's a lot of data being looped over, it can be very tempting to add new features to the existing loop, rather than adding a new pass (inefficient) or trying to collate the results after they've been spread out (complicated & error-prone). This can turn a simple parser into a core piece of business logic, and of course it's then only a matter of time before custom, XML/JSON-incompatible 'directives' start creeping in to control that business logic.
Worrying about the bloat is premature optimization. If your process is actually too slow and parsing is the bottleneck then by all means rip out the XML parser library and roll your own.
Especially since the latter seems to have a much deeper and irreversible impact on the engineering culture.
> If it's a core business function -- do it yourself, no matter what.
Sure, and that point is pretty easily defined: the point at which the existing solution has produces real, tangible, definable problems with achieving your current goals is the time to invent a replacement that "does it better", with better defined in exactly the terms of the problem that you are addressing.
Its also okay to do it as something exploratory, off the critical path of a real project for understanding/skill-building.
What's not okay -- and is instead a dangerous, expensive (both in the short-term and in maintenance terms), ultimately unproductive diversion of effort -- is making a reinvention of something that works adequately a dependency for some other effort.
That said, I do agree with the author that, in the context of a programming class, writing the 'LabStack' class is absolutely useless because it teaches you nothing about actually using a linked-list to implement a stack, which I'm sure was the real intent of the assignment. The 'LabStack' class is useful in the context of being used in a larger program, it's not really useful for learning how to actually implement a stack because it just passes the implementation off to 'LabStack'. (That said, because Java doesn't support defining objects as value-types, using the 'LabStack' interface adds an extra level of indirection you may want to avoid - You have to access 'this.list' to get the LinkedList object, instead of just accessing it directly. This is more a fault of Java then anything else though.).
The note about programming class isn't really relevant to programming as a whole though, which is where the disconnect happens.
Other things are also hard to avoid by myself, like
- God Classes (I could refactor here, but then the anti-pattern is fear of refactoring)
- Management by Numbers (I'm not a manager, what can I do about it? Actually I'm happy if my management is already that good that they have numbers. Nothing is more horrible than management who doesn't tell you explicitly what they want and then complains about details for half a year, then starts with the next project without declaring the last one to be finished or anything. I'd say numbers are good, finding the right ones is tricky, though.)
- Useless classes are avoided by simply not writing them? No, they develop when code gets refactored and nobody had looked at the responsibility of that specific class for some time. Nobody writes classes without having a goal for them in mind. Mostly calling methods of another class is fine. It happens in design patterns like Observer, Delegator, or Compositor.
Edit: I can't find it. Anyone else remember reading this? The mysql job queue reference was from a popular article that went out about it at the time
Edit 2: the job queue article is from 2011.
Edit 3: still can't find it. But I know every part of this article before seeing it ... I tested myself and remembered all the details... Where is this thing from?
Still can't find it. My confidence in the permanence of the web is pretty destroyed right now. I can't even find references to it existing.
This isn't the first time I wanted a searchable version of archive.org that effectively timecapsuled the interactive internet. Things fall off the net far more than common folklore suggests
Anyone got $10 mil to drop on that project?
Those enormous URLs that describe nothing but a document GUID are also a sin.
I can strongly recommend using such a tool to make sure you don't 'lose' stuff on the internet that is important to you.
Right now i am using chrome, to save whole page as zip, but that becomes in-convenient when you have several hundreds of them.
I don't understand how to avoid this even if you break it up in smaller classes. You still need a point of entry where the logic begins and where it is decided which components to use. ThIs always means some kind of Manager or Main class for me. How do I fix it?
It's important to keep this graph as acyclic as possible, and to try not to have any one component directly interact with too many other components, because then your abstraction is too fragile.
Here's a (vaguely) contrasting view from John Carmack:
http://number-none.com/blow/john_carmack_on_inlined_code.htm...
The main takeaway is that inlining a bunch of code to flatten the dependency tree may actually reveal problems and reduce your maintenance burden.
The worst for me these days is analysis paralysis. I'm currently working on my own so I don't have anyone around to bounce ideas off. For smaller tasks it's not a problem, but there are bigger design decisions that have ended up taking longer than probably should. When you're left to figure them out on your own it takes a lot longer to convince yourself of the "better" way of doing something.
I like to think that the best way to choose between seemingly equally advantageous designs is to start with the one whose first step(s) is(are) the most straightforward. The thing is, while I prepare myself for implementation of that first step, I've a background loop in my mind that constantly checks against other implementations choices. What I am losing here, what I would gain otherwise.
In the end, I never really make definitive decision before starting. I start with that background brain noise on the “most simple first step design”, and when the background noise stops and the raw pleasure of coding kicks in, I known I'm on a good track.
For any feature / bit of work that's in isolation, I don't really worry. I figure out an approach and I implement it. When it comes to changing data structures / larger changes within the product there's a fear of getting it wrong and having a mess to unpick later.
Great quote - It reminds me of the story of when Michelangelo unveiled the statue of David, someone asked him how he managed to create such a masterpiece and Michelangelo answered "It was simple, I just chipped away all the rock that wasn't David".
I don't think reduction only leads to perfection. If it would be, then our starting point must always be the right one. But in practice you start somewhere, reduce, move forward, add stuff, move sideways, try out and test. Sometimes you need to add stuff to make it clearer to your users.
- why waste time with creating a function when I can copy-paste?
- but I'm never going to have to change this code anymore
- but it's more readable else I have to figure out what the function will do, browse to it, ... (well, good luck navigating any codebase at all)
- but it occurs only twice (ok, twice is debatable, but still: before you know it the same is used again in another place an there we go again..)
- but it's faster because it doesn't require an actual function call (this is by far the worst one: they claim this without having even measured it, without having checked if there even is a bottleneck, without realizing that in compiled laguages any decent optimizer would inline and generate assembly anyway)
I don't know about this one...we need something have a chuckle (or cry) about over beer after work, and to make sure software archaeologists in 100 years time have an interesting job (not to mention plenty of opportunities to write blog-spam)!
I think the idea that "explicit is better than implicit" is a good idea for more than simply "magic numbers and strings".
I much prefer seeing the code in front of me than having to jump through multiple functions across several source files + tracking down the configuration values I'm looking for because someone decided to use DI.
I'm not saying DI is a bad practice, but unless you really need it, prefer explicit dependencies.
And on and on with just about everything you do, if all things are equal, prefer explicit over implicit.
My default these days is to write numbers as numbers, plus maybe a comment if things are unclear. I only turn numbers and strings into named constants if they are used in more than one place, or if they are true configuration values that need to be tweaked.
This makes the program easier to read (the ultimate goal) by avoiding the need to jump around in the codebase. Comments are better at explaining things than cryptic variable names anyhow.
On the other hand, tools like CheckStyle can make people do some strange things when they take it too literally.
I once reviewed code that defined: static final int ONE_HUNDRED = 100;
Actually, the whole example smells ripe for refactoring. Where are there multiple windows throughout the app all sharing a default starting size? Sounds like it's time for a base class or something. Once you've eliminated the redundant sizing code, you will find that your default size values now exist in only one place (the base class). Since they are only in one place, you can write them as numbers again instead of named constants.
I've played this pattern out more times than I can count. The problem isn't magic values sprinkled through the code; it's having redundant code in the first place. Once you solve that, the magic numbers tend to evaporate.
static final int ONE_HUNDRED = 200; // Changed to 50 to optimizeAnd it's funny because it's true.
And ... the funny could have been eliminated if the named constant had just had a useful name, or been replaced by a function, rather than a passive-aggressive response to a code analyzer or coding standards ("There, I fixed it."). The conflict between the name and its value would have been eliminated, and the comment may have never been written in the first place.
I saw:
static final int ONE_HUNERD = 100;
a while back, which really made me laugh.
static final String HTTP = "http";
static final String COLON = ":";
static final String SLASH = "/";
String url = HTTP + COLON + SLASH + SLASH + ....;
Which I guess isn't wrong, but isn't right either :-)I'd pick
const DEFAULT_FONT_SIZE = 42;
// ... lots of code in the same file
var font = new Font(DEFAULT_FONT_SIZE)
over // 42 is the default font size
var font = new Font(42);
any time.Two reasons: version 2 encourages others to copy-paste out of laziness, and it does not resist refactoring very well.
" Writing a task scheduler for your web-server in PHP"
Is the author talking about "at" and "cron" instead ?
I think the problem lies in that PHP wasn't designed to be a long running process, and for a scheduler you need something long running.
The master interface to the database was a series of PHP functions since the data massaging that had to occur (to facilitate correct interaction in multiple human languages across disparate timezones versus user preferences and time of day, etc.) ruled out direct DB access from anywhere else, back-end processes included.
I think the end solution was something like 'every minute run PHP X from cron, which checks if there are jobs, if so successively spawn children to handle'. It was basic but it worked. Godawful pain with MySQL replication over lossy Chinese internet WAN ... never again.
Anyway, neither decent ruby unicode handling nor RoR nor sidekiq existed back then. I actually considered rewriting the whole system in ruby at one point, but the unicode support was still dodgy.