Is it a Good Idea to Write Tests for Legacy Code?
stephenhaunts.com
stephenhaunts.com
The author mentions "characterization tests," which document what the code actually does as opposed to what it "should" do, or what the documentation says it does. These kinds of tests are gold, especially if you go back through the bug tracker and create tests for the major bugs that have been fixed. Doing this gives you a good framework for constructing a regression test suite, which is what you really want when working with legacy code. It's like the opposite of TDD, because you want a green light to start, rather than RED -> GREEN -> REFACTOR.
OTOH, some code just isn't all that testable by its very nature. I'm thinking of stuff that requires expensive, custom hardware that would be difficult to mock out, for instance. Or GUI code. Or, if you're unlucky, like I was in a previous job... GUI code that requires specialized hardware.
Also keep in mind the amount of work it takes to make the code testable. You can easily end up breaking things just doing this if you're not careful or if the code is just not written with testability in mind.
(RED shows you that the test fails when it should: there is good reason why this step should not normally be skipped; then GREEN shows you that the test passes when it should)
Although if you're fixing bugs, red-green-refactor can work well.
It's not enough to know that the test pass, you need to make sure the test would fail. Otherwise it's like having a cybernetic canary that doesn't need air to survive (great example, isn't it?).
And the problem then becomes "okay, now, which of the other 16,000 untested files in this codebase are relying on that broken behavior." It's not a fun place to be in (welcome to my day job) but you have to remember--You're writing characterization tests, you're trying to capture what it currently does. Not what it should--every difference between those two is, unfortunately, potentially an important one.
It's not worth fixing one minor thing if it breaks 14 other things down stream (that you then won't catch, since they also don't have tests).
I had an e-commerce system from a long time ago that was developed pre-TDD mania. My post-TDD approach was to develop a suite of acceptance tests for it testing a wide variety of typical use cases (add to cart, use discount codes, check out, tax treatments for various countries, etc.) and this has since saved my ass in a variety of ways without needing to go right down to the unit testing level.
(As an aside, if you're working on a project, need to move super fast and simply feel you haven't the time to do "proper" testing, be sure to do some high level acceptance tests at the least. They will save you time because when the inevitable problems occur, you can just code the process to run automatically rather than be clicking 101 times in the 23rd hour of that 24 hour hackathon ;-))
What also works well is to give someone new to the project some initial tasks to write and fill in gaps for missing unit tests and coverage. Just enough to get them productive from the start, but don't task them with writing all the missing tests. I've found that for most developers that it helps them come up to speed on getting their environment configured, looking at the code, and introduces them to your build and version control process from the start.
Ultimately it's going to depend on the project and how hard it is to go back and add the tests.
If a particular application is just milking out another year or two until it's replaced, it's probably not worth erecting unit test scaffolding around the entire app. Sometimes it's safer to just change a couple constants and do a smoke test than to really get into it.
That said, my experience was that it was immeasurably helpful in breaking apart and working with older code - just taking a portion of my time devoted solely to writing unit tests helped my understanding beyond what you'd expect.
Without the tests, refactoring is dangerous. This means that the systems are encouraged to remain static while the world around them changes. Systems that do the job, but for the wrong reasons, are legacy.
That last one was me. :)
It's relatively up to date in regards to technique and UML use but doesn't overpower the reader with too much.
Written to a decent level (it's not a Knuth book).
One main critique, it covers a lot of common issues but no complex corner cases, e.g. what to do about network architectures i.e. CORBA.
Should be required reading at the University level.
Yes, it's a good idea to write unit tests for the existing code base, though be careful with setting unrealistic expectations with regards to test coverage. The goal is to better understand the existing code base and have tests around the most critical parts of functionality to provide evidence they haven't been negatively impacted by the code changes.