Exploding Software-Engineering Myths (2009)
research.microsoft.com
research.microsoft.com
In fact, taking the geometric mean of the man month values (which, I think, but am not sure is the right metric to use when comparing ratios), we find the average ratio to be 0.947x. Make of this what you will, given the size of the sample.
[1]: http://research.microsoft.com/en-us/groups/ese/nagappan_tdd....
However, on a closer reading, even that paper isn't really assessing TDD. It basically defines any test-first coding practice to be TDD, including those that still spend considerable time and resources on up-front design.
What we really need is some properly structured study that compares professional practitioners (not just students or recent grads) over extended periods (not just a brief project that doesn't show the long-term benefits or costs in terms of quality and maintainability) who are using no unit tests but otherwise common development practices (control group), the same but writing unit tests after the main code, the same but with unit tests written before the main code, and full test-driven development where other aspects of the development process, particularly design, are effectively replaced by the test-writing activities as advocated in well-known TDD books and other training materials.
Of course it's all but impossible to really achieve that sort of like-for-like comparison with usefully controlled conditions, which is why these issues are so hard to debate objectively. But try finding any study with pro-TDD conclusions that doesn't obviously fail the objectivity test in at least one of the three ways above (not a sample of professional programmers, too short term to draw general conclusions, or conflating TDD with some other unit-test-related development process).
Another way of describing this is that the non TDD teams weren't actually done when they shipped.
Was ANY code, TDD or not, ever "done" when it shipped?
Of course, you'd better be darn sure the update works (firmware updates always scare me a little, never mind on a pacemaker!)
Sometimes partially perfect and released beats better but later. (Perhaps not for Microsoft, but certainly for lots of their competitors) How many more bugs can you find with software "in the Wild"?
This isn't to say it's worth tossing crappy code over the transom, just that real tradeoffs do exist.
Especially since "later" implies a new market, with new demands on the product. "Better" then becomes a moving target, and with "Best" you might as well say "Stop building the road when you reach the horizon and the sky stops you".
In the short term, TDD will look like it's costing more (especially if the product is not released). But if you want developers to maintain confidence in their code or for your team to depend upon the dev team's ability to deliver robust code, it's an important safeguard to have.
Mappings + conversions are such a case.
For example if you're using an ORM being able to test the mappings in seconds will increase velocity. Or if you've written a function to convert type T0 to type T1, then automatically testing all combinations of T is very effective.
On average though, it takes longer (as summarised in the article).
It's a question of tradeoffs, something I've said many many times, and the research appears to be bearing that out.
No one who is being honest can really say with a straight face that development with TDD is actively faster unless they're naive or they're using cherry picked examples.
If you want it out now, but have loads of bugs sure. But i'm pretty sure most managers don't mean that, when they say they want it out sooner.
TDD has both costs and benefits. Like most everything else.
In that light, TDD might just be a ploy to buy more time.
> Organizational metrics, which are not related to the code, can predict software failure-proneness with a precision and recall of 85 percent. This is a significantly higher precision than traditional metrics such as churn, complexity, or coverage that have been used until now to predict failure-proneness. This was probably the most surprising outcome of all the studies.
Which I take to confirm my belief that the freedom and responsibility given to developers, and how good they are, is more significant than the specific testing policy enforced :)
A developer with a short deadline is going to design things differently than one with a longer deadline. That superior design can often result in better long term maintenance.
There's also the idea of requirements gathering. Some people take it for granted, but if your developers aren't in a position of power in your company, they can't insist on good requirements for software. I've experienced this firsthand.
I once worked for a company in which the field engineers brought in a large chunk of the money for the company. They were the "big dick swingers" in the company, if you will, while the software development team was an attempt at the company to automate a lot of the work done by engineers and then sell services with the gathered data.
This presented a conflict of interest in that the software would have actively made the company need less engineers as they could do their job faster, but because of the two groups relative status within the company, there was no way for the software team to insist on accurate specifications.
The day I walked out of that company I had been asked to rewrite a module of the software for the 3rd time and requested written documentation. I was on the phone with both my manager and the engineer we were supposed to be coordinating with. The engineer flat out refused and told me I didn't need it. I asked him how many years of software dev experience he had.
That's one anecdote, but the point is, politics in a company severely affect the performance of developers and not simply because of testing. I honestly think that's an awfully naive view of the world.
Freedom gives freedom to adjust anything as required, not just testing.
Having freedom means being able to ask for longer deadlines to do things properly, regardless of politics for example.
It's also interesting to combine metrics like cyclomatic complexity, or relative churn rate, with coverage. Files/modules with high complexity/churn and low coverage are much scarier than those with low complexity and low coverage
Test coverage has had little correlation to code quality on the projects I've worked on.
I would agree that well-written tests of the most complex parts of code help contribute towards a quality system. Most of the other tests -- sometimes more so, and sometimes less so, depending on other aspects of the project.
TL;DR - it's a matter of context
But here's an anecdote where two complex pieces of code interacted...
A few years ago I was working on a fairly big project (~50 developers) and was responsible for much of the architecture, and on an implementation level, implementing some of the lines that connected boxes in said architecture. One of those was a line that crossed what can loosely be thought of as a system/application boundary, this involved both implementing the application hosting layer and writing the marshalling code that moved data in either direction, there was a bunch of complexity because of not quite convergent type systems, some threading issues etc...
So all this is done, not too many tests have been written, but hey, it works, if it didn't nothing in the system would work... a year or so into the project, the most senior technical person on the team has a flaky test that once in a while just completely fails to make progress, he investigates, and comes to the conclusion that once in a blue moon, the transport layer (my code) is dropping a message.
Needless to say I'm totally flummoxed, there is no code path that can drop messages without also creating a huge stink in logs etc. I spend two months investigating this, eventually I have a consistent repo, a fairly big hole in the other guys state machine that causes it to just stop making forward progress.
So despite the fact that a test exposed the bug, neither piece of complex code really had proper tests, but I would assert that the "line" didn't need the same level of tests, it was being exercised constantly by dozens of "boxes", but the unicorn "box" with the logic error could sure as hell have used some unit tests to validate its state transitions.
That's a fundamental problem that neither TDD or any other methodology can really solve.
At some point, an over-zealous administrator can disable the rights for the account that my application is using inside a client organization. Then I end up on the phone for an hour or more, figuring out why an account that used to work no longer does, and getting my clients to raise a change request with the admins that broke our application in the first place to put things back.
Some stuff you test in unit tests, other you do functional testing for. What you describe is probably higher level than unit tests.
>At some point, an over-zealous administrator can disable the rights for the account that my application is using inside a client organization. Then I end up on the phone for an hour or more, figuring out why an account that used to work no longer does, and getting my clients to raise a change request with the admins that broke our application in the first place to put things back.
So? That's orthogonal to your app. You test that your app works if the rights are working OK, and you test what it shows to the user if the rights are not OK.
Other than that, obviously you cannot ensure it works without things it needs to have, nor do you need to test each and every externally dependent failure mode (e.g. user wiped out his hard drive).
It could be a lot of work, of course, but that just reflects that you're building a complex system.
Test 1: Look up AD prop with user credentials Test 2: Look up AD prop when user credentials don't work
You'd almost certainly want to mock out the environment. Don't use the actual services.
It's clear to me what my code needs to do (Above)
How do I even know what correct outputs to code interacting with those components would be? (From your original comment)
If you know all the possible environmental factors that you need to deal with, then you can write up front tests for them (whether that is economical or not, is a different question).
However, if the complete set of environmental factors is unknown (though potentially discoverable) then you have a requirements problem, not a code problem. TDD can't solve that problem (although it might help make it obvious that the problem exists).
TDD is based on the assumption that it is possible to know when your code "works", and requires you to define that criteria up front.
But any development process you follow is must have some way of answering the "am I done yet?" question, or you'd never ship anything.
Perhaps your concern is that TDD doesn't help solve the "I don't have a complete set of requirements" problem. If so, then that's true, but it never claims to. If you don't have the requirements yet, then TDD says you're not ready to write code, and some other process has to be followed in order to produce requirements.
It is possible to do TDD if you only have some (but not all) requirements - but you can only write code for the requirements you do have.
Saying that "no methodology can solve this" is defeatist: there are lots of things you can do. One technique is to write your tests to use the actual component, instead of mocks. For example, I never mock the filesystem: I always design tests to run using the real thing. This reduces the chances that I've encoded some bogus assumption into my tests, and makes it more likely that I'll encounter the individual quirks of the filesystem.
That's hardly OPs fault, and neither more design or more TDD can solve it.
I wrote a program that used the computer's MAC address. The API that reports the MAC strips leading 0s, but my MAC had no leading 0s and so I I did not notice this behavior during development. If I had just captured and replayed this result, my tests would have passed on all machines, even affected ones. But because my tests exercised the real API, I encountered the bug and was able to work around it.
In practice, I tend to do a mixture. I will mock/fake the external services and then build adapters and then finally remove the mocks/fakes. In the final part you may have to stub the adapters. The unit tests in the adapters will alert me to changes in the external services, so this is relatively safe to do. Your main sources of error will be subtle bugs where you have insufficient coverage in the adapters. You must also use restraint in working on the adapters independent of the code that uses it, because you have no end to end tests.
Building fakes instead of mocks in this situation is also a good way to really test your understanding of the external services. It takes more time, but it can pay dividends. Many times I will write a fake and then think, "Does it really work that way?" only to realize that I can't actually use the external service in the way I envisioned. With a mock, you completely divorce yourself from implementation so it is relatively easy to assume that the external service can do impossible things. You can go a long way before you realize that your code is unusable.
One thing I will caution -- mocks, fakes, stubs and integration tests tend to attract people with almost religious views in how they should be used. I tend to believe that this is stuff is hard and that current industry practice is very naive. I think there is a lot of room for experimentation with a very big upside. But you are likely to attract a lot of criticism if you wander outside someone's boundaries, so it is best to spend significant amounts of time pairing to make sure that everybody is comfortable with the approaches you use. Nothing is worse that "test wars" where people are going in 100 different directions in the tests.
http://stackoverflow.com/questions/2611280/how-can-i-effecti...
[1] http://henrikwarne.com/2014/02/19/5-unit-testing-mistakes/
For me the definite example that TDD for algorithmic design/exploration is not very useful is the infamous Sudoku Debacle, but I've also never seen anything more than toy examples (e.g. fibonacci) of algorithmic exploration using TDD. Is there any real-world example of complex algorithmic design using TDD where the author doesn't cheat by using hidden domain knowledge to make "magical" leaps of inference in some of the TDD steps?
PS: just read your blog post, but you seem to be talking about unit testing with JUnit & mocking frameworks. Just to be sure we're on the same page: we do agree that TDD and Unit Testing are not the same and do not even share the same goal, right?
Like, does anyone feel really good about saying there's good test coverage of a class with 50% code coverage? 60%?
After I hit 80-90% coverage, I start wanting to know about other attributes of the tests. Are they really really really testing that all the outputs are exactly what we expect? Are they controlling dependencies in a useful, realistic way? Are they exploring a wide range of inputs?
Before that, though, all of the above seem a bit premature. I don't deeply care if you've done a really thorough job of testing 20% of the code if the other 80% isn't even run.
We have several with 0% and I'm just fine with that. They're exceedingly simple. And/or they're 15 years old and only change in very slight ways every few years.
There's other classes that are 100% that I'm not fine with. Because while the code is short, the implications are complex.
I feel like blindly talking about these metrics is like saying "hey, I bought an earring for $100, did I get ripped off?"
I'm talking about to what extent code coverage can and can not be used as a proxy for "well tested," not to what extent "well tested" can and can not be used as a proxy for "I feel good about my code."
No, sorry, not even close.
To me, measuring the value of a test suite by code coverage is like measuring the productivity of developers by lines of code. It's just not what actually matters.
For example, it's well known that bugs in code tend to cluster. Given a large system where some parts are inherently complex and therefore error-prone but most of the code is routine boilerplate stuff, I would much prefer to see test effort concentrated on the challenging areas.
Consider this as well: many interpreted languages count the definition of a method/function as a "covered line of code". In other words, if a method/class is defined then you have at least one (maybe several) lines of code "covered". Add to that things like declaring variables and you can easily get to a point where (in good code with very short methods) you have 60% coverage if the method is simply defined and run. If you are writing really good OO code that has very few conditionals, then you may get very, very close to 100% test coverage simply by running each method.
Also, tests are code too. You are just as likely to write a bug in your test as you are in your production code. In fact, I've noticed that in some systems, I actually end up writing more lines of test code than production code. I might actually be more likely to write a bug in the tests. Especially for bugs that are due to problems with requirements, you are very, very likely to put the same bug in the tests as you are the production code.
This is why I don't like calling these things "tests". IMHO the very best "tests" document assumptions that the original programmer had. If you break the assumption, the the test should break. It may or may not be a bug, but it is a big hint that you should be thinking very carefully about what you are doing.
Code coverage is a good statistic in the reverse. It doesn't tell you much if you have good code coverage. On the other hand, it can lead you to find places where your tests are unlikely to be sufficient. Depending on the language/environment you use, if you have less than 80% code coverage in an area, it might indicate that nothing of value is being tested. It is a good place to start looking to check.
Note that setting a code coverage target will defeat this strategy, so I recommend never, ever doing it. As I said, getting nearly 100% code coverage without actually testing anything is relatively trivial if the code is in any way decent (and if it isn't, then you have bigger fish to fry). Setting a threshold above the place where poor testing is obvious just means that you no longer have a means of finding places that are poorly tested.
Like, I agree with every claim about fact that you're making. Of course you can have high code coverage without having good tests. If all you do is, like, touch lines of code, all you're really testing is that they don't raise an exception. That's why I said that once you hit 80% code coverage or so, I start to want to know not "what your code coverage is," but "how good are your tests?"
But if you don't even run the majority of your code, then your tests are kind of by definition... not testing your code well.
Depends. 20% might be enough too. You can just test the 20% harder and more failure dependent paths in your app, not the 80% trivial BS code.
Remember the 80/20 rule?
Like, seriously guys, I'm sure we can all come up with some kind of degenerate case where 80% of the code is in exception-handling or some other really, really, really rarely called piece of code, and feel okay with not testing that. But that's hardly the ordinary scenario.
And, for example, in the messaging code I've recently been writing, it certainly is the case that some of the code is exercised a lot more than others. For example, reads are done an order of magnitude more than writes, and writes are done an order of magnitude more than deletes.
But that doesn't make me feel okay about bugs in my delete code. It's a whole feature. It needs to work.
Then maybe you don't understand the notion of "opportunity cost".
Perhaps you feel "0 is the only acceptable amount of bugs", but
a) nobody cares for an imaginary bug-free product unless it ships. And shiping, for programs intended for the market, means cutting corners and, yes, shipping bugs too.
b) programs that have 0 bugs are so few as to be statistical noise, including among TDD/full coverage programs.
TDD will basically help you cover that the behavior you expect is happening and help you track breakage when you need to refactor. It wont eliminated all bugs.
>But that doesn't make me feel okay about bugs in my delete code. It's a whole feature. It needs to work.
Perhaps you conflated untested with "not working". People created succesful software for ages, software that changed the world, with no or less than 50% "test coverage".
TDD "code coverage" always seemed like a misguided breadth-first approach that needed to be depth-first primarily and breadth-first as needed.
But I think the latter approach then leads to integration testing as the primary approach, so TDD zealots will hate that.
I wouldn't want people to conclude that coverage was a useless metric based on this, but it does seem to support the idea that chasing code coverage dogmatically is not necessarily the best use of your time.
But a coverage ratio and/or browsing annotated source is a good feedback mechanism for how many tests you've written.
The article does prove that some people tried to impose (via management or structure in large organizations) metrics on coding as a way to rationalize the development quality and miserably failed but clearly lack the counterpart: what happens when you don't impose these metrics? Experience is important but that's not new is it? In the end it just feels like: small programs won't have much problem as do small organizations - Big organization should function like small companies. Not much of an answer.
Also, I'm not sure that a manager saying something like "according to data, we need one more month to make the program 35% more stable" is a good argument (or even a new one). What would be interesting is to try something like "ship the program one month earlier and dedicate the two remaining month to fix what's broken according to feedback", then compare with the first approach.
Perhaps they should measure the coverage by taking the number of tested paths through the software divided by the total number of paths.
"The assertions are primarily of two types of specification: specification of function interfaces (e.g. consistency between arguments, dependency of return value on arguments, effect of global state etc.) and specification of function bodies (e.g. consistency of default branch in switch statement, consistency between related data etc.)."
from http://research.microsoft.com/pubs/70290/tr-2006-54.pdf
Is there somebody here, who could elaborate on this, ideally with some small examples?
Consistency between arguments: memcpy "The memory areas must not overlap."
Dependency of return value on arguments: max returns one of its arguments. r = max(x,y); assert (r == x || r == y);
Effect on global state: printf might change the global errno variable.
I'm too lazy to check the actual paper, if they clarify their understanding. ;)
I've never heard of this before. Where does this come from and why did they suspect this would be indicative of anything?
I think it goes deeper than "the conditions you forgot to write" although that's part of it. A developer's test suite tends to make assumptions and these are the same assumptions that exist in the production code. If your assumptions about usage, or input, or edge cases are poor then the tests you write will be poor too and in that case 100% coverage isn't worth much. Big reason I'm a fan of things like QuickCheck in Haskell.
People who like to talk about the "proper" way to develop software tend to like it.
People who run open source projects tend to either like it or bow to peer pressure.
But people working on in-house and enterprisey stuff? Yeah, not so much there.
This depends a lot on your corporate development culture, even for in-house enterprisey software.
I suspect you are working from a limited sample size there.
That is absolutely incorrect, in my experience.
If you don't have tests, it is an extremely manual process to change anything about your codebase (including dependencies) and have any confidence that you haven't broken things. Once you write tests, code coverage lets you figure out what code is not being tested.
It's defining the term "code coverage" to be the ratio of lines covered by tests to the total lines.
> why did they suspect this would be indicative of anything?
If anything, it's a useful feedback mechanism. e.g. "Gee, we wrote sixteen unit tests for this module but we're still only at 34% coverage. Ohh, I see it now -- framistan.cpp doesn't get used by these tests at all. We should write a test that enables framistan mode!"
Early 1960's or so!
IMO, proper "culture" is a significant part of almost all aspects of project management and no matter what, without it, the teams will never reach their optimum potentials, unless some highly motivated dev takes upon their shoulders to implement these base tenets for everyone else.
I'm not saying it's always easy to implement or it guarantees success, but that the lack of it always rears it's ugly head at some point in the development process.