Lean Testing or Why Unit Tests Are Worse Than You Think
blog.usejournal.com
blog.usejournal.com
At the end of the day, there is a spectrum of tests going from unit tests to end-to-end tests. The spectrum represents several trade-offs such code locality vs. coverage. In my experience, the most economical approach is to write a balanced mix of tests the lie along this spectrum.
With bite-size integration tests, I find it's generally not too hard to isolate the cause of a failing test, because the code it's testing tends to be straightforward, and fairly easy to step through, if necessary.
I frequently have a harder time with unit tests. The code ends up involving a lot of extraneous abstractions that I need to think through. The test code is often so heavily mocked that it's hard to distinguish the behavior under test from stuff that's just being mocked or stubbed to get the SUT to run cleanly, meaning I've got to start with trying to figure out whether the bug is in the test or the code being tested.
It gets worse in long-lived code bases, where the unit tests are often subject to significant bit rot on account of how brittle they are. I've definitely had some code archaeology excursions reveal that the reason an entire suite of tests were tautological is because someone was doing a nominally unrelated refactor, and just put in the minimum effort necessary to get the tests to go green again.
You can argue that developers need to be more diligent. Me, though, I figure it's sort of like those lines of bare dirt you see criscrossing the lawns of university campuses: when things get to that point, it's a sign that the official way of getting around isn't appropriate to most people's real needs.
I do think it's important to have unit tests when the unit's behavior is complicated. Where I start to get worried is when there are unit tests being written against classes that have very little behavior that doesn't involve interaction with some other module.
I think this is precisely where unit test suits start to have problems. Good, flexible unit testing requires a lot of judgement about what will be useful to test and what will be too much of a burden in the future. Unfortunately judgement is hard to acquire and even more difficult to teach, and a lot of teams want to create and enforce over-dogmatic testing "standards." When unit testing, you have to balance:
1. What testing do I need to have confidence my code is working?
2. What testing do I need to catch likely regressions?
3. What kinds of tests will just get in my way in the future or are literally useless [1]?
[1] E.g. unit tests that essentially only test core language functionality, once you take out all the mocks.
So for a little passthrough/orchestration class, it probably doesn't make sense to do much testing. For something that actually performs business logic, that's a prime candidate for testing. I've seen plenty of tests that just seem to aim to increase coverage, heck, I've written plenty of those myself - but at the end of the day, the benefit they serve after being written is probably minimal.
You're only unit testing the code how you 'intended' for it to work at that time. Even though the tests are written, it probably wouldn't be uncommon for a bug to slip through when running your code, what you then can do is write another test to account for that scenario, then repeat and your code becomes more robust as a result.
Correct that the goal is to prevent regressions. I claim (I don't know how to study this) 80% of your tests will never fail and so they could safely be deleted - but I have no insight into which tests will fail so I say keep them all.
Incorrect because in fact it isn't hard to localize failures: it is something in the code you just touched!
Yes, but you also need to localize the effect of the bug to know why is it that the code you changed broke the program (and remember that we're talking about a case where the rest of the program is not familiar to you enough). Good unit tests can help you find the immediate effect of the failure, rather than the ultimate one.
Keeping below 1 bug per 100 lines of code is viable simply by being careful and thinking things through. That's a long way from perfection, but it really helps.
Once you can generally write mostly correct code then you can work on improving your process. But, until most of your functions are working when you right them just work on improving that.
Edit: And yea quality may drop as part of this transition, but you need to get a feel for how much you can pay attention to.
The tests are there specifically to help find the errors that “trying harder” didn’t catch. You don’t get a higher quality result by cutting QA.
You can build a factory that has a Great QA process that finds all faults and corrects them. But, inside the factory you still want to minimize the amount of defects for QA to find. Either way you still need QA, but it's easy to get into the habit of improving QA vs improving the initial process.
PS: One exercise I did for a while was every time I found a non syntax bug in a function I just wrote (aka bug at run-time) I would start over and rewrite it. It's painful, but 'Build' wait let's look at this again if I hit run and it does not work that's going to be painful. So is this actually correct? Anyway, doing that showed me how important it was to just focus on the code (and what it means) to exclusion of everything else.
But you will not drive significantly higher quality into earlier phases by telling your factory workers to try to do a better job. You drive higher quality by analyzing where your processes are failing you and changing them so that they no longer do.
If your workers are consistently failing to tighten bolts sufficiently, you can add a QA step to check the bolts every time, or require your workers use a torque wrench, or use a robot to make this more consistent. But you have to do something other than telling them to try harder. You undoubtedly told them to try harder already and it didn’t work. Now you need better tools or processes.
However, the lady designing the factory in the first place should know if tightening bolts or using an arc wielder has more issues. So sure, QA should include checking the oil level inside a robot in the factory to avoid problems in the first place. But QA is not all inclusive it's not for example part of the collage education that's training your workforce other than perhaps a requirement that they have said degree.
QA is the system of trying to produce at the desired quality. Anything done to improve quality is QA, including whether you hire better employees or train the ones you have. You can define QA more narrowly if you like, but then we're arguing semantics and we're already off in the weeds.
Heh...irony alert. You're advocating focusing on just getting it right the first time and you didn't get write right.
But seriously, the people you're responding to are correct. Any process that relies on humans being less error proneis bound to fail. You need to either create a process that makes humans less error prone (e.g. checklists) or embrace our propensity for making errors and plan for that eventuality.
> I find it nearly impossible to write a "unit" of code without also writing one or two bugs along the way
Finding bugs early is great, but minimizing bug creation is even better.
1 3 2 5 9 2 8 7
vs. 1 3 2 5
9 2 8 7
Memorizing the first sequence is trivial, but dealing with each half independently is much less mental effort. Now each idea can be more complex than just a number, but even still you make fewer mistakes by removing mental overhead.PS: Now, their are a lot of tricks on how you can get better at this stuff. But if the parent poster tends to write bugs in most functions then simplifying things may help.
Sure, it's less efficient to write unit tests ahead of time for every case, but in a lot of cases fixing a bug faster when it's discovered is much more important than even double the hours spent during less "pressing" times.
I tend to agree - in my personal experience in organizations with strict TDD culture, a perverse incentive often emerges to preserve existing flawed architecture over obviously better solutions just because it's so painful to deal with all the tests.
One of software development's most powerful properties is the ability to iterate quickly: it's foolish to prioritize dogmatic beliefs about testing over that quality.
It's just an extension of your code, if you are going to throw some code away to change an interface, then throw the tests away too. If you are afraid to do that because of the time spent, then you probably spent too much time on writing tests.
While I am dogmatic about tests, I also believe that around 50% code coverage is normally enough in most application codebases. Cover the important parts, the parts that are hard to test manually, the "core" pieces that lots of other areas rely on, and some tests for bugs that you want to prevent happening again.
If you want to quickly iterate, go for it! You shouldn't have all that many tests in the parts you are changing frequently. But to change a core aspect of the codebase, or a really complicated aspect of it, then the extra work of rewriting the tests shouldn't be all that bad.
I'd shorten your last statement. It's foolish to prioritize dogmatic beliefs. Everything has a time and a place, and moderation is key.
This metric means very little. It does not measure the extent of code path coverage and much more.
Code coverage can tell you what code is not tested at all.
Now, this is very useful. But:
Code coverage can't tell you what code _is tested_.
Code coverage can't tell you how _well_ code is tested.
IME unit tests work acceptably in one very specific scenario and fail pretty badly in all others. That scenario being:
1) You're surrounding a self contained block of code that interacts "with the outside world" via a code API.
2) That code API is a very stable and clean abstraction.
3) It has minimal interactions with modules outside of it and those interactions that it does have are tightly scoped (i.e. minimal to zero mock objects are required to write the test).
4) The logic of the code is relatively complex and most bugs that crop up are logical in nature (e.g. off by one, things getting swapped around, incorrect calculations, wrong behavior with negative numbers).
Meanwhile, integration tests (at varying levels) work well for pretty much every case apart from this and still work okay for this type of code. They make much more sense as a go-to default.
I've also worked on several projects where there was little to no code that it actually made sense to unit test. It's not uncommon that an entire codebase is predicated mainly on hooking systems together and doing some shallow calculations. IMHO, having zero unit tests in that environment is actually desirable.
The worst unit tests I've seen have been written when two or more of those preconditions have failed. They would fail constantly, require massive maintenance and, somewhat comically, almost never fail in the presence of an actual bug.
But it doesn't follow that changing a unit's interface means unit tests suddenly become just a burden. Ideally unit tests are, well, testing a bunch of core functionality of the unit under test. You adapt them to the new interface. Then you're back to having a quick, automatic sanity check you can run against the unit whenever you have to make a change.
I don't understand people bemoaning this 'cost' of unit tests when the benefits they provide typically far outweigh the costs. It's possible broader functional/integration tests have a better ROI in certain situations, but they come with a maintenance cost as well.
If your refactoring is largely centered around changing unit interfaces (not uncommon) then it means that those unit tests are 100% overhead because most of the time they fail just because you changed the code.
>I don't understand people bemoaning this 'cost' of unit tests when the benefits they provide typically far outweigh the costs. It's possible broader functional/integration tests have a better ROI in certain situations, but they come with a maintenance cost as well.
I've certainly found that the ROI is a lot better. I find that integration tests have a higher up-front cost but maintenance-wise they're the same or cheaper. % of failures that are actually catching bugs is higher too.
Another condition I would add to the list, which the component of mine I am thinking about meets is:
5) Failures in this component have cascading effects to multiple other components, causing seemingly "impossible" failures that are not obvious to others that they were caused by this component.
That needn't be a unit test. It could simply be a lower level integration test.
Maybe that should be a number 6): easily tested invariants.
Not sure about other types of development, but in web development, TDD can be a good way to have automated tests without the additional cost.
And indeed, tests don't take much time to write once you get used to writing them. It's like anything. The more you do it, the better you get at it.
If you’re very familiar with testing and/or doing TDD, you might include your testing costs in your estimate, but you still have the cost to pay. And if you write good testing up front, it will cost more up front.
We have a full continuous integration environment at work, and the tests run there fine, but trying to reproduce that on a local machine is often a fairly difficult experience.
We have maybe 30 components in our system, so often I haven't touched the component before and I am asked to fix a bug in it. Sometimes they are using standard testing libraries, other times there are a lot of extra libraries that I haven't used before. Getting everything to play nicely isn't always trivial.
This drastically scopes down the surface area for CI breaks.
[1] There are a few exceptions for tests that specifically cover environmental config/behavior that cannot be fully tested locally.
> invested heavily
pick one.
But also, testing in browser is more expensive long term that writing proper automated tests. Up front it’s cheaper, like any other technical debt.
Sorry for any confusion for anyone reading my previous comments.
Cake + Eat-it, too.
In the browser, testing edge cases sometimes requires custom headers, encoding of data to create authorization tokens, etc. This means you have to either have amazing browser tools (which I've yet to find) or use a combination of browser and shell to achieve what I need.
In TDD, most setup can be automated with simple function calls. In addition, well made frameworks, such as PHP's Symfony, have utilities that even avoid making real HTTP requests, so tests run faster than using a browser, but the result is the same.
I'm not saying everyone should do TDD, but from my experience, it can lead to increased productivity and fewer bugs in some cases.
By the way, if you know of browser tools that make testing easier, I'd be glad to learn more. I use TDD because I lack the browser tools. If I had the right tools, maybe I'd consider going back to in browser testing.
When my function under test calls sort() I don't fake that call out, so technically I just wrote an integration test. (In fact I work with one group that will inject a fake sort())
If I write a library foo which has sub-module bar which has class baz and fuzz, is the test for bar that tests both baz and fuzz a unit test of an integration test? Of course if you are a user of my library tests for foo are unit tests to you...
There's no agreement what is "unit". Few classes interacting with each other could still be a unit.
I didn't find any strict definition which would be useful in practice. I just write tests and guess they are mostly integration tests, some end-to-end tests and a few unit tests.
Since we live in a world of limited budget and time to spend, I agree with the conclusion of the article: "use unit tests where it makes sense". It is what we try to implement on our new code (and when refactoring legacy code).
That's completely at odds with my experience.
I find that for local bugs, the cost of locating them grows with O(log n) of the amount of code. And for non-local bugs (interface bugs, system bugs, incompatible specs...) unit tests don't catch them anyway.
I don't think this is true. Fixing a bug is comprised of four parts. 1. Understanding and reproducing the bug.
2. Finding the code responsible for the bug.
3. Coming up with code that fixes the bug.
4. Verifying that your fixing code does not introduce any new bugs.
#2 is the only one that could even theoretically be exponential. The only bugs I've found where I've spent days on 2, were hard to reproduce, intermittent, race condition bugs, which units tests aren't very good at finding anyway.
A unit test, or a component test, would tell you that something failed, and it's in this general area, which narrows down the search quite a bit.
They're both useful, but I've seen far more problems with people arguing that end-to-end test is more than what they need, while a bad conversion from seconds to nanoseconds would be caught quickly if a unit test was actually written.
Unit tests also have the advantage of being cheap to set up and cheap to run, so their limitations in whole-system testing are offset by the fact that you can have many, many more of them. You can easily blow through a couple thousand unit tests in a matter of 10 minutes or less, whereas even a modest battery of integration testing will be an overnight process at best. That means your OODA loop is a lot faster on unit testing.
(the ideal is of course structuring your application so you can deploy many tests in parallel, but that's not always possible.)
Fixing Selenium so WebDriver instances don't take 60+ seconds to launch would be a major help. Until then, integration tests that work at a controller level (minus the UI) are probably the best compromise between speed and integration. You can pass in a mocked request as input, and assert on the output. Runs much faster than a "true" integration test with a driven browser client, but hits the whole stack better than a unit test. They are also really easy to set up, since you can just have a tester or an integration test go through the motions and then dump request data to JSON, then feed it back as a test. The lack of state persisted between requests makes implementation a lot simpler.
They are not cheap to develop and maintain however. And given enough developers a certain percentage of them will completely distort the design of the code to enable unit testing making the code itself much more complex.
If it's because you need to change them after a feature changes or breaks, isn't that the entire purpose of testing and not exclusive to unit tests?
Others see units as the smallest bits of irreducible business or tech knowledge. So if you have 20 methods in a class, you test the 5 public ones, or the methods that compose a useful "unit". This is the definition used by the original test driven development book.
Many of the unit test advocates I've read use the second definition because testing private functionality incurs all the problems the article points out.
Using the second definition leads to the testing pyramid, which is a thing of beauty. Unit tests have a few mocks to the unit's partners. The unit tests test for the correctness of the unit. The integration tests are one level up in terms of code coverage in a single test, and they ensure that the mocks used in the unit tests are correct. The e2e tests are then used to test that the app starts correctly and is configured correctly. The obvious happy and sad path scenarios are also tested to ensure that dependencies are functioning correctly (in a complex application there can be hundred of "obvious" e2e tests).
In my career, every single instance I've seen about someone complaining about writing unit tests, they're either using the first definition I mentioned above, or their code is not composed correctly. I'm sure exceptions exist somewhere, but I've yet to see them.
Using integration tests like the article describes leads to a host of problems. For starters, it's extremely difficult to get complete code coverage without huge or redundant test suites. These large test suites are harder to read and take longer to write because more complex testing requires more complex mocking. Regression is also more likely during large refactors since you haven't clearly enumerated expected behavior.
Unit tests should only change if you also change public apis (I have one exception to this rule, which is very complex private methods which could easily be their own unit).
Then if you're changing the tests, you're also refactoring everything else.
>For example if you use mocks to invocation counts, refactorings can be even more costly.
These should only be costly if you're over-testing. You shouldn't be mocking and asserting calls on every call within a system under test, you should be checking the ones that matter.
Correctly written unit tests are a refactoring aid.
It is OK to be economical in what one would see as a cost of testing a product end-to-end. Call it whatever term you like, but functionally testing a product end-to-end has always followed an economical approach, as all such testing is a factor of time & resources. You'd perhaps be less economical if you were testing life-safety systems, but you could choose to be more economical if you considered your software to be not-that-critical, or, you needn't have to set exceptionally high standards.
I find it extremely odd that unittests are even considered into this equation of 'testing-costs', despite decades of improvements in software development processes. Unittests should be part of 'core' engineering & a part of development. Not a task that's added to testing costs. If you aren't writing unittests (irrespective of whether its TDD or not), you are not developing/engineering your product right. You are only building a stack of cards _hoping_ that it wouldn't collapse at some point in the future.
It's a sad state of affairs if one gives code 1st level treatment, but treats unittests, documentation, build scripts and other support infra as something less important; worth economizing. This is really what _matters_ when it comes to overall quality of a software.
One must remember that software is always improved, refactored, expanded, ported or worked on in some way or another. When unittests are missing, then the very boundaries that were meant to dictate the rules of the software don't apply anymore. This leads to human errors causing portions of software to break.
Critical pieces of software that exist today, exist strongly because they were engineered right (Linux kernel for example). Not because their developers followed an _economical_ approach to testing. If an OS disto claimed to perform economical end-to-end testing (of a potential user's most commonly performed paths), would the author of this article want to use that OS over one that has had strong ground-up unittests and testing of interfaces where they matter?
Just a quick google and not even the paper I was thinking of: https://softwareengineering.stackexchange.com/questions/6050...
The rule of thumb I always come back to is "the minute you decide it's time to attach a debugger, be prepared to burn an hour." I used to not mind firing up the debugger -- heck, I used to run my code with a debugger attached every time I ran it. These days, I write tests. I don't aim for 100% code coverage, though I aim for a majority and 100% coverage of critical paths. When I encounter a problem that isn't covered by a test (rare), I write a test centered around the problem. This usually surfaces the bug before I've even ran the actual test -- and once that test is there, any refactoring that affects that code will avoid that bug if the test passes.
This rule has been so effective for me that I've now gone through four different teams, encouraging each to do "the next project" with automated testing. The result has been a commitment to working that way from that point forward[1].
[0] Assuming some baseline quality requirements -- automated build, and a general requirement to publish a mostly functional product. :)
[1] I can't take full credit for the idea -- it was a blog post that convinced me to try it on myself and I advocated that idea after it convinced me.
I'm a big fan of "economic testing". Test that are not going to "be economically viable" to write during the dev't of new software, usually do not get my permission to be written in that phase.
Unit test thus get written for part that we absolutely want to be maximum certain of that it works correctly. Mostly that functionality can be extracted into a library. For instance a lib that deal with scheduling of events of arbitrary length: we do not want to have issues with those getting messed up, so we try to get 100% unit test coverage of the functionality in that lib.
Another thing with unit test is that I found they mean something different in dynamic languages than in strongly typed languages. In the first unit test are often something to provide confidence that the type system can give you in the latter. Sure both will also test the logic. But this understanding has pushed me to love strong types (as a rubyist), as I can go with less unit test and still be more confident about my code base when it is being heavily worked on. Refactoring becomes so much easier with strong types, as all is "in the code".
When it comes to integration tests I usually want to make a biz case for them. How much does it cost to have it all manually tested, and how much do the automatic integration tests cost to make. While including the benefit of being able to run the automated integration tests at near zero cost (thus several times per day if we need to).
I think this is a partial answer.
Yes, I write unit tests for such pieces of software. However, what I also write a lot of unit tests for a pieces of software which are hard to write good integration test for because of combinatorial explosion. If I have a part for integration method which consists of a chain of 4 methods, each of which can have 5 different results, that'd mean I'd have to write 4^5 integration tests, which is clearly not possible. If I were to write unit + integration tests, however, I could do with 4x5 unit tests plus a couple of integration test which prove that the overall combination of those methods work well together.
In practice I've also run into a lot of combinatorial explosions for which unit testing is worse than useless. E.g. say you have a library that has two dependencies and interacts with browsers you can have:
* 30 different versions of library A
* 50 different versions of library B
* 25 different versions of the language runtime
* 30 different browsers
That's over a million different potential variations for every test. In practice though, there's a sharp power law. You don't have to test all of those variations, you just have to randomly throw variations at the code until bugs stop cropping up. 99% of the bugs you find probably be found in the first 1,000 iterations.
* (lib A, lib B)
* (lib A, language runtime)
* (lib A, browser)
* (lib B, language runtime)
* (lib A, browser)
* (language runtime, browser)
in 50*30 = 1500 tests (or a little more to make the analysis tractable)I hear this kind of rhetoric a lot, but rarely do I hear of code that doesn't need to work correctly.
The difference is having code for which having some bugs is acceptable, vs having code for which bugs are not acceptable.
There is definitely code for which having some bugs is acceptable: think about some small feature that the user can totally do without. I had to do one recently, and bugs were acceptable, as long as it was shipped in time, and working for the client specific devices. That code shipped with no automated test (testing is done manually), and a few known bugs not worth solving. On the other hand, anything that touch core features must be tested.
An issue with that is bus factor when it needs to be changed in the future. What if the original dev isn't around any more?
Also, the code has been reviewed and is documented.
Bus factor really isn't a problem in the instance I was talking about. In fact, the original dev won't be around in a few months, and that's fine. The rest of the team will be able to modify the code.
A graphical error that prevents users from being able to use the application at all can be basically the same as server downtime in terms of impact to the end user.
On the other hand if the code responsible for the styling of some menu is not working correct, the functionality may still work, and once we discover the bug we can simply fix it.
> rarely do I hear of code that doesn't need to work correctly.
So there are different grades. That's why I used the word maximum. The cost for these unit test should pay bank, even during the dev't phase, IHMO. There are often areas in a large code base for which this hold true, and often they can be identified as such beforehand; then writing these tests should be a priority.
While I use all sorts of testing, I still write a good amount of unit tests. For whatever kind of tests I write: not a day goes by without tests or the act of writing them exposing bugs that would otherwise have moved on to the "next level". Doesn't mean they'd have ended up in production, but they'd still have been a lot more expensive to fix.
> Kent C. Dodds claims that integration tests provide the best balance of cost, speed and confidence. I subscribe to that claim. We don’t have empirical evidence showing that this is actually true
And
> There is the claim that making your code unit-testable will improve its quality. Many arguments and some empirical evidence in favor of that claim exist
Why rely on evidence when conjecture an opinion are so much easier?
So here's my own limited evidence to toss on the other pile. Our team has started to unit test everything, 100% coverage. We don't take this approach to testing lightly, it is a holistic approach that includes the code itself, focusing on architectures that enable testing and testability. Doing so has vastly improved code quality in measurable ways.
Code reviews are faster and easier as a starting point. It is easier for another person to understand quickly what each component does and what it is expected to do. Code is composed using small autonomous functional components that are uncompleted and easy to reason about and test. These make debugging easier because they can be isolated to simple components and tests. Bug fixes are faster to implement and test since the test infrastructure is already in place and does not need to be remediated in later.
Yes, we do also write integration tests. They tests the end-to-end scenarios we are designing the system to support. But unit testing is still the primary approach to how we design, while integration provides support by verifying complete scenarios.
However our approach admittedly doesn't have "Lean" in the name, so that's one downside I guess.
You are engaging in a strawman. I don't claim anywhere that these are the only dimensions to care about, I simply claim that these are some things we've noted that testing leads to improvements.
Ultimately though my primary claim is simply that I see less value in opinions vs evidence.
Though I'd like to think that I presented clear arguments even though I didn't present _empirical_ evidence for my arguments. I presented ideas and logical steps how to arrive at them. Which, to me, makes it more than "just an opinion". Sure, you may disagree with the ideas and or the steps–it's not a mathematical proof, it's much softer. Sidenote: Nonetheless, I provided two links to empirical research finding counterintuitive results from TDD and from unit tests.
Clarification: I am not engaging in a strawman. I didn't claim that _you_ said these are the only dimensions. I tried to express, generally, that software quality is one of several important dimensions. I could have phrased it in a clearer way.
This statement is an implied counter argument. If not then what is it? If so, then I suggest it is a strawman as I never argued against that.
> I'd like to think that I presented clear arguments
If you are referring to the original article I was not aware you were the author, so my response above simple responded to your statement in context and not the article as a whole. I'll do so below.
I don't argue that the article presents it's arguments clearly, though I don't necessarily agree with them. Actually, my main issue is with the rhetorical style it is presents it in.
It tries to hard to sell less testing as a solution to testing pain, rather than exposing it as a possible idea to be discussed. Using mostly a combination of appeals to authority and association with the term "lean" as rhetorical devices. This, to me, smells like the last generation "Agile" snake oil.
On separate level I do disagree with presenting logical reasoning as a counter argument to empirical evidence. Logical reasoning may be a compelling reason to seek empirical evidence through trial and experimentation, but reasoning alone can't invalidate evidence. With that said, I do believe some other commenters here mentioned that there is evidence on both sides of this debate, so I won't say your idea's are not worth exploring, though they definitely contradict my own experiences with testing.
Integration tests, in contrast, require attention more frequently -- because not all subcomponents are stable and when one changes, multiple integration tests need superficial readjustment.
I'm skeptical that my own ROI analysis is more "right" -- it probably varies by project and team. The conclusion I draw instead is that ROI analysis is important in general when writing tests, but that one particular ROI analysis is unlikely to apply universally.
My feeling about unit testing / integration test is you do as much of it as you can, trying to sincerely learn the lessons of how to do it it effectively, until you know why you don't need to do it. Or until you know what the right mix is. I started with Extreme Programming and testing around 18 years ago and my thinking has constantly evolved over that time. One of the biggest factors, I think, is project context and the nature of the software you are developing.
I find in embedded systems, unit tests done extensively are very very useful. As is integration testing.
In web systems, I find integration tests are often most useful as so much is all about gluing together various tech, and it's often the gluing the comes unstuck. However there is often chunks that have some interesting logic that lends itself to more extensive unit tests.
Sometimes testing is just a mess because of the design. Too many things that are not really needed, too many layers, too many levels of abstractions, too much "future proofing". Then you see people start trying to throw unit testing around it, and it is mostly just useless and explodes out the amount of work to do. Often I think articles like the original are reactions to this scenario.
Not advocating writing 100% test coverage, but just some code that shows how you will be using the new feature as you add it is enough to help you structure your code.
Most are related to DBs. I NEVER mockup. NEVER.
NEVER.
I just rebuild the DB each time the tests start (and is PostgreSQL. I found that DROP/REBUILD the SCHEMA is fast than drop the database and rebuild, so is ok).
This always leave me with all the data in the DB, no matter it fail or not. This is the main reason to never mock. Mock data get lost after the tests end. The db instead persist the data and allow to inspect after the fact.
Plus, fill me with the data necessary to let the UI run without re-create again here.
This simple rule is what help me most. If I wanna the test to run faster, I can't do anything else that write a good database layer and not call the DB when not needed...
I have a mail server that I test by cloning a root (that's in git) and chrooting there. I also have created a language for testing network protocols once.
Try building a tennis game, just in code with no gui, and take note of how much you would either write tests for things like the score, set points etc., or how much time you spend debugging to justify if your code works.
Unit tests are a great documentation tool as well, both for developers but also for the business needs of the software — the value.
If a test fails and it breaks CI, there should be a problem in the product. Move everything else to some "additional testing" bucket and not block developers.
That means you're either testing the return value or testing the inputs which were modified by the unit.
If the behavior of a unit changes in such a way that it does something else with its inputs, it's a specification change which requires other test cases.
The article makes a lot of assumptions like "In practice, most agree as most projects set the lower bound for coverage to around 80%". Most projects where, which industry, who are these 80% that agree? This is just taken out of thin air in this case. It might be true but should we take the author's word for it?
There are also some psychological arguments. For
example, if you value unit-testability, you would prefer
a program design that is easier to test than a design
that is harder to test but is otherwise better, because
you know that you’ll spend a lot more time writing tests.
This is very subjective without examples. The opposite argument can also be made that code which is easier to test is sometimes better. It felt like reading a collection of quotes and articles by other people.I think TDD, unit, integration and E2E testing works. How much of which to apply is entirely project and industry specific and it's up to teams to decide what testing strategy works best for them.
https://medium.com/@TuckerConnelly/94-gems-from-code-complet...
personally I like write some cursory unit test because it ensures the code logic is correctly decoupled from data retrieval, but most of my unit tests are written from user reports, as part of a large no regression suite.
If integration tests can't cover some part of your code, it's because that part does not matter. Get rid of it.
I also disagree that whether an integration test can cover some part of your code is the litmus test for whether the code is relevant. There are types of errors that are relevant but complicated to simulate in an integration test, or at least in an automated integration test.
An example would be a network partition. It's involved to code a test to cause the partition, wait for the various timeouts to take place, then heal the partition, then wait some more for things to start talking and recovering. And depending on who owns the test environment and how your organization is managed, you may need to negotiate with a DevOps / Operations / Security person to get permission to run iptables in the appropriate machines.
On the other hand, unit-testing what happens when a connection fails is probably pretty easy.
Yet, on the layer that is charged with surviving a partition, that's exactly what you should be doing.
Your unit test can verify that some machine sends the designed message down the wire, it can check if it responds to the expected message, but it can not even check if the software can detect a partition, even less that it can survive one.
I would recommend you make those timeouts configurable from the user point of view.
That's not a refactor. That's a rewrite or rearchitecturing. Of course you have to redo unit tests, you've eliminated the units that were being tested.
> What if you decouple two components that are always used together.
If they're always used together, then they were a unit. Why were they decoupled? See the other front page post about Dijkstra's parable for a rather apropos example of where coupling (perhaps not his intended moral) makes sense.
And we made the choice not to have unit tests. Everything was done through integration tests. What made up for it is that we had very good traceability. So for we knew which lines of code were related to which requirement and which test. So if we changed some code, we knew which test to run, and if a test fails, we knew where to look at.
For the edge cases, we had the ability to inject data directly into memory, triggering the otherwise unreachable error conditions, and the very last cases where done through code analysis (we needed 100% coverage).
On the other side, if you use dynamic lang (Ruby, Javascript, etc), or if you are developing a library or framework, then a good coverage for unit tests is a must. I guess the developer's context (whether he/she is developing a framework, library or a simple CRUD app) plays a bigger role.
Every time I hear this kind of statement, I think of Bob Martin's article: https://blog.cleancoder.com/uncle-bob/2017/05/05/TestDefinit...
> The author tells of how his unit tests are fragile because he mocks out all the collaborators. (sigh). Every time a collaborator changes, the mocks have to be changed. And this, of course, makes the unit tests fragile... the author of the above blog has given up on badly written micro-tests, in favor of badly written functional tests. Why badly written? Because in both cases the author describes these tests as coupled to things that they should not be coupled to. In the case of his micro-tests, he was using too many mocks, and deeply coupling his tests to the implementation of the production code. That, of course, leads to fragile tests.
> What the blog author does not seem to recognize is that first class citizens of the system should not be coupled. Someone who treats their tests like first class citizens will take great pains to ensure that those tests are not strongly coupled to the production code.
Testing is hard for a number of reasons, not least because given five engineers you'll get six definitions of Unit vs. Integration test. I think the term "micro-test" is useful as a contrast; most beginners sit down and write tests for every significant function, be it public or private (i.e. they write micro-tests), and this produces tests that are tightly coupled to the shape of the code at the time of writing.
However if your code is well-encapsulated, and just tested through the public API of your class, then you should be able to make fairly significant changes to the structure of that class without actually changing the tests. (Of course if you change the API then the tests must change; it must always be so).
In the initial run on this code base we went a bit unit test crazy, which was fine at the time, but I'm seeing the drawbacks now where I'm doing a lot of surgery on the components.
I've ended up just beefing up the current integration tests with some extra edge case testing and, after the refactoring work, they still pass, which has given me a lot more confidence.
I think a lot of this comes to head when you unit test the glue logic that orchestrates the entire program, unit testing those becomes painful with mocks etc because they become fragile whenever you do any refactoring work.
I'm not sure what to think really.
My advice is usually to do tdd for new dev have good coverage of full function tests and delete unit tests if the get in the way/lose value on refactoring. No one likes deleting tests but they must add value or its just more debt.
If you can take a 5ms unit test and verify with 80% confidence for a given feature, run that first before running at 5s system integration test that verifies it with 100% confidence.
You run both, because you run unit tests with a build/linting stage (something that can run synchronously with every commit, or close to it) and you run the integration test less frequently because of time/cost constraints (stable branch merges).
Even though they test they same thing, they serve different needs.
"I suggest that the programmer should continue to understand what he is doing, that his growing product remains firmly within his intellectual grip." [1]
[1] E.W. Dijkstra: https://www.cs.utexas.edu/users/EWD/transcriptions/EWD03xx/E...
So, your argument is true for individuals, but for the community it is void. Together we should focus on identifying correct, effective, simple and reliable building blocks on which we can build our larger systems.
Perhaps you haven't seen "Simple made easy" by Rich Hickey, which illustrates this concept rather well.
And in my experience most of the defects which aren't caught by our quality control process wouldn't have been prevented by switching from OOP to FP, or from mutable to immutable state, or anything like that. So I'm not convinced.
I think "Why Most Unit Testing is Waste" is a better article on this subject. https://rbcs-us.com/documents/Why-Most-Unit-Testing-is-Waste...
Here's an isolated overkill example of how unit tests should not be done on one of my codepen files (its a javascript calculator). All tests passed but the calculator is partially broken still. I made about 20 unit test case scenarios when in reality it should have been significantly less (~ about 5 given the complexity of the problem)
https://codepen.io/vincentntang/pen/XqNGqv?editors=1100
I ended up getting very little useful feedback on codereview / stackexchange, but this is how I felt about unit-testing in general.
https://codereview.stackexchange.com/questions/193128/shunty...
I think unit testing is great in some applications (e.g. working with teams as more things can be missed, where a bug can be exponentially cause issues, systems with larger complexity). Doing it in isolated solo projects is such a waste of time IMO, good commenting / structurally organizing code/ naming conventions suffices here. Most things don't work in a functional way anyways that makes unit-testing easier. E.g hardware calculators don't have logic the way I wrote it in my code above.
Unit testing is OPPURTUNITY COST. For a business to implement this the risks of exponential failure and its demand for risk-mitigation should far outweigh the time and developmental practice of implementing it in your training / management / workflow / toolchain.
...disclaimer: I might have no idea what I'm talking about I'm a fairly inexperienced dev
This article is to be about the kind of test you should use. It's just that unit tests are often not powerful, but always expensive. So, you should use other kinds of tests.
https://rbcs-us.com/documents/Why-Most-Unit-Testing-is-Waste...
Unit tests can efficiently handle, for example, libraries, error handling which cannot normally be reached, complex algorithms which are not AI based and require a large number of test cases, fragile pieces code and code patterns, and program design patterns which can be tested using inexpensive common tests. Consider the business value of quick tests in each of these situations. It is high.
While your units have to come together in a way that allows your users to actually use your software, your business logic has to be represented in whole or in part by one or more unit. Testing this logic earlier on in the process and faster than other types of tests is better.
This is probably not the same story for something like a react or swift application. Good read, thanks for the solid points.
Long-time experience of debugging proves that untested code is worthless. And indeed most people start programming by testing some test case or use case.
I would rather recommend starting with use case testing. Data driven testing.
Never believe that you can get a code without testing. I guess I finally need to write the article about my diamond testing routine, so that people have something better than unit test swarms ;-)
But if symptoms would likely be too subtle to stand out, e.g. a typical bug would manifest itself in results that seem plausible despite being wrong, then it's a strong indication for unit tests.
Being in between does not mean it gives the highest ROI or the best balance. They can easily share the worst of both.
> Lean Testing takes an economic point of view to reconsider the Return on Investment of unit tests.
But does not in any way actually try and work this out, which means that you can't make the conclusions that are in this.
> The Return on investment (ROI) of an end-to-end test is higher than that of a unit test. This is because an end-to-end test covers a much greater area of the code base. Even taking into account higher costs, it provides disproportionally more confidence.
And yet no figures or evidence.
Many e2e tests can cover the same path segments, meaning adding a new test may not increase your confidence by much at all. But they're still taking longer to run, and despite the articles insistence that integration and unit tests are brittle, I've found E2E tests are also brittle, just in different ways.
I shouldn't have to change my string normalisation tests just because the website changed moving the resulting output field to a new place on the page. They shouldn't break because the login flow changes.
I have definitely seen (and written) E2E tests that provide almost no real value, and unit tests that provide huge amounts.
> Plus, end-to-end tests test the critical paths that your users actually take. Whereas unit tests may test corner cases that are never or very seldomly encountered in practice
Unfortunately low proportions of errors still can account for extremely large actual numbers of problems. 1/1000 bugs will happen all the time if you've got a reasonable number of users. Those bugs also may happen to most of your users if they're somewhat random, or worse may heavily impact one group.
> For many products, it is acceptable to have the common cases work but not the exotic ones (‘Unit Test Fetish’). If you miss a corner case bug due to low code coverage that affects 0.1% of your users you might survive. If your time to market increases due to high code coverage demands you might not survive.
Or if roughly one in 1000 customers keep taking your production server down due to a bug you might not, and you may have been just fine waiting another week to go live.
I agree with the idea, you should consider what it is you're trying to achieve and what the best ways of doing that are. What you absolutely shouldn't do is use terms like ROI and economic without any analysis, purely to justify not doing something you don't really want to do.
I'm a huge advocate of automated testing and with the available tools, like docker, it's relatively painless to get the pieces together to sort out automated testing. Often the tooling you use to enable automated testing is tooling you end up needing anyway -- it's dual purpose. Before the "first run" of a set of code, I'll create Dockerfiles that make up a complete, local, development instance of an application along with some boilerplate tooling that I include to make debugging easier on me. When setting up the production build, the final version is usually this same environment with fewer lines in the file. Because my environments tend to be similar, I have a zsh script that strips out lines in the Dockerfile to get me 90% of the way to a production container.
For me, it's always worth it. I came to this conclusion after spending a few months forcing myself to test rigorously[0], starting with unit tests written often and early and ending with a small number of integration tests and a much smaller number of end-to-end tests. I don't find any of these particularly difficult to write.
The benefits, however, are vast: (1) Avoiding the debugger time-sink: The #1 thing that I always come back to is that I generally end up never having to fire up a debugger. I noticed that every time I encountered a bug in code that was poorly covered, the first instinct was to attach a debugger and peek at locals to see what was going on. This rarely resulted in spending less than an hour troubleshooting. Sometimes you get lucky and you find more than one issue in that debugging session, but often it landed in at burning an hour on ever bug and way too often it was an hour spent debugging production code and the bug was customer impacting[1]. At the same time, it's rare for an automated test to have a time cost that high. (2) Refactoring - Since "premature optimization is the root of all evil", that necessarily means that a performance bug is going to involve injecting complexity into a running codebase and this often comes with high-impact refactoring. Unit tests, specifically, are incredibly helpful here. This is often an argument against integration/end-to-end test automation since refactoring regularly breaks these brittle tests, however, I've found in practice that this isn't the case at least half the time. Of the times that it does affect those tests, the practice of refactoring can surface subtle bugs (on a few occasions it surfaced a subtle race condition that might have been missed if a few of the integration tests covering a subset of the functionality hadn't broken). (3) Design - more for integration and unit tests, thinking about testing while writing code can result in a less brittle design[2]. On integration tests, it means writing SQL scripts and migrations to ensure that a fresh environment can be spun up on-demand instead of using GUI tooling (or, using the GUI tooling to generate said scripts/migrations). (4) Build automation - I'm somewhat surprised at how often I encounter a customer project where I have to follow a 20 step process to get things functional in a development environment. It seems like if CI isn't involved, people figure a README.md with a mess of shell commands and button clicking is OK. Scripting out environment configuration and build was already one of the first things I did when I began running the code I'd written, however, I find I no longer have to argue in favor of this when testing is involved -- everyone wants a single command to execute tests and once integration and end-to-end tests are involved, it just makes sense to add standing up the docker parts, too[3].
I get why there's resistance to doing these things -- getting people to simply write any automated tests seems to be the most difficult hurdle. Throw in "learn docker" and other technologies to make automated end-to-end testing easier and the barrier is even higher. And hey, there are some times that the time spent writing tests doesn't pan out to a time savings. For the unconvinced, I can only strongly recommend: try it on your next big project. There's no need to change the way you think -- skip TDD if it doesn't work for you -- but write unit tests over your, public facing surface area. Write integration tests over the most important parts of your codebase -- those which if a bug were to be encountered would have the greatest impact on reliability. Write a few end-to-end tests of major functionality. Keep track of the time spent from the first line of code to the final, released, product. If your experience is like mine and the 4 different teams I've done this exercise with, you'll end up doing things this way from that point forward. If not, you have a gift that I lack -- you write incredibly bug-free code "the first time", every time. :) Then try switching to a single 1080P monitor[4].
[0] I tend to code first, test second. Though, on paper, TDD looks like a good idea in that it forces you to think about the desired outcome and write methods in a manner that guarantees the ability to test, I don't find it difficult to write things in this manner from the beginning and I find it more natural to write the actual code first and I'll often write a large footprint of code before writing the first test -- I don't find as much value as others in frequent, instant, feedback but I recognize the value that others find in that.
[1] And bugs resolved during an outage are duct-tape "there, I fixed it" kind of repairs.
[2] Provided you don't like having to figure out new and creative ways to mock complex god-objects/routines. Maybe you dig that sort of thing?
[3] I reload enough that I have a script that automates installing and configuring docker in the most common manner if it's missing or the configuration isn't complete.
Or how about another project that had lots of real tests, that were old, unmaintained, did humongous amounts of work in each single unit test and took a weekend to run, and of course, fail.
Or perhaps the team that obsessed about the testing for each user story, spent all the time doing that, then running out of time to deliver the actual stories.
And those were the less egregious examples!
TDD for music composition: add some notes which jar the ear (red test). Then massage them until they sound pleasing (green test). Make some adjustments for better flow with surrounding notes (refactor). Now repeat. Gee whiz, you're gonna be the next Mozart or Bach!
TDD does absolutely nothing to hinder your creativity. All it does is help you to think about your code's behavior more thoroughly and catch inadvertent bugs. Nobody has so much talent they don't write bugs.
For music composition: composers generally start out with an idea of what they want to achieve, whether at the scale of the symphony, movement, section, passage, phrase, or bar. Then they try different solutions until they satisfy the idea they want to achieve. And ensure it flows with the overall structure. And repeat.
Beethoven didn't write out symphonies fully formed. He spent months and years experimenting with ideas until they satisfied what he wanted, e.g. the "test".
I agree that every software development team has to make their own decision to balance development time and maintenance cost but I don't agree with this low entitlement mentality of writing correct code in this article.
Of course there are also costs when implementing tests and your unit tests are rendered useless on design changes. But so what? You're also decreasing code assurance and it's better to increase it with tests so that you won't encounter bugs in production which is much more expensive than writing tests.
Software development is not only about adding functionality but also about ensuring its correctness.