TDD, where did ‘I’ go wrong
frankcode.wordpress.com
frankcode.wordpress.com
I am not convinced of this argument. Obviously, it's important to heavily test the endpoints of any API you make public. In fact, that's a situation where you want 100% coverage. However, testing the endpoints of an API does not allow you to forego testing its internals. At best, a passing suite of API endpoints should give you reasonable confidence that the internals are functioning properly.
An API is an abstraction over many moving parts. Any good API will consolidate thousands of lines of business logic into a few endpoints. There is a lot of room for error in the layer of abstraction between API and internal logic. It's entirely possible that a API endpoints could appear to be functioning properly, but actually be relying on broken internal code.
For example, consider a fruit basket API. You can insert fruit into the basket, and check what fruit is in the basket. A suite of tests for the API endpoints could insert fruit, and then check that it's there. In this case, the API is hiding a lot of internal logic. Storage mechanisms, data persistence, fault tolerance, and a slew of other logic decisions are completely opaque to the API consumer.
What if the internal code incorrectly stores the fruit in a temporary file? The API test will pass if it inserts fruit and then checks that it's there. But it is not going to check for that same fruit in an hour. What if it's gone?
Internal logic like data persistence is (rightfully) opaque to API consumers. That means that no testing of API endpoints can validate all internal logic. Therefore, you cannot ignore testing the internal logic in favor of only testing API endpoints.
Basically anything more complicated than a "unit" is impossible to test for every permutation and as a result you get bugs when things happen in production that nobody anticipated a test for.
In short, been there, done that. Yes, way easier to write less tests that don't test everything. Witnessed plenty of six figure+ bugs as result.
Meanwhile, if refactoring means breaking a load of tests then you're doing it wrong. TDD is actually pretty hard to learn and it seems to me most people give up when the have to refactor code that they've tested internal details. Given their knowledge they conclude TDD is bogus, rather than wonder if perhaps their tests are bogus.
In your fruit basket example, there are two fixes for this. Either improve the public API tests to verify that an hour later the file is still there. Or you use a storage technology where you can assume it will work correctly (Postgres, S3, etc) and then have a code review.
There is definitely an exception to "only need to test the public API". That is when a very complex component has to be built from scratch or almost scratch in order to support the public API. For example, I personally need to build a Lucene-based search server (Solr and ES don't fit my use case). While I shouldn't need unit tests for Lucene itself, I do need unit tests to verify the threading model, file structure, and couple other things I write from scratch are correct because the complexity is so high and I don't trust myself or anyone else to write it correct the first time.
Genuinely curious - how would you write a lower-level unit test to account for this? If my "testAddFruit" method directly hits the temporally flimsy datastore, how is that any better than the API doing the same, albeit through multiple abstraction tiers?
Otherwise, I quite agree. I don't write unit tests. A complete set of integration tests gives me about 80% confidence in the correctness of code. Unit tests might bump that up to 90%, but you'll never get to 100% from testing alone and the substantial effort necessary to get there just isn't worth it.
Testing follows the classic 80/20 rule. Integration tests only take 20% of the effort required for full unit test coverage, but can give you the bulk (80%) of the benefits. For all but the most sensitive of applications, it's probably not worth it to put in the additional 80% effort for a mere 20% gain.
When I write/modify a piece of code, I of course must see so it works. I run the new code, feed in some data and see the results.
If I separate UI and backend/model, I can just write code that feed data to the model to see so it works (instead of doing it ad hoc, by hand). Then I save that as a unit test. It is part of the documentation too (a use case).
Cheap, easy and with good value for effort. (Depending on problem domain.)
Edit: I can see where the "test induced design damage" comes from, mocking is bad, but I think code also often get better when it is made testable (dep injection, think about fan out/in, etc)
That's all easily accomplished with integration tests. Generally, any good architecture makes a separation between the UI and the backend. You should absolutely have integration tests for the backend by itself, but I don't think it is necessary to get to the level of unit tests (testing every single method used in the backend).
Personally, I try to separate the backend into its own separate service. The first step is to write tests against the API for that service, and then to make those tests pass by completing the service. This has the added benefit of letting the tests serve as a spec documentation, which makes it very easy to farm out implementation to employees or contractors.
But sure, it is a discussion of what we should call the useful tests. I have (also) burned out on > 80% coverage for tests when the specifications aren't written in stone.
The problem that the article doesn't address is if your test fails finding the failure point isn't as straightforward as it is with smaller units.
However the points made here and in the discussions linked are valid and I think this is the right approach to testing. But I think that there's still lots of value in writing and maintaining tests around smaller units of code. Balancing the two approaches brings us solid test coverage along with easy to debug test failures.
I think unit tests do have one important use: testing library functions, e.g. cryptographic primitives. Most people aren't writing "library code", though; they're employing it.
It's correct that by reducing your number of units and making them more broad scoped, you potentially reduce the number of undetectable faults you are catching (but need not), but you also make crafting the input that detects the remaining faults much more difficult.
The correct approach is (it seems obvious to me) to mix the two. Don't force tests to be on a per-method basis, but also don't neglect to individually test those methods that can be better tested by individual testing. It seems like common sense - following a policy such as this so rigidly clearly leads to either wasted time and inferior APIs or missed test cases - the former I believe the OP has discovered already, the latter I'm sure they'll discover in due time.
http://www.quickmeme.com/img/0f/0fb4fa35fad1b9ed112dc7584f47...
(I think the writeup in the Ship It! book was better than this post)