Most Unit Testing Is Waste (2014) [pdf]
rbcs-us.com
rbcs-us.com
Regarding wasting time on 50% tests... it depends what your system is. In most of the systems (the no mission critical ones, no general libraries) I like to use the 80%/20% rule. I always ask: what are the 20% of the tests that deliver 80% of the value?
You learn with experience but if you do not have it, this is a simple way to learn (as a team).
During the sprint, have people write down:
1 - The tests that actually failed and found a real problem before the code was shipped.
2 - The tests you wrote for the bugs that were found in production during the sprint. If you haven't written the test yet, ask: what kind of test would have found this bug?
Have people in the team share in the retrospective meeting. If you need more time at the beginning you can hold a meeting just for bug sharing. Discuss about the patterns. Commit to write more of these kind of tests.
In a few sprint, you should see regression bugs going down. If you also take a look at the tests than never fail, you will also have a sense of the bugs you should not write.
Let me know how it goes!
Just checking: Did you mean "the tests you should not write"?
I don't personally believe this is the case, but I also can't tell you I know it to be false. What I do know is that types have a non-zero cost, and I think sometimes that's ignored, and only their benefits are acknowledged. But the efforts to bring types to historically dynamic languages show that not all common dynamic language idioms are amenable to reasonable type signatures. While some idioms (e.g. monkey patching) are problematic, that's not the case of all of them.
That aside, that "bit" of typing can amount to a lot. Python programs are usually very compact compared to statically typed languages. Yes, this means writing more unit tests to cover the bits the type checker would have caught, but in my experience the total LOC with unit tests for 100% coverage in Python is still less than a statically types language.
That said, I still prefer statically types languages. The big difference between statically typing and unit tests is when the error is caught. Static typing catches the errors very early. In fact in modern IDE's that do incremental compiles it's literally after you have typed the statement. Errors that hang around tend to ripple through your code because your mental model is wrong, so catching them early means less errors in total. The effect can be large.
The way I see it there are two kinds of costs type systems may impose. Very simplistic type systems, such as the ones in C, Go, and even C++, Java, C#, often force you to write code that is too specific, thus forcing duplication (C and Go especially suffer terribly from not having generics - in Go especially its impossible to write a well typed function that can be passed an array of any kind and return an element of that array).
On the other hand, of your type system is powerful enough, you still get a similar problem when trying to write code that is abstract enough to be reusable, but specific enough to be safe.
For example, if you were writing code for a physical simulation system, you would want to have all quantities carry their measurement unit in their type, to avoid mismatches.
However, you wouldn't want to rewrite maths code for each measurement unit. However, to properly keep track of types in complex maths code gets pretty ugly pretty quickly. For example, a function that wants to compute the scalar product of 2 2-dimensional vectors would only work if the first values of each vector have the same/convertible units as the second pair, in any order (m,s dot s,m should be ok). If you move to matrix multiplication, the types become even more complicated, and this is still very basic algebra. If you want your function to be applicable to arbitrary matrix sizes you're already writing a complex library.
First, static typed languages often simply subtract features that would be available in dynamic languages. So, that's a cost that's paid on day zero, and it can be easy to forget about.
Secondly, there are absolutely times when something that is conceptually easy to describe is a real beast to model in a type system. Think DSLs for SQL, for example. A non-trivial amount of time can be lost fighting these battles.
Dynamic languages are not helpful just because they let you write bugs faster.
Shit at 50% off is still shit. You're not being paid to produce shit, even if you might be really, really efficient at doing so.
Now try serializing things in XML instead of JSON. But not one type of message, multiple types, with flexible formats.
"Oh but this never happens" except it does.
Not to mention the iteration speed whenever you need to fix something.
Apart from both json and XML are a PITA to work with in the first place
How would I do that with a type system?
I also agree, as years ago I wrote a large amount of javascript which had type assertions pretty much everywhere. They didn't take long to write but I'm very glad I did them, and I'd rather have the language do it to save me time, clutter and maintenance.
Lax typing is a real cost. I've been working on some SQL I inherited and even SQL's not-too-bad typing allowed the original coder to mix types where I wouldn't, and create potential runtime errors eg. to assign the contents of a 64-bit integer field to an 32-bit integer field. That's legal and gives no warning. It's also suddenly a runtime error when it exceeds 2^31 (and there are other possible consequences such as screwing up the optimiser). I'd much rather it forced me to match types exactly, or explicitly cast.
In a typed, compiled language you get some pretty nice guarantees just by getting your code compiled. In a dynamic language, you have no way to know if your code even runs until you have 100% code coverage.
Maybe my understanding is wrong, but what you described doesn't seem like a unit test. I think the appropriate term for a test like this at the junction of UI, business logic and code-level triggers is 'feature test', 'integration test', or maybe even just 'test'.
You could use Coq, Isabelle, or Lean. Their type systems are powerful enough to allow such checks.
You can then assert and type against the warning function, the printing function you could only assert against with output buffering or something like that but in my mind that's less likely to go wrong anyway
Keep your type system simple, custom types beyond in built primitives cause head aches imo
Given that transferFunds shows a warning over that 95% threshold, what type of change is going to break that logic? Some sort of botched code rewrite where values doesn't retain their meanings? That batchMode shouldn't trigger the warning? That race condition between the check and the actual transfer?
Most of the unit testing I see is more like test-of-defintions (is fullName still 80 characters wide and doesn't accept null bytes?) which confuses me because a definition can only be correct or incorrect in context.
I accept the usefulness in dynamic languages in the absence of types. My Personal preference is for tests to be a bit higher level (does login with a username containing null still cause an exception?) but I have come to terms with the fact that few people agree with me across multiple organizations, so there must be some point to this trivial testing that is completely lost on me.
These tests are great, but what if username has a bunch of validations on it?
For example, it's reasonable to think a username field might be validated with:
- Must be required
- Within 1-32 characters that match a certain pattern (let's say a regex to limit it to lowercase letters, numbers and -)
- Must not be a blacklisted word (admin, administrator, etc.)
- Must be unique (enforced with a database index)
Pretty standard stuff. Are you going to write 5 integration tests for this? 4 to test each validation and then the success case? These would be tests that exercise your entire web framework's routing stack from request to response (ie. the user visiting a /register URL and then submitting the form).
Personally I would not. I would write 1 unit test for each of those things (4 unhappy cases where I assert a specific validation error for each invalid input and 1 happy case where with valid input I expect 0 validation errors). In this case, the "unit" would likely be a `register_user` function that accepts params as input and either aborts with validation errors, or succeeds by writing the record to the DB.
Then, for an integration test I would have 2 tests. One to make sure with invalid input I end up with some type of error displayed in the HTML response (it doesn't matter which one), and another test with the success case to make sure things work when they should (such as the user is registered and a new record was created in the DB).
So I end up with a tiny bit of overlap in tests. Technically the unit test for the success case doesn't need to be there since the integration test covers it but I usually include it for the sake of completeness because it's usually like 4 lines of code to make that test but I'm not 100% opposed to someone saying it should be left out.
Yes, exactly. It's supposed to catch the rewrite where somebody turns (a * 100) / b > 95 into (a / b) * 100 > 95 and suddenly, the code doesn't work anymore.
Showing a warning is full scenario - it has a whole UX. If this warning is important enough to have a PM, a UX designer, and be translated into 20 languages, it's important enough for the engineer to make sure it actually shows up when its supposed to.
What you describe is a real use case that I personally want to have tested on a real system with all the stuff in place (but test database instead of the real one). It has nothing to do with one unit you can test on its own.
And the crucial problem here is that you tie your testing code on the inner workings of transferFunds, when you are mocking stuff.
> and verify that the warning is triggered when the value being transferred is over 95% of the amount in the account.
Please also note that you cannot verify this property using testing. You can only verify that this works for one particular set of values (or a big but finite set of values). Testing never can verify that something always works. That's the wrong tool for this job.
I, as a user of your software, only care about the feature, and that the feature works. I don't care about testability, whatever that means. And since I case about features, I think this entails that features have to be tested end to end, and not some unrelated classes ("mocks") that are not even used in the final product.
> is easy and makes the resulting code better (eg, easy to replace the notification manager).
It might or it might not. I have seen too much code which was written way too complex which a lot of unnecessary classes and interfaces, just for the sake of "testability". However, some features did not work every once in a while.
If you have the need for different notification backends, go ahead. Even if you need a notification manager (whatever this is) for this. But don't make code more complicated (sometimes called "test-induced design damage") without need.
Or are you talking here about completely different type system ?
"Software development has been characterized from its origins by a serious want of empirical facts tested against reality, which provide evidence of the advantages or disadvantages of the different methods, techniques or tools used in the development of software systems."[1] In my view, the best thing we could do is to adopt "Evidence-Based Software Engineering"[2] as other disciplines have. This is more likely to have a major positive impact than the newest and hottest language, tool, or technique.
[1] Reviewing 25 Years of Testing Technique Experiments. https://www.researchgate.net/publication/220277637_Reviewing...
[2] Evidence-Based Software Engineering. https://dl.acm.org/citation.cfm?id=999432
This is the norm in the industry, "you're doing it wrong, the best way is this way..."--based on what evidence that pertains to this case or has proven generalized success applicable here?
In most cases, it's someone's successful anecdotal experiences which worked for the specific cases they were involved with. That doesn't mean those approaches can be abstracted away and generalized to all cases, but many in this industry do that regularly and critique others' approaches based on that. It becomes this competitive ego contest: well my work was at x, y, z solving q and was successful--making me an authority on a, b, c's problems solving p because x, y, z is a leader (appeal to authority fallacy)... etc.
It's one thing to treat development as much of an art (which to me, it very well still is) but once you start treating it as a concrete discipline, you need to provide the lyme, cement, aggregates, water... evidence and studies showing approaches and how they faired across controls and varied cases.
On one of his visits he started interviewing engineers individually to see how things were going and I asked him what the point of it all was and I could tell he was quite taken back. But then, not surprisingly, he said to be more efficient developing software. I then asked him if he or any organization he'd worked with had ever actually tracked the time it took to implement the process, to which he answered no. Then I asked him, then how the hell do you know if any of this is making software development more efficient?
Though, to counter my own counter-point, it would be nice to have better analytical tools; more informative & accurate ways to understand & visualize how resources are actually being used when a program runs.
I'm in automotive [0]. There is somebody who has to sign that the software is safe. Without someone to make that signature the car does not get released.
For Free Software, look at Sqlite. There are some nice slides from 2009 [1].
In general it results in some simple rules like "100% test coverage". Of course, these rules (while simple) are not easy to satisfy and certainly expensive.
[0] http://beza1e1.tuxen.de/aspice.html [1] https://www.sqlite.org/talks/wroclaw-20090310.pdf
Unit tests for the most part aren't about "testing". They are developer tool. To verify rthat modifications (refactoring, additions, bug fixes, etc) doesn't break contracts etc. Oh and showing that your code is a codependent mess of poorly isolated spagehtti, if your unit tests are hard to write, the code under test is a mess. Inittests are more useful in languages with loose or poor type systems.
Many unit tests are not written well. testing more than the interface/contract. Or full of complexity boilerplate and mock. Which means code needs fixing.
We moved to a behavior-focused testing style, and it has been much more robust to non-functional changes and refactorings in the code.
Part of the problem is that most test fake libraries support mocks (usually they even have it in the name), those mocks have complex “verify the 3rd call was function doDoAction with parameter (‘yes’, ‘really’)” and then people think that full-blown mocks are the only way to guarantee 100% correctness.
Just think about it at a high level. If your module depends on another module, is it really always necessary that this module be substitutable for another module that has the same interface (e.g. another similar module or a mock)? 99% of the time, the answer is no! In fact it's often a problem because this wrongly assumes that if a class exposes the required interface, then it is compatible. There is more to compatibility than just interface; if the submodule maintains its own state (OOP), then its behaviour can change in a way which could break the dependent logic. You can't have Dependency Injection for everything because it's never so simple; these customizable parts of the code have to be designed very carefully.
A unit test should only test a single class so that means you need to mock out all other classes which your class depends on. Mocking out dependencies in the test code is difficult or not feasible and that's why developers often resort to Dependency Injection in their source code but as mentioned before, it is an antipattern.
Also, I noticed that a lot of developers confuse unit tests with integration tests but the definition is quite clear: If your test covers more than one class without mocks, then it's an integration test. Integration testing does not necessarily mean end-to-end testing. It could just be a single class with its internal dependencies.
The only reason dependency injection gets a bad rep is magical frameworks which obscure the actual wiring and end up causing bigger problems.
I think that the class which ineteracts with the database directly via the client should be tightly coupled to the database client. It's not very often that you change database and when you do, you can just swap out that entire class completely. Classes which interact with the database should expose simple interfaces for performing actions against the database and those wrappers should be replaceable.
In my experience, after code is checked in, unit tests have two main purposes:
1. Checking that functions aren't completely broken in dynamically typed languages (i.e. a check so basic that a static type checker can do its job) 2. Allowing programmers to refactor code without being terrified it will break something far away in the execution path
And like, that's it. That's enough to make them useful to have around.
Unit tests are supposed to work with the internals of a module and not just its boundaries. Since refactoring normally tries to significantly change the internals without changing the boundary, it usually also results in a rewrite of the unit tests. Integration tests are the ones that usually help me sleep better after a refactor.
Some people argue that a unit test of a module should mock out its dependencies and not rely on them. For unit tests like that... well, I endorse the article.
Now because you have to test it anyway, you can just simply spend the same time writing a unit test instead of executing the code with manually configured non-repeatable test cycles, a.k.a clicking through the UI, or sending Postman requests, etc.
Also it's kind of selfish not to make something repeatable by others.
And as someone pointed out before me, unit tests are not about testing at all. It's a documentation about the system, and what it's supposed to do. Also it's the way to stop the next person who works on the code to ruin something by not knowing the business rules.
First, in most cases you're going to test your code manually anyway- whether you write unit tests or not. So writing unit tests is just in addition to the time already spent testing manually.
But most importantly, unit tests often do very little to prove "your code actually does what it claims to do". They prove you have tested some cases, that's all. In many occasions, the code under test doesn't even "claim" to do anything in particular: the claim is that you're delivering a feature, not pieces of code that behave as expected. All your code fragments can behave exactly as expected and still the feature might be broken, or miss implicit requirements, or ill conceived.
Finally, unit tests as documentation are disastrous, as they document only details of specific implementations, as opposed to the feature you're expected to deliver. Good comments are a hundred times better that tests in that respect.
Only if you are coding things with UI's. If you are creating a new API endpoint on some service or creating a new service or adding some new parts to a maths library there is no point in manual testing.
I feel a lot of the disagreements people have on development methodologies comes down to people working in different domains. When I do UI work I do mostly manual testing, when I do library work I do only unit testing.
And if it's some poorly coded legacy application or something that's mostly wrapping a proprietary black box, working around other peoples bugs, and things like that then system or integration tests may be easier than attempting to mock a gordian knot's tangles.
Unit tests work best for properly encapsulated pieces of code, utility classes, libraries and such.
The arguments in the paper are anecdotal and I took them to be “question the cargo cult” which is a valuable thing to do. The author is merely encouraging you to think because “automated garbage is still garbage.”
Good read.
Does having tests make it any different? You would need to test those tests to make sure they are testing the right thing after all (ad infinity).
Also, tests are actually subject to very basic testing in the red-green-refactor TDD cycle: they start out failing so you know that they actually test something when they start passing.
While the tangible benefits of unit tests are very important, there are other intangible benefits that are equally important.
The single most important reason to test your code seems to be that it allows you to refactor your code while ensuring it still meets the original outside expectations.
The author’s main gripe seems to be about tightly coupled tests that make refactoring larger systems more difficult, and about prioritizing meeting arbitrary metrics (lines of codes covered) rather that thinking critically about the actual benefits.
This is in line with his emphasis on integration tests determined by the business. I think that same thought process can be applied to internal components of code as well, and you can empirically determine the quality of your approach (roughly) by evaluating how often your tests change. If nearly every change breaks a test, that probably means they’re low value/too tightly coupled.
Excessive mocking seems to be the biggest source of evil in that regard.
The confusion here between unit testing and automated testing in general kind of illustrates the author’s point: we’ve become so obsessively focused on one kind of testing that we aren’t thinking critically about its alternatives (of which there are many others besides “no tests”).
Cannot tell you how many times I have seen projects with hundreds of green unit tests at 90% line coverage and yet regress on, miss important and obvious cases, or simply have never had their headline functionality.
Sounds more like a condemnation of OOP than of unit testing, and I do genuinely feel sorry for the unit testing, OOP purists out there. I prefer to design more functional methods, which operate on parameters and injected config instead of instance state and/or globals (cringe). Incidently, this approach makes full coverage attainable.
"Large functions for which 80% coverage was impossible were broken down into many small functions for which 80% coverage was trivial. ... Of course, this also meant that functions no longer encapsulated algorithms. It was no longer possible to reason about the execution context of a line of code in terms of the lines that precede and follow it in execution"
I can reason about such an implentation MUCH more effectively by glancing at the small bit of higher level code which integrates everything (and as mentioned above by foresaking instance state and polymorphism). This strikes me as a bit like advocating a flat directory structure cause it's important to be able to see all your files at once.
Robert Martin sketched 2 diagrams in [1] that elegantly illustrate these two different design patterns and how testable usually means composable and more tractable:
[1] http://blog.cleancoder.com/uncle-bob/2017/03/03/TDD-Harms-Ar...
- Trying to use coverage percentage metric as a sign of quality. As if a simple percentage means anything about the quality of the tests. It's as useless as using lines of code as a way to measure progress.
- Not recognizing that useless unit tests are harmful for the code maintenance. It makes refactoring code into better structure difficult and developers just give up because it's too much work to fix the tests that weren't even providing any value in the first place. This ignorance is expressed in statements like how there's a testing "pyramid" with unit tests at the bottom and end to end tests at the top. Which is a nice sounding soundbite and image, but is useless. Forget pyramids, just write tests where it makes sense.
- Code review comments that try to look smart with "where's the unit test?" Then the developer doesn't want a long drawn out fight about how a unit test would be useless here since there's a huge crowd just cargo cult yelling "code coverage!" "unit tests are good!" "pyramid!" So the developer just writes the stupid test to get the code merged. This is also an example of how harmful code reviews can sometimes be, when there's a popular stupid idea, code reviews perpetuate it because the developers who know better just get tired of fighting the same fight over and over again.
I really hate useless unit tests. I've seen tests that setup a mock, configure it to return a value when called, call the code using the mock, then check the returned value is the mocked value. This tested absolutely nothing! I've seen unit tests that verify every line of the method was called, completely pointless. The point is supposed to be to verify that for a given input, it has a given output, not lock it into a specific implementation by verifying every line of code ran a certain way.
There is an idea for structuring software "functional core, imperative shell". Write the software this way and the natural place for unit tests and integration tests becomes obvious. But nope, the industry is all about unit test coverage percentage, stupid pyramids, looking good in code reviews. It's all quality theater, not actual focus on quality.
I've had a similar disagreement about the validity of such a test.
The other mocking issue I've seen, is people mocking blackbox third-party APIs over which they have absolutely no control, which sometimes leads to passing unit tests, failing integration tests and head scratching all around.
Mocks often contain the same bad assumptions and misunderstandings about the mocked API which the developer used during the implementation of the unit they are trying to test.
If you feel the need to mock something then you should first ask yourself whether an integration test can do the job for you. Actually, I would generalise this advice to: Don't write a unit test when you can write an integration test.
The emphasis here should be on the reason for splitting up functions. Long, complex functions can be difficult to understand, and removing a few lines in exchange for a (well named) function call is very beneficial for the reader. The opportunity for testing comes from this delegation. A function call is a contract, and the test ensures it complies. Now the reader can comprehend what the code is doing at a higher level; trusting the sub functions do what they intend.
* Integration tests that test the API I get a ton of value out of it though it's sometimes hard to guarantee a certain state at the API's end. But I use them while building the API's but they also can detect any kind of problem after an API upgrade etcetera
* Unit tests that test data transformations This is stuff like date formatting, building headers into a request given certain input, building more complex ViewModels that take a few structs as input and turn them into something that reflects the actual process happening in the view. They're valuable while building the logic, help separating the logic since you need to make it testable and also help a ton if you find a bug since a bug simply means adding another test case to see if something goes wrong and then fix it.
I don't think I even get to 30% code coverage but I think what I cover is super valuable and the other 70% is usually mostly CRUD boiler plate.
Granted, that's from a unit test perspective. Integration tests of APIs are invaluable. I wish we had them already where I currently work.
As someone who's been working in corporate integration for the past 6 years..
Never trust a contract. Be in WSDL, OpenAPI spec, Word documents or otherwise. I've worked with large tech vendors, I've worked with finance, I've worked with large consultancies. The only people that seem to get it right, is the people you don't want to work with because it's soul-crushing - think HL7 et al.
That's why I value integration tests over most thing. I can see immediately that something is broken at a high level, what business impacts it has and explain what systems are affected.
The logical contradictions eventually overcame me.
The point is that when you do integration tests, you will test the underlying classes the way in which it actually matters. You can't test every possible condition a unit can have, but you can test for the most likely.
IME unit tests can only effectively substitute for integration tests where you're testing logical/algorithmic code with simple function inputs/outputs.
If there were enough code involved in a test that we could meaningfully refactor it while keeping the test green, we would call it an integration test.
When testing NAPI like to use edge cases like the company with the longest name eg "Donaudampfschiffahrtsgesellschaft" or the Famous Welsh location "Llanfairpwllgwyngyllgogerychwyrndrobwllllantysiliogogogoch"
I am pretty sure I crashed an A17 clearpath mainframe by doing aggressive testing
* Have at least one unit test for each nontrivial unit, but no unit tests for the trivial ones. But, in these unit tests, use mocking and stubbing only when using real objects isn't feasible, or very inconvenient - isolation is nice to have, but generally overrated in my experience
* Create as many unit tests as necessary to feel confident about the edge cases for the most critical units
* Implement higher level tests covering the main use cases for the interaction of all the units
> In a given computing context, the exact function to be called is determined at run-time and cannot be deduced from the source code [in OOP languages] as it could in FORTRAN.
This is not necessarily true, and I don't think this was true when this was written, either. (The Wayback Machine first saw this in 2014, and the paper doesn't date itself.)
Most member function calls in C++ (non-virtual ones) can be resolved at compile-time. Where idiomatic C++ could would use a virtual function would require some equally uninspectable construction in any language, because one would use it when one needed run-time switching. Rust lacks OOP in the usual sense, but if I needed a virtual function of sorts, I might reach for a trait type, which also wouldn't be inspectable (beyond whatever semantics the trait establishes).
There's Java, JS, and Python, I do admit. And I find my unit tests often fill in the static analysis I wish I had. But the truth is more nuanced than I think the author conveys.
> Unit tests are unlikely to test more than one trillionth of the functionality of any given method in a reasonable testing cycle. Get over it.
> (Trillion is not used rhetorically here, but is based on the different possible states given that the average object size is four words, and the conservative estimate that you are using 16-bit words).
This is as vapid a definition of test coverage as the line coverage the author derides in the prior paragraphs. A trivial function,
fn add_these(a: u64, b: 64) -> u64 {
a + b
}
could not be adequately unit-tested, because I could never cover the many trillions of input states it has.I think there's always an implicit line to "well, if the code under test is absolutely nuts", it's not the test's fault for not catching it, and that if the function under test (with the same contract as the function above) is errantly coded something like,
fn add_these(a: u64, b: 64) -> u64 {
if a == 11771008849893880921 && b == 5668622331333919113 {
0
} else { a + b }
}
then what amount of testing is going to catch that?> I define 100% coverage as having examined all possible combinations of all possible paths through all methods of a class, having reproduced every possible configuration of data bits accessible to those methods, at every machine language instruction along the paths of execution.
While this somewhat contradicts the earlier example of states. I don't think we should aspire to this, since I think many "combinations" are not interesting to test; they're just pieces of the code that really don't interact; testing them is redundant. Yes, this hurts our ability to formally define a protocol for testing or not testing, but it prevents the inevitable combinatorical explosion.
> Business people, rather than programmers, should design most functional tests.
Absolutely not. A large part of my job is coming up with the actual requirements of the system from the vague machinations of business people who barely know how a computer works, let alone can specify salient requirements. How are they to write the tests, when they can't clearly articulate the requirements?¹
> Turn unit tests into assertions.
But I want to ensure these don't fire under normal circumstances. We do this, in my current code base, and we get notified of them, and of any crashes in general. But it's a failure that was noticed, and might have downstream consequences, as opposed to one caught by a unit-test.
Then the parts about Eastern Europeans with bad Internet make better programmers because they had to think instead of using the Internet, and how he grew up under similar circumstances. Real Programmers!
¹I'm not saying business types "should" (or should not) be able to articulate requirements/specifications.
Additionally, some fuzzing tools can actually to branch analysis and target your "rare" codepaths. Now, sure, figuring out how to hit some codepaths basically requires brute force - it's not going to magically figure out an efficient way to collide SHA1s if that's what it takes for your test to fail - and you need a way to differentiate good from bad behavior (this might be a reference implementation, or this might be as simple as "does it crash?") - but you do have to get crazier than add_these before you hit the limits of even existing tooling to catch bugs.
Mutation Testing could. :) https://ai.google/research/pubs/pub46584
At least, I didn't take away more than 'throw away tests that haven't failed in a year' and I don't agree.
Redundancy vs. Dependencies: it's the dependencies which kill you. Redundancy is often a good thing.
If you state an algorithm only once, as the implementation, then the next programmer only knows what it does, not whether it's correct.
This is, in my opinion, the main value of unit tests: state the algorithm twice, once as implementation and once as expectations, and if they don't agree then something's wrong. While the odds of a bug existing in any line of code haven't changed, the odds that the exact same bug exists in both sides are much lower.
Any bugs which survive that are probably in design rather than implementation, ie. my mental model of what the module should achieve is wrong somehow. Catching that is a job for integration tests.
I definitely agree with the 80/20 rule here. 100% code coverage is neither necessary nor desirable, and 20% is fine if it's the most valuable 20%.
My second theory would be that the unit tests are written by the inexperienced developers while the good developers write the other code.
Isolated units that are well tested usually don't need to be updated in my experience so they rarely break.
I feel like there is some subtlety here that most will miss. Specifically, "and consider". That is the part we are really bad at.
Tests also help in general development. Make some changes, then fix the tests. If you are running tests locally on your machine (and offline) how can you be sure that every time a test fails locally the failure is logged?
I've lost count of how many times I've gone into an old test suite at work and found that the tests were still passing even though the code they were testing had completely changed or been removed.
Sometimes tests are written very poorly. The codebase benefits from their removal.
But I've also never written a unit test that didn't expose bugs.
That's mainly because I know what parts of the code will be dicey, and focus unit tests on them.
The goal is to prevent a future programmer, including myself, from breaking any of the declared properties.
Too often I’ve seen “tests show it’s correct” suites horribly fail to provide value when the behaviour changes in a unit of business logic and the brittleness needs to be unwound and replaced with robust assumptions.
If I'm doing I/O on untrusted data, of course I'm going to fuzz test it thoroughly and probably find bugs.
If I'm writing new RAII types - containers, smart pointers, etc. - I'm going to write tests in an attempt to exercise double frees, or any kind of rule-of-3+ violation I can think of, and I'll probably find at least one thing I overlooked if it's complicated enough. The people building code need to be able to trust some of their foundational tools, at least, to be nearly bug free.
If I'm abstracting system APIs into a cross platform representation, I'm going to write unit tests to compare behavior, because the abstraction is probably leaky and fails to fully abstract in some edge case - possibly due to bugs in one of the N system APIs I'm targeting, or possibly because I simply forgot to implement a codepath, or possibly due to strange edge cases.
If I'm writing SIMD abstractions, I'm going to write some basic math unit tests to catch the occasional compiler codegen bug for the less thoroughly used and tested intrinsics.
If I'm writing code that I simply know to be fundamentally brittle, I write tests to catch when it eventually breaks. For example, I wrote some unit tests for rust, to catch when the module standard types are implemented within internally change, which will in turn break the .natvis files relying on those internal type names.
If I'm writing lock-free multithreaded code, I'm going to write a metric ton of tests to try and suss out edge cases, because I know a single bug slipping through can result in weeks of time lost to chasing heisenbugs, and I know my reviewers probably can't catch everything either, and you'd better believe it's going to catch a bug at some point.
I'm always weary of arguments that rely on 'unit tests are more complex than the code'. If that's true, then it's correct. Tests have to encapsulate the software contract that's being enforced and should be more complex than the underlying code.
At the very least testing helps communicate to a code reviewer "hey this thing does what I think it does".
I do a similar thing with the parser tests -- have a test for each valid complete symbol production, and then tests for as many error cases that the parser can handle for that symbol.
With that base, I can add additional tests for bugs and additional error recovery cases I implement.
I never see unit tests as a waste as they provide a set of regression tests that are invaluable for refactoring and making other changes to the code, like implementing new features or better error handling.
What I've found in my time is that unit testing can be good, but like anything it's not a panacea. It requires discipline, and like normal code, it has code smells.
Black box unit tests are the most likely to be good tests, and white box unit tests are the most likely to be bad tests. The more you depend on the inner workings of a function in order to test it, the more likely it is that you are coupling your test to the implementation rather than the purpose of the unit being tested. Once you tie to the implementation, refactoring becomes a LOT harder, because changes will break the tests even if they don't break the functionality.
Mocks are also a major source of trouble, and more likely to be a code smell. If your tests are using a mock to test how many times your unit called it, either your tests are bad or your architecture is wrong.
There are three main kinds of code:
- Code that fetches data
- Code that stores data
- Code that transforms data
Mocks are necessary when you mix these. If you have a function that opens a DB connection, fetches data, transforms the data, and then stores the data, you now have an extra problem to deal with (the database), when all you wanted to do was test the transformation. Things would be far easier if you separated the transformation out, tested that in isolation, and then integrated that encapsulated functionality with fetching/storing code. This also improves separation of concerns and code duplication, since now your fetching/storing code can be generalized and also tested in isolation.
Actually I lie. There is a fourth kind of code: code that modifies state. This is the evilest, smelliest code around, and it's also something that unfortunately we can't get completely away from. But we can manage it, by isolating state, reducing the need for or scope of the state, and providing "configuration object" function entry points to make testing these monstrosities less nasty.
Code coverage is not just a measure of quality, but also of waste. If your code is not being called, then one of three things is happening:
1. It's error checking code for another API it's calling, which you normally shouldn't be writing tests for (unless that API is known to be buggy and you need to guard against it).
2. It's not contributing to the goals of the program, and can be taken out.
3. It does contribute to the goals of the program, in which case you need a test for it.
You can't reach 100% code coverage because of (1). But you absolutely should check WHICH code is covered in your tests because of (2) and (3). Anything higher than 80% coverage is pure luck, and tells you nothing about quality or wastage. In many cases, even 60-70% is sufficient.
There's a limit to this, of course, but keeping this principle in mind in your design also has the effect of bringing coupling to a minimum or even zero in some cases, and promoting re-use of these now completely independent components. When components don't care where or how they get their data, they become a lot simpler due to the elimination of state (implicit and explicit). Elimination of state also facilitates making functions idempotent, which massively reduces cognitive load when reasoning about a system.
If you are depending on another component of the application (a lexer, a JSON class, a maths function, etc.) don't mock that because if that class breaks you want tests to fail instead of silently passing because the broken class/function was mocked.
If you are depending on thirdparty libraries, don't mock those unless you have to (i.e. if you cannot run the tests). This will help avoid unexpected bugs after upgrading libraries or if supporting different versions of a library.
If you are writing code for a complex infrastructure (e.g. an plugin for an IDE), try to use test-specific versions of enough of the infrastructure to get the rest functioning, and use the real versions of as much as you can. This will help pick up issues in your code when new versions of that infrastructure make changes -- you want your tests to fail in this case, as the real code would fail.
Understand why your tests are breaking and address that. Use @Ignore as a last resort. If APIs or behaviour has changed, update the code and tests to reflect that. If you are supporting different versions, create version-specific compatibility layers.
Same goes for application code in the same service or whatever. Mock those calls and only test your unit of code.
If you don't do this things can be fine. But at some point the code base will become unwieldy and changing code in one place will break tests all over the place.
Integration tests are fine and they have their place but are not a substitute for proper unit tests.
But the problem, is API change is very often, and we don't want to change both implementation and test just for the sake of API change.
Integration test is enough in most case.
Even API is the implementation details of an abstraction.
It seems clear that the author has a strong opinion on this and perhaps that has been formed by exposure to unit tests done wrong. I suppose his article is worth reviewing and asking if any of my unit tests suffer from the problems he identified. I think most of his complaints relate to badly applied unit tests and I think we can all agree that any methodology can be badly applied. That does not provide grounds to condemn the methodology.
Having used Fortran (and Macro-11) as the first languages I employed professionally, I do not recall any thrust for unit testing at that time. Maybe it was just the shop I worked in. More recently I have used unit tests for Go, C/C++, Python, Perl, Java, shell scripts and probably some I'm forgetting. I wouldn't consider coding anything w/out some kind of unit test.
Back to the quote "Testing does not increase quality; programming and design do", I disagree vehemently with the claim that testing does not increase quality. While true that one cannot 'test in' quality, I find that designing code to be testable provides higher quality results. This is particularly true in languages such as shell scripts that often start small and grow until they are hundreds of lines of in-line code. Testable code is generally better partitioned and structured than it would otherwise be.
A second benefit to unit testing is immediate feedback and completion of parts of the system. I get a feeling of accomplishment when I complete something that passes its tests and prefer that to deferring satisfaction until the entire thing works.
Finally, if I test the bits in isolation, I can provide them data for all of the corner cases I think could cause trouble and make sure they work for a wider range of inputs than could easily be done during integration testing. When I do get to the point where I put the bits together, I have a much higher success rate with integration testing.
https://www.reddit.com/r/ProgrammerHumor/comments/5pbl2q/two...
This is very much in line with DHH's opinion about testing in general (the discussion "TDD is dead" is about this). He said that he doesn't want to split logic for the sake of testing, so HE drives the design NOT tests. I very much agree with this. I think a human can design better code (meaning that's easier to use, so easier to consume by other humans) than any automated process (like TDD).
I noticed his blog post quotes this PDF: https://dhh.dk/2014/tdd-is-dead-long-live-testing.html