Why Most Unit Testing Is Waste [pdf]
rbcs-us.com
rbcs-us.com
So is a big/bad test suite a burden? Sure. But is it a burden compared to maintaining specifications of other kinds? Is it a burden compared to working in a system with neither type of specification?
Further, the people writing long articles like this are (or at least were) very good developers. There is often an element of "good developers write good code, so just use good developers" in them. But writing good software with good developers was never the problem. The problem is making software that isn't terrible, with developers ranging from good to terrible, and most being mediocre.
We tend to put a high value on code review, yet some don't value unit testing. Lean in and I'll tell you a secret:
... unit testing is automated, repeatable, code review!
1) A relaxed requirement was agreed to by the PO, but the remote developer wasn't aware of that and wrote code to fulfil the original complex requirement. The implementation was harder to understand and touched more areas of the code, leading to issue number 2.
2) By analysing the interactions between multiple units, I was able to determine that the code as implemented was in fact reading a setting too early, when it was not available, thereby having no effect at all compared to the already existing code. Ironically, this exact concern had prompted the negotiation with the PO which resulted in the relaxed requirement.
Interestingly the error was not caught by unit tests, because it didn't happen in a single unit. It wasn't caught by integration tests, because that particular online component was mocked and it passed code review by two other developers.
Except when it isn't.
"Unit tests can be one form of code review (that also happens to have the advantage of being automated and repeatable). But should never be mistaken as substitute for actually understanding what the code does, why it does it and how it got that way. Which requires an incompressible amount of effort, analysis, and plain and pure grit", would be another take on the matter.
Unit tests are better than nothing, but integration tests usually serve as a better specification still.
Unit tests are ok at specifying low level side-effect-less modules though.
In addition, if you depend on unit tests, you cannot tell if a change has broken the system - you can expand or rewrite your tests to cover the change (if we put aside the question of what requirements these tests are testing), but that cannot tell you if you have, for example, violated a system-wide constraint. Only testing at higher levels of abstraction, up to the complete-system level, can help you with that.
Therefore, to the extent that the "the code is the specification" argument is a valid one, it actually leads to the conclusion that Coplien is right.
The most famous property-based testing framework is QuickCheck, but there are plenty of others, e.g. FsCheck for C#/F#, ScalaCheck for Scala, Hypothesis for Python, etc...
The general idea being that you write out a specification that the code should meet, and the test parameters are determined at runtime, rather than being hardcoded. There's more to it than that, but that's the general idea. If you want a simple introduction, I'd recommend this video:
1.4 The Belief that Tests are Smarter than Code Telegraphs Latent Fear or a Bad Process
Your other point about only good programmers not needing unit tests is moot as you haven't followed it through to the conclusion. Namely, if bad programmers write bad code that needs unit tests, they're also going to write bad unit tests that don't test the code correctly. So what's the point?
> Software engineering research has shown that the most costeffective > places to remove bugs are during the transition from > analysis to design, in design itself, and in the disciplines of > coding. It's much easier to avoid putting bugs in than to take > them out.
Everyone knows that. It's just beating a dead horse. But companies demonstrably want to ship buggy and feature rich software fast, rather than ship well designed and lean software. And b) don't want to pay for developers of the kind that write good code from the beginning. So with that out of the way, the whole point of real world development methodology is to make sure that the feature-bloated code hastily written by average developers doesn't lack a specification, doesn't deteriorate over time into something that has to be abandoned, and doesn't have such a high risk of modification/refactoring that it can't be maintained.
Now there are some good points to 1.4 "code is better than tests" actually applies: make asserts in the code rather than in the tests, where possible. Fully agreee with that. Or even better I'd say: don't "assert" things, just make the invalid state impossible. This is what types are for. If you can't take a null into a method, don't even accept an object that could be null by accident. Accept an Option type. Yes even in Java. Even in bloody C. Anything that a compiler could have caught shouldn't be either a runtime assert nor a test.
> if bad programmers write bad code that needs unit tests, they're also going to write bad unit tests that don't test the code correctly. So what's the point?
Well firstly they are usually better than no tests. And second, they tend to fail too often, rather than too little. This is what causes the enormous cost incurred by these tests. But apart from that 90% unnecessary cost (test failures are failures of the code rather than the costs, so after changing the code you have to change the test) - they still do add a lot of confidence for refactoring. Because in my experience with a large but bad test suite, there are often false positives but false negatives are much rarer.
I don't think that conclusion is any more valid than the proposition that if we just skipped the developer alltogether and generated both the business code AND the test, we'd have something of value.
> Coplien's claim that higher-level tests are better stands, so this article cannot be dismissed as "beating a dead horse."
i think higher level tests are good, but I don't think that makes lower level tests bad. I think there should usually be a health mix, with systems tested on several levels.
That would have some value, as evidenced by all the useful spreadsheets produced by end users, but the the value of the conjunction ("the business code AND the test") lies entirely in the 'business code' part, which is why we do not automatically generate unit tests for spreadsheet macros, or any other code for that matter.
>I think higher level tests are good, but I don't think that makes lower level tests bad.
That's not my point; my point is that unit tests are not automatically useful. Once we accept that, Coplien's argument that there are better ways to spend our limited resources cannot be dismissed (and resources are always limited, because of the combinatorial explosion, which is most evident at the unit level).
Actually, this is a surprisingly useful strategy! There are two common versions of this idea:
1. Invariant testing with random data. This usually goes by the name of "QuickCheck", and it allows writing tests that say things like, "If we reverse a random list twice, then it should equal the original list." In practice, this catches tons of corner case bugs.
2. Fuzz testing. In this case, the invariant is, "No matter how broken the input, my program should never crash (or corrupt the heap, or access uninitialized memory, or whatever)." Then you generate a billion random test cases and see what happens. This usually finds tons of bugs, and it's a staple of modern security testing.
So yeah, even random test cases are very valuable. :-)
Virtually nobody just makes a change and ships it users without some kind of testing. Maybe you run your program by hand and mess around with the UI. Maybe you call the modified function from a listener with a few arguments. Maybe you have sadistic paid testers who spend 8 weeks hammering on each version before you print DVDs. (And those testers will inevitably write a "test plan" containing things to try with each new release.)
The goal of automated testing is to take all those things you were going to do anyway (calling functions, messing with the UI, following a test plan), and automate them. Different kinds of testing result in different kinds of tests:
1. Manually calling a function can be replaced by unit tests, or even better, doc tests. Everybody loves accurate documentation showing how to use an API!
2. Manually messing around with the UI can be replaced by integration tests.
3. Certain things are much faster to test thoroughly in isolation, so you may mock out parts of your system while testing other parts in great detail, then tie everything together with a few system-wide tests.
People love to invent elaborate taxonomies of testing (unit, request, system, integration, property, etc.). But really, they're all just variations on the same idea: You'd never ship code without testing it somehow, and it's boring and error-prone to do all that testing by hand every time. So automate it!
What I am doing here is using the idea of automatic unit test generation to show that Coplien's argument cannot be dismissed with the claim that all unit testing adds value.
It takes some proper skill, vision, and maturity to ship version 3.0. And you probably won’t get a big payday if you can’t.
Without getting into the greater unit testing debate... the point would be having a testable code base.
Those who test-first guarantee that tests exist and that code can be exercised in a test-harness. Those who do not generally have code that cannot be exercised in a test-harness. Granting bad developers who write bad code and bad tests: they're at the least writing tested bad code. That lowers the burden to fixing it, and ensures that any given issue will be solved in isolation instead of requiring a major, risky, rewrite to get the system to the point of further maintenance.
Having that test apparatus puts your maintenance developers on a much nicer foot and provides the blueprints for migrating, abstracting, or replacing most any system component, along with meaningful quality baselines. It's a mitigation technique for the high-level big-bang rinse-and-repeat system cycles that pop up in the Enterprise space, but no silver bullet.
Bad tests can be deleted or improved with impunity, but a bad logical system core made with no eye towards verification is expensive and risky to even touch. For a small website that means nothing, for an complex legacy monster system about to be rewritten for the 4th time it can make all the difference in the world.
1) If you decide to change your implementation (refactor low-level code to make it more performant/re-usable/maintainable) then you won't have to change your tests.
2) It will simplify your implementation code.
The second point is as important as the first: In general your code should have the minimum number of abstractions needed to satisfy your requirements (including requirements for re-use/flexibility etc). Testable code often introduces new abstractions or indirections just for the sake of making it possible to plug in a test. One or two of these are OK. But when you add up all the plug points for low level tests in a system it can make the code a lot more complex and indirect than it needs to be. This actually makes it harder to change which was not the idea behind making it testable in the first place.
Gui code may be an exception here. Because gui's are often slow and unstable to test it is worth cleanly separating the gui from the model just to make the system more testable, even if you never intend to use the model with any other kind of ui (the usual motivator for separating model and gui). Any model-view system will help here. If you have a system where the gui is more or less a pure function of the model then you can write all your scenario/stateful tests at the model level and just write a few tests for the ui to verify that it reflects the model accurately
Modifying complicated tests can be unbounded, and I think a lot of the pushback comes from these situation.
It is very possible to create tests that are worse than doing nothing. It can be a roadmap for learned helplessness, too.
Perhaps I misunderstand you, but I have to disagree with this.
Testable code is absolutely a good thing in and of itself.
When something breaks, being able to isolate each piece of the codebase and figure out why it's breaking is essential.
Far too often I've got bug reports from other developers saying "System X is broken, what changed?". When I ask for more detail, a test case proving it's broken, etc - they can't come up with that. Their test case is "our entire service is broken and I think the last thing that happened was related to X". Sometimes it is a problem with System X, but more frequently it's something else in their application.
The smaller you can make your test case, the better, and the easier it is to get a grasp of the acutal problem.
Let's assume you hire good programmers, because otherwise you're doomed. But oftentimes, the "good" programmer and the "bad" programmer are the same person, six months apart:
1. John writes some good code, with good integration tests and good unit tests. He understands the code base. When he deploys his code, he finds a couple of bugs and adds regression tests.
2. Six months later, John needs to work on his code again to replace a low-level module. He's forgotten a lot of details. He makes some changes, but he's forgotten some corner cases. The tests fail, showing him what needs to be fixed.
3. A year later, John is busy on another project, and Jane needs to take over John's code and make significant changes. Jane's an awesome developer, but she just got dropped into 20,000 lines of unfamiliar code. The tests will help ensure she doesn't break too much.
Also, unit tests (and specifically TDD) can offer two additional advantages:
1. They encourage you to design your APIs before implementing them, making APIs a bit more pleasant and easier to use in isolation.
2. The "red-green-refactor-repeat" loop is almost like the "reward loop" in a video game. By offering small goals and frequent victories, it makes it easier to keep productivity high for hours at a time.
Sometimes you can get away without tests: smaller projects, smaller teams, statically-typed languages, and minimal maintenance can all help. But when things involve multiple good developers working for years, tests can really help.
> Jane's an awesome developer, but she just got dropped into 20,000 lines of unfamiliar code. The tests will help ensure she doesn't break too much.
In my experience, a well tested codebase lowers the risk and decreases the time to productivity of new hires.
I also have no problems designing APIs, experience is what you need and the experience in asking the right questions. No amount of TDD will solve you getting half-way through a design and then finding you needed many-many because you misunderstood requirements.
Other than that API design is (fairly) trivial. I spend a tiny amount of my overall programming time on it. I will map out entire classes and components for large chunks of functionality without filling in any of the code in less than an hour or two and without really thinking about it. Then the hard part of writing the code begins which might take a month or two, and those data structures and API designs will barely change. I see my colleagues doing similar check-ins, un-fleshed components with the broad strokes mapped out.
I'd say it's similar to how we never talk about SQL normal forms any more. 15 years ago the boards would discuss it at length. Today no-one seems to. Why? Because we understand it and all the junior programmers just copy what we do (without realizing it took us a decade to get it right). API design was hard, but today it's not, we know what works and everyone uses it. We've learnt from Java Date and the vagaries in .net 1.0 or PHP's crazy param ordering and all the other difficult API designs from the past and now just copy what everyone else does.
I've gone back through most of the PDF, and I can't figure out what section you're referring to. There's a bunch of discussion of perverse incentives (mostly involved incompetent managers or painfully sloppy developers, who will fail with any methodology), but I don't see where the author addresses what I'm talking about.
Specifically, I'm talking about the hour-to-hour process of implementing complex features, and how tests can "lead" you through the implementation process. It's possible to ride that red-green-refactor loop for hours, deep in the zone. If a colleague interrupts me, no problem, I have a red test case waiting on my monitor as soon as I look back.
This "loop" is hard to teach, and it requires both skill and judgment. I've had mixed results teaching it to junior developers—some of them suddenly become massively productive, and others get lost writing reams of lousy tests. It certainly won't turn a terrible programmer into a good one.
If anything, the biggest drawback of this process is that it can suck me in for endless productive hours, keeping me going long after I should have taken a break. It's too much of a good thing. Sure, I've written some of the most beautiful and best-designed code in my life inside that loop. But I've also let it push me deep into "brain fry".
> No amount of TDD will solve you getting half-way through a design and then finding you needed many-many because you misunderstood requirements.
I've occasionally been lucky enough to work on projects where all the requirements could be known in advance. These tend to be either very short consulting projects, or things like "implement a compiler for language X."
But usually I work with startups and smaller companies. There's a constant process of discovery—nobody knows the right answer up front, because we have to invent it, in cooperation with paying customers. The idea that I can go ask a bunch of people what I need to build in 6 months, then spend weeks designing it, and months implementing it, is totally alien. We'd be out of business if it took us that long to try an idea for our customers.
It's possible to build good software under these conditions, with fairly clean code. But it takes skilled programmers with taste, competent management and a good process.
Testing (unit, integration, etc.) is a potentially useful part of that process, allowing you to adapt to changing requirements while minimizing regressions.
Were the failing tests due to John/Jane's correctly coded changes, regressions, or bad code changes? The tests provide no meaningful insight into that - it's still ultimately up to the programmer to make that value judgement based on the understanding of what the code is supposed to do.
What happens is that John and Jane make a change, find the failing unit tests, and they deem the test failures as reasonable given the changes they were asked to make. They then change the tests to make them pass again. Again, the unit tests are providing no actual indication that their changes were the correct changes to make.
WRT the advantages:
1. "design your APIs before implementing them" - this only works out if we know our requirements ahead of time. Given our acknowledgement that these requirements are usually absent via Agile methodologies, this benefit typically vanishes with the first requirement change.
2. "makes it easier to keep productivity high for hours at a time" This tells me that we're rewarding the wrong thing: the creation of passing tests, not the creation of correct code. Those dopamine hits are pretty potent, agreed, but not useful.
No, the failing tests indicate something changed, and where it changed - what external behavior of a function or class changed. Is the change the right thing, or a bug? Don't know, but it tells you where to look. That's miles better than "I hope this change doesn't break anything".
> "design your APIs before implementing them" - this only works out if we know our requirements ahead of time.
No. "API" here includes things as small as the public interface to a class, even if the class is never used by anything other than other classes in the module. You have some idea of the requirements for the class at the time when you're writing the class; otherwise, you have no idea what to write! But writing the test makes you think like a user of that class, not like the author of that class. That gives you a chance to see places where the public interface is awkward - places that you wouldn't see as the class author.
> This tells me that we're rewarding the wrong thing: the creation of passing tests, not the creation of correct code.
We're rewarding the creation of provably working code. I fail to see how that's "the wrong thing".
If we were discussing integration tests, where the interactions between different methods and modules are validated against the input - I'd agree.
But this is about unit tests, in which case the "where" is limited to the method you're modifying, since it's most likely mocked out in other methods and modules to avoid tight coupling.
> includes things as small as the public interface to a class
If we're talking about internal APIs as well, we can't forget that canonical unit testing and TDD frequently requires monkey patching and dependency injection, which can make for some really nasty internal APIs.
> I fail to see how [provably working code is] "the wrong thing".
So, thinking back a bit, I can recall someone showing me TDD, and they gave me the classical example of "how to TDD": Write your first test - that you can call the function. Now test that it returns a number (we were using Python). Now test that it accepts two numbers in. Now test that the output is the sum of the inputs. Congratulations, you're done!
Except you're not, not really. What happens when maxint is one of your parameters? minint? 0? 1? -1? maxuint? A float? A float near to, but not quite 0? infinity? -infinity?
Provably working code is meaningless. Code that can be proven to meet the functionality required of it - that's what you really want. But that's hard to encapsulate in a slogan (like "Red, Green, Refactor"), and harder use as a source of quick dopamine hits.
Sure, but the consequences may not be. Here's a class where you have two public functions, A and B. Both call an internal function, C. You're trying to change the behavior of A because of a bug, or a requirement change, or whatever. In the process, you have to change C. The unit tests show that B is now broken. Sure, it's the change to C that is at fault, but without the tests, it's easy to not think of checking B. That's the "where" that the tests point you to.
> Provably working code is meaningless. Code that can be proven to meet the functionality required of it - that's what you really want.
Um, yeah, of course that's what you really want. Maybe you should write your tests for that, and not for stuff you don't want? If you don't care about maxint (and you're pretty sure you're never going to), don't write a test for maxint. But it might be worth taking a minute to think about whether you actually do need to handle it (and therefore test for it).
I feel like you're perhaps defining unit tests to exclude "useful unit tests" by relabeling them. Yes, if you exclude all the useful tests from unit testing, unit testing is useless. Does a unit test suddenly become a regression test if the method it tests contains, say, an unmocked sprintf - which has subtle variations in behavior depending on which standard library you link against? No true ~~scotsman~~ unit test would have conditions that would only be likely to fail in the case of an actual bug?
> But this is about unit tests, in which case the "where" is limited to the method you're modifying
That's still useful. C++ build cycles mean I could have easily touched 20 methods before I have an executable that lets me run my unit tests, telling me which 3 of the 20 I fucked up in is useful. Renaming a single member variable could easily affect that many methods, and refactoring tools are often not perfectly reliable.
Speaking of refactoring, I'm doing pure refactoring decently often - I might try to simplify a method for readability before I make behavior changes to it. Any changes to behavior in this context are unintentional - and if they occur, 9 times out of 10 discover I have a legitimate bug. Even pretty terrible unit tests, written by a monkey just looking to increase code coverage and score that dopamine hit can help here - to say nothing of unit tests that rise to the level of being mediocre.
Further, "where" is not limited to "the method you're modifying". I'm often doing multiplatform work - understanding that "where" is in method M... on platform X in build configuration Y is extremely useful. Even garbage unit tests with no sense of "correctness" to them are again useful here - they least tell me I've got an (almost certainly undesirable) inconsistency in my platform abstractions, likely to lead to platform specific bugs (because reliant code is often initially tested on only one of those platforms for expediency). This lets me eliminate those inconsistencies at my earliest convenience. When writing the code to abstract away platform specific details, these inconsistencies are quite common.
This is probably the root of most of our disagreement. I belong to the school of thought that says (in most cases) "Mocking is a Code Smell": https://medium.com/javascript-scene/mocking-is-a-code-smell-... And dependency injection (especially at the unit test level) is giant warning sign that you need to reconsider your entire architecture.
Nearly all unit tests should look like one of:
- "I call this function with these arguments, and I get this result."
- "I construct a nice a little object in isolation, mess with it briefly, and here's what I expect to happen."
At least 50% of the junior developers I've mentored can learn to do this tastefully and productively.
But if you need to install 15 monkey patches and fire up a monster dependency injection framework, something has gone very wrong.
But this school of thought also implies that most unit tests have something in common with "integration" tests—they test a function or a class from the "outside," but that function or class may (as an implementation detail) call other functions or classes. As long as it's not part of the public API, it doesn't need to be mocked. And anything which does need to be mocked should be kept away from the core code, which should be relatively "pure" in a functional sense.
This is more of an old-school approach. I learned my TDD back in the days of Kent Beck and "eXtreme Programming Explained", not from some agile consultant.
Even if all you know is "something changed", that's valuable. Pre-existing unit tests can give you confidence that you understand the change you made. You may find an unexpected failure that alerts you to an interaction you didn't consider. Or maybe you'll see a pass where you expected failure, and have a mystery to solve. Or you may see that the tests are failing exactly how you expected they would. At least you have more information than "well, it compiled".
And if "the unit tests are providing no actual indication that their changes were the correct changes to make", even the staunchest proponents of TDD would probably advise you to delete those tests. If there's 1 thing most proponents and opponents of unit testing can agree on, it's probably that low-value tests are worse than none at all.
I have no idea if this describes you, but I've noticed a really unfortunate trend where experienced engineers decide to try out TDD, write low-value tests that become a drag on the project, and assume that their output is representative of the practice in general before they've put the time in to make it over the learning curve. People like to assume that skilled devs just inherently know how to write good unit tests, but testing is itself a skill that must be specifically cultivated.
Err...just do a quick search in the comments and you will see that this is hardly the consensus.
There probably are good resources about this somewhere, but compared to the impression of "TDD promotes oceans of tiny, low-value tests" they lack visibility.
Unfortunately, the existing architecture doesn't make dependencies obvious. So simply knowing that "something has changed" is very, very helpful.
I don't think one necessarily follows the other. Consider the examples of test pseudo-code:
EXPECT[image_to_text("photo1.jpg") == "cat"]
EXPECT[image_to_text("photo2.jpg") == "dog"]
EXPECT[image_to_text("photo3.jpg") == "apple"]
EXPECT[image_to_text("photo4.jpg") == "Barack Obama"]
A novice bootcamp graduate could write that test harness code.
However, it requires an experienced programmer with a PhD in Machine Learning to write the actual neural net code for image_to_text(). (Relevant XKCD[1])(Yes, there's also variability of skill in writing test code. An experienced programmer will think of extra edge cases that a novice will not and therefore, write more comprehensive tests.)
That still doesn't change the fact that writing code that tests to check for correct results is often easier than writing the algorithm itself.
It also contradicts the notion that refactoring code requires that the test code be refactored. That's not always true. Since the jpg file format came out in 1992, that test code could have been written in that same year and it would still be valid today, 25 years later. We still want to ensure the image_to_text() function will return "cat" for "photo1.jpg" even if we refactor the low-level code to use the latest A.I. algorithms.
Another example of a regression test to test performance:
TIMING[EXPECT[image_to_text("photo1.jpg") == "cat"] < 500 milleseconds]
Again, it's easy for a novice programmer to write a "performance timing specification executes in less than 500ms". It's much harder to refactor the code from slow code O(n^2) to faster code O(n log n).To do a meta-analysis of the debate with Coplien: a big reason we're arguing his conclusion is that he complains about this abstract thing he called "unit tests". If he had copy-pasted concrete examples of the bad unit testing code into his pdf, maybe more of us would agree with him? In other words, I'd like to see and judge for myself if Coplien's "unit tests" look like the regression tests that SQLite and NASA use to improve quality -- or -- if they are frivolous tests that truly wastes developers' time. Since he didn't do that, we're all in the dark as to the "test code waste" he's actually complaining about.
It should be short, plain, simple and make it obvious what is being tested, how and why.
Because when that test fails you want to quickly understand what the problem is.
“Smart” (or just too big) tests all detract from this.
On 2 separate occasions I had an app with extensive unit tests that seemed to work fine but also seemed to have strange rare bugs.
Both times I wrote a single additional "unit" test that fired up the environment (with mocks, same as the other unit tests) but then acted like a consumer of the API and spammed the environment with random (but not nonsense) calls for several minutes. These tests were quite complex so basically the exact opposite of what you're suggesting.
Not only did I immediately find the bug, but in both cases I found like 10 bugs before I got the test to even pass for the first time.
At the same time all the small unit tests were happily passing. Because they didn't hit edge cases (both in data and in timing) that nobody had thought of.
In fact that’s usually one of my default tests to ensure that I have a well-defined behavior even in these cases.
http://whatdidilearn.info/2017/10/23/writing-documentation-i...
Maybe, but there's no reason to write such tests by hand. Just automatically test each method against the previous version and see what changed.
I think the generated results would be too noisy and developers would stop paying attention to test failures, but maybe a strict enough code review process could prevent that.
”The code is the spec” is just a way of saying “we don't know what the code is supposed to do, and the intent was never, even initially, properly validated by domain experts, or, if it was, we tossed the knowledge generated in that process.”
This is especially the case if that phrase is used about a large system in a complex domain.
And even with all the (very real) downsides of it, it still a huge plus.
Besides, a lot of persons tend to see tests in a very rigid way. You don't have to do tdd. You don't have to do have full coverage. You don't even need all your tests to be unit tests, you can mix end2end and functional tests with it. You can have hacky mocks in it too. It doesn't need to be perfect for your team to benefit from it.
Actually, I would say that you probably want it not to be perfect so that the benefits outweight the costs. Cause they, tests have a huge cost.
But I would argue that represents in my experience just 10% of the unit tests I see. Instead I see a lot more testing of implementation details.
You'd think, if this was the desired model, you could partially automate the "writing unit tests" part of this process. (Integration tests no, but unit tests yes.) The "spec" for the unit tests is already there, in the form of the worktree of the previous, known-good commit.
That means that, in a dynamic language, you'd just need a "test suite" consisting of a series of example calls to functions. No outputs specified, no assertions—just some valid input parameters to let the test-harness call the functions. (In a static language, you wouldn't even need that; the test-harness could act like a fuzzer, generating inputs from each function's domain automatically.)
The tooling would then just compare the outputs of the functions in the known-good-build, to the outputs of the same functions from your worktree. Anywhere they differ is an "assertion failure." You'd have to either fix the code, or add a pragma above the function to specify that the API has changed. (Though, hopefully, such pragmas would be onerous enough to get people to mostly add new API surface for altered functionality, rather than in-place modifying existing guarantees.) A pre-commit hook would then strip the pragmas from the finalized commit. (They would be invalid as of the next commit, after all.)
Interestingly, given such pragmas, the pre-commit hook could also automatically derive a semver tag for the new commit. No pragmas? Patch. Pragmas on functions? Minor version. Pragmas on entire modules? Major version.
Ah, but the spec is often really poorly defined. That is all sorts of edge cases and uncommon code paths do not work as imagined. And, of course, it is full of bugs.
I have found countless "bugs" and issues when writing tests for existing code. I also often have to refactor the code to make it clearer what it's doing or to make it testable.
Unit tests are a development tool more than they are a testing tool. They are a means to produce correct, well documented, well speced, cleanly separated code.
I believe I came across this document (or something close to it) a few years ago, and it changed my life. I was able to get great velocity and low bug-per-feature count (1 or 2), using systems/integration testing combined with exploratory testing by a QA team. I was even able to find subtle bugs in the unit-tested back-end via integration tests from a mobile app. At the end of the day, if your testing isn't serving the business logic, it's wasted effort.
The most bad-ass thing I remembered was the part where if you couldn't elicit a code path using a systems test, it's a good candidate for deletion.
These days, a large majority of the frontend tests in my codebase aren't tests that I've written myself. They're automatically generated Jest snapshot tests that capture the virtual DOM output of entire components in different states as dictated by product requirements, using the StoryShots plugin for React Storybook.
These tests prove to be extremely effective in catching bugs, because they directly reflect product requirements, and because they end up exercising a large majority of the codebase for very little cost compared to covering the same amount of code with individual unit tests.
The latter point is an important consideration too, because a lot of the code these integration-level tests exercise are code that I'd have considered too trivial to have been worth the ongoing maintenance costs in building unit tests for, when considered in isolation.
This means it'd have been easy for me to be satisfied with selectively and arbitrarily deciding which pieces of code should qualify for test coverage, and leave test coverage at some arbitrary number, as opposed to what I do now, which is to always strictly require 100% code coverage, but explicitly and deliberately exclude code that I have good reason to not test, which feels like a much more solid framework for ensuring code quality.
I still write unit tests for functionality that component snapshot tests can't reasonably cover, but usually only at a level of granularity where the tests themselves can map cleanly to product requirements as well, instead of painstakingly testing every little function in perfect isolation.
I dealt with some of this, including a rather bullying type who tried to browbeat people into writing tests first. And while that may be a bad example, I'm afraid it really wasn't unrelated to the movement itself.
When people, a while ago, said "TDD is dead", they didn't actually argue that the technique wasn't useful. They were saying that the whole "you're wrong, I'm right, if you want to stay employed you'll do as I say and practice TDD", that is no longer defensible.
It's probably time to leave that in the past, and hope that people who promote a methodology this way have learned the mistakes of aggressively cramming things down everyone's throat. Truth is, I was actually a frequent practitioner of TDD, and I was pretty appalled with how it was getting pitched to the programming community.
Now, let's leave that in the past and take a fresh look on whether TDD is a very beneficial practice in many contexts. I certainly agree it isn't crap.
On a side note, unfortunately I think the software industry tends to be particularly bad about this pendulum swinging back and forth between extremes thing. There is so much emphasis on innovation that people are biased toward making radical changes and believing they are going to revolutionize everything. And the culture is such that the more you go out on a limb and push something radical, the more you are respected, because we often value guts more than good judgement.
> Hopefully we'll arrive at a reasonable middle ground at some point, and unit tests will be valued (and prioritized) neither too much nor too little.
My coding style has very much gone through a similar evolution. When I first learned about TDD, I went in whole hog - test everything, ui tests, request tests, unit tests - tests for everything.
Then I started to notice that the development costs associated with extreme coverage did not, in fact, pay off.
The number of bugs found was vastly outnumbered by the time wasted dealing with the peculiarities of various ui testing frameworks, let alone the amount of time wasted waiting for those tests to run.
My metric is now that I've written enough tests to feel confident that the code works. It's an extraordinarily qualitative metric, and I cannot find a way to objectively quantify it, yet it very much works for me.
Really, it all comes down to visibility. Is there somewhere that is up to date, that will tell you what the code is supposed to do? Do you have tooling in production that will alert you when it's not doing what it's supposed to do? Do you feel confident that the code works?
I agree in principle, but IME:
* This kind of algorithmic code often makes up a fairly small proportion (~5%) of code written for most business driven applications - it's usually mostly coordination. Some domains may be different (and for them, unit testing will appear to be much more effective), but I think this is the norm.
* Integration tests can test algorithmic acceptably well, but unit tests do not test integration code in an acceptable fashion.
* Where you have algorithmic and integration code smushed together (this type of technical debt is, sadly, the norm 'in the wild'), again, integration tests work acceptably well whereas unit tests require a mess of mocks.
* Integration tests do not have to be high level or slow.
"Refactoring breaks tests. Sometimes when you refactor code, you break tests. But my experience is that this is not a big problem. For example, a method signature changes, so you have to go through and add an extra parameter in all tests where it is called. This can often be done very quickly, and it doesn’t happen very often. This sounds like a big problem in theory, but in practice it isn’t."
This is a big problem in practice when you are working on large scale code bases in which the developers have not been super strict about decoupling everything.
True, but you can have both types of tests, which lets you test algorithms (unit tests) and higher level combinations (integration tests) both in the best way. The right tool for the right job.
> Where you have algorithmic and integration code smushed together (this type of technical debt is, sadly, the norm 'in the wild'), again, integration tests work acceptably well whereas unit tests require a mess of mocks.
In many cases, you'd be better off separating the logic, which would let you have both types of tests, each for a different part of the logic.
> Integration tests do not have to be high level or slow.
No, but they are a different tool.
You can, but if your code is 95% coordination and 5% algorithmic, at most you'd want 5% unit tests.
>In many cases, you'd be better off separating the logic
Which is refactoring and if you're refactoring you need to have your code surrounded with tests to do it safely.
And, once you've done that, if you've already written integration tests for that logic and they perform acceptably well, there may be no point in rewriting the integration test as a unit test.
The combinatorial complexity argument does apply more (much more, in many cases) to unit testing than to testing at higher level of abstractions, broadly for the same reasons that abstraction ameliorates the design-complexity problem. Consider a program using heuristic methods to find a near-optimal solution to an NP problem: not only would testing that it produces valid solutions at the program level be a lot simpler than unit-testing all its components, no amount of the latter would establish that any of its solutions are valid. The point is not that integration testing allows you to cover the whole state-space, but that it is a more effective way to select what you test.
The point of assertions is that they are tested in the execution of integration tests.
It has been my experience that most of the effort expended in the debug-fix cycle is a result of errors in decomposing the problem (or, equivalently, composing the units) - for example, an incorrect assumption about the data a unit uses, rather than the handling of the data in accordance with the assumptions made about it. Discrepancies between what is expected of a component and what it does also seem to be common, such as when one component is predicated on a constraint being respected by another, but the latter does not always do so.
https://news.ycombinator.com/item?id=7353767 (268 comments)
https://news.ycombinator.com/item?id=11799272 (280 comments)
https://news.ycombinator.com/item?id=13815779 ("only" 24 comments)
So maybe this document should be the focus of the discussion now.
However, the title "Seque" is completely unspecific and hence totally useless. Given that title, I don't see how to submit this to HN in a way that it attracts readers.
This is really a pity.
The bigger problem, I think, is the velocity of articles appearing and disappearing from the front page makes follow-up articles have an inherently smaller starting audience.
Mods appear really clear about this. They want you to use the original title unless it's clickbait, or inflammatory, or misleading. Only then can you change it, and they say an informative sentence from the article might do.
This means that articles with terrible, unclear, titles get posted to HN all the time.
See eg this thread: https://news.ycombinator.com/item?id=15540388
(Of course, I'm not a mod, so maybe I'm wrong.)
Coming up with a title was indeed a bit difficult.
Sometimes I find that unit testing is an attempt to spend as much time as possible experiencing the first reality (you are a smart, competent programmer) as a relief from the pain of the second reality (you have a gnawing fear that your understanding of the context is fatally flawed.) Can't figure out if the work you're doing is constructive? Anxious about the lack of requirements? Go soothe yourself writing unit tests. Kick back and watch the integration tests run. You're doing your job, and the rest will take care of itself, right?
Just be prepared for a lot of shocked reactions when you say 'Most Unit Testing Is Waste' to people who have not been softened a bit to this concept.
I really like this way of thinking. The next chapter of getting into this subject are the live discussions (5 videos) between Kent Beck, David Heinemeier Hansson and Martin Fowler talking about TDD (which relies a lot on unit tests): https://martinfowler.com/articles/is-tdd-dead/ I enjoyed these a lot too and the combination of these sources has really improved my testing.
Unit tests are unlikely to test more than one trillionth of the functionality of any given method in a reasonable testing cycle. ... Trillion is not used rhetorically here, but is based on the different possible states given that the average object size is four words, and the conservative estimate that you are using 16-bit words)
An int may contain "four billion states", but for the requirements, its highly likely that we can classify the integer into three states: less than zero, zero, greater than zero. As a bank, I might not care how much money you have, only that you have more than zero. In a transaction, I don't care how much money changes hands, as long as no money is lost.
Pointing at memory-as-bits, as if we're still using punch cards, and then hand waving "I can't possibly test this", ignores sixty years of progress. The refusal to imagine a class with range checking is a damning statement about the author's own ability as an engineer.
Programmers have a tacit belief that they can think more clearly (or guess better) when writing tests [than] when writing code, or that somehow there is more information in a test than in code.
Consider writing a sorting algorithm vs testing a sorting algorithm. Would you feel more confident writing the test for a sorting algorithm than writing the algorithm itself? The test is simple: is every item in the list less than the next item? The code is far more complex. We're in the same realm as NP problems. I can write a test to verify that a graph is correctly 3 colored, but the code might be a bit harder.
Perhaps, then, the author's experience of other developers believing they can "think more clearly" is actually his observation that the developers are solving simpler problems, and are thus more confident. And that is the point of tests: it is easier to verify than solve.
In short, every conclusion in this article begs the question, "Might there be another explanation?"
Being able to inspect the state of the program in real time is invaluable and gives me a lot of confidence that I understand how the code works. For some reason most programmers I see don't even run their code locally and just use log statements to guess at state when it invariably doesn't work as expected.
Another big problem is test data. I see way too much naive mocking. You really need to exercise your code with data that is as close to real-life input as possible, ideally it's a sanitized version of production data. Other tests are great too (i.e large lists of 'naughty strings') but if you're manually specifying your test data you are a) spending a lot of time doing something that should be automated and b) are only exercising your code with what you think it might see which is usually not good enough.
Working in an GUI debugger / REPL is great for visualizing code I often code that way myself, but let's not fool ourselves, it still requires manually setting breakpoints, and pressing keys to step through code. You can't do this for every method after every change, whereas unit testing has that advantage. I do agree with most things mentioned in the article though. A lot people end up writing tests that hit databases and are too slow or fail randomly due to chained state, or are over specified & end up just getting in the way. What it comes down to is you can have good tests or bad tests, and its still entirely subjective just like whether the code itself is good or bad. I recommend the book xUnit test patterns, its basically a bunch of "rules of thumb"
Despite the inflammatory title, the point the PDF is making I believe is that unit tests should be kept short and simple, and convey the intent of your code.
Intent is akin to the 'why' - written code explains the 'how' and 'what' extremely well, but without the 'why' it loses all meaning. Writing tests to convey intent is essential because no computer or programming language can do this for us currently, so it's left up to us.
What the Article describes are tests but I would not call them unit tests.
Unit tests are great when you have a large code base with dozens of apps and need to modify a core library to remove side effects. How do you know you didn't break one of the apps?
But if we go deeper, why does the core library have side effects? Because it was a crummy, poorly designed piece of crap to begin with. If it was originally written with high quality it wouldn't need the refactor now. The author noticed this. He wrote good code to begin with so it didn't need a lot of testing.
When you have a large team of mediocre engineers you need unit tests to guard against more bad code from getting in. They might even be brilliant engineers, stuck in a horrible process of churning out features to unrealistic deadlines.
If you have a great team, a great process, and a great budget with realistic goals, you can crank out amazingly good software without the need for a lot of unit tests. But that's not the real world. In the real world we have lots of unit tests. They are a band-aid over the other problems without addressing them specifically.
It's easy to disagree with someone who portrays an enemy which doesn't exist.
I disagree with this piece everywhere it manages to land somewhere concrete, where its claims can be verified and assessed.
> But if we go deeper, why does the core library have side effects
Let's not get ahead of ourselves in theoretical, non-applied functional programming.
As even Haskellers recognize, all computing would be meaningless if there ultimately was no side effects. In fact, we use computers for their "side-effects".
Sometimes you need side-effects. And sometimes you unit-tests are great way of automatically verifying that you have the right side-effects.
Complex running dialog with a hardware serial device, for example, could be rearchitected as a component that has a clear-text dialog with a serial device proxy. This would create functional, side-effect free, testable units of code while also maintaining rich side effects in production (handled by pure integration tests).
Essentially: if you _have_ to use an integration test you should almost always be refactoring into something that you handle with unit tests in conjunction with your integration tests. Otherwise you're leaving maintenance developers with too much risk when making changes and no clear separation of your domain model from your application code.
Absolutely, but at that point you're moving the side effect out to the boundary and separating it from the logic. Then you can unit-test the logic away from the side effect (but the test of the side effect itself still has to be an integration test).
class AccessCountedInteger {
private final int value;
private int accessCount = 0;
AccessCountedInteger(int value) {
this.value = value;
}
public int getValue() {
accessCount++;
return value;
}
public int getAccessCount() {
return accessCount;
}
}
This unit has side effects on the method getValue() that impact what is returned by getAccessCount(). While it is contrived, if you are making something following the builder pattern, you will in general have a lot of side effects from all of the methods called on the final build() method. This is quite amenable to unit testing.We use computers for their effects. "Side-effects" are things that happen by accident, or incidentally, when we were trying to do something else.
Maybe except of a one guy who I talked to a while ago. He does not write unit tests, because static type checking in C++ takes care of everything that unit tests do (according to him), but I actually have seen his code that contained few system tests, so there are exceptions.
I guess the main reason for my condescending sentiment is due that the author gives 0 examples of code that requires tests vs code that does not so it is actually not.
Why do you think they aren’t? I don’t write unit tests except when I’m specifically paid extra for that (yes, coding strongly-typed languages too), but I do write other kinds of tests as I see fit.
Btw, what type of software do you write and how do you know if it actually works?
I find unit tests are useless for most cases. But this doesn’t stop me from writing more complicated tests. Even unit tests when they’re useful. E.g. for sufficiently complex low-level SIMD math routines: inputs & outputs are simple, no IO, no multithreading, no large dependencies, no side effects, just some computations, and tons of weird _mm256_verb_typesuffix intrinsics everywhere i.e. it’s very easy to make mistakes writing these.
> what type of software do you write
Lately Windows CAD and Linux embedded. Previously mobile software, PC and console videogames, CNC & robotics, GIS, WinCE/embedded, multimedia/video codecs, utilities for Windows administration, lots of other stuff.
> how do you know if it actually works?
I use strongly typed languages that catch 99% of stupid errors at compile time. I test it manually while using debugger at the same time to inspect internal state. I design my software to be testable, e.g. logging, asserts, configurable debug dumps, well-defined interfaces between components so they can be tested in isolation, etc.
P.S. Today I’ve fixed 3 bugs in my code.
[Windows] A user tried using my software on old AMD CPU, it didn’t work ‘coz SSE 4.1 instruction set is not supported.
[Windows] Under some conditions, my code reads value from a depth+stencil D3D texture while the GPU is still rendering previously submitted commands into the same texture, read fails.
[Linux] When both HDMI and DSI displays are connected, my DRM/KMS client code selects the wrong one.
All three are caused by environment or external hardware, and you can’t unit test these.
This is, imo, the most important takeaway and it is, or should be, obvious, but it often isn't (maybe because "common sense is the least common of senses" or something like that). It's also what I strive for, with one caveat[1].
> One is to use it as a learning tool: to learn more about the program and how it works.
Another great takeaway. Good tests should work as documentation, imo. I consider this another test quality metric, even: If looking at the tests only confuses me further, those tests need to be changed (and the code they're testing too, most likely).
[1]: Striving for this can turn code coverage into an useful metric. Not of correctness, of course, but on how much code we're writing that doesn't solve a business requirement. That code should be refactored away to separate libraries or replaced with third party libraries that already do that. I'm firmly in the "avoid NIH" camp.
That's not why you write unit tests. They are, in a roundabout way, programming's way of double-entry bookkeeping. Tests and code both are there to double check the other one. Neither one is "smarter" than the other.
Integration tests, however, should not have to change, since they are typically run against the external interface to the software, which shouldn't change if you're just doing a refactoring.
They absolutely do, but at higher levels of abstraction, in terms of interfaces, interactions, constraints, obligations and responsibilities. In fact, unit and integration tests make you think about design in essentially the same way.
Having good integration tests means that you're free to completely re-organize your code, and you can still test that it works as expected.
Given that requirements change over time, even a well thought out code structure will likely need to be refactored at some point, and being able to have a test suite that doesn't have to be re-written during a refactoring is extremely valuable.
I've also seen code bases become unnecessarily complex because so much thought was put into how to unit test it, that not enough thought was put into writing clear, easy to understand code.
But, I mostly code C++ and C#, both are strongly typed, so the compiler helps a lot with refactors.
Most proponents of unit tests use horrible programming languages; they are afraid to change code because anything could break at any time. Stop using those languages and most of the problems 'fixed' by unit testing just disappear.
A lot of us would consider SQLite to be high quality and relatively bug-free. The extensive test suite they've built up that exercises each release is a huge reason for it.
Of test code quantity, Coplien writes:
>If your coders have more lines of unit tests than of code, it probably means one of several things. They may be paranoid about correctness; paranoia drives out the clear thinking and innovation that bode for high quality.
> - Keep regression tests around for up to a year
> - Throw away tests that haven’t failed in a year.
SQLite appears to keep their tests. Somebody files a bug; SQLite writes a test that reproduces that bug; the test code remains long after the bug is fixed. This prevents the bug from reappearing. (Isn't that the main purpose of "regression" in the phrase "regression test"?)
I also think it's better to keep relevant regression tests for years. E.g. Consider that there's piece of code that's been working correctly for 5 years with 5 years of passing regression tests . Imagine a new programmer wants to rewrite the code to optimize it for speed and reduced memory usage. I think we'd feel much more confidient if the new code passes those same regression tests that were accumulated over 5 years.
As for test code size ratio... A lot of good comprehensive tests will have LOC outnumbering the actual code being tested. This is especially true for library code that's used in many places up the stack. I wrote string parsing routines and a reverse Boyer-Moore search routine where the test code (test edge cases, test nulls, test string sizes at 2^32 boundaries, etc) was 10 times larger than the actual code.
Of testing's utility, Coplien writes:
- Testing can’t replace good development
- [...] Tests don’t improve quality: developers do
... which looks like a strawman and a false dichotomy. Can anyone cite a credible development philosophy that believes testing can replace bad developers or bad process?We could say that about <ANYTECHNOLOGY> such that <ANYTECHNOLOGY> can't replace quality developers. Garbage Collection doesn't improve quality, developers do. Array boundary checking doesn't improve quality, developers do. And so on.
Or maybe there's a difference in terminology? I wonder if Coplien considers SQLite testing "system test" or a "unit test"? Does he consider SQLite "white box testing" or "black box testing"?
None. But that's hardly the issue, I too often have seen people championing heavy testing as a way to deliver quality software, because it's the agile way. The problem is not with the philosophies but the way people interpret them or misinterpret them.
The sarcastic quote in the article sums my feeling on the behaviour I've seen in a lot of shops: "I find that weeks of coding and testing can save me hours of planning.".
I've had the hardest time convincing people to spend a week thinking about our approach and solution, but adding months of development time to expend test coverage and suddenly everyone thinks it's worth it (And usually with nothing to back that up).
Scrapping unit tests is scrapping time.
The correct solution to this is to not throw passing tests out, but to stop doing white-box testing.
(Not to mention that a test that may pass in your CI environment may fail - frequently - in your local workspace.)
1) examine the entire code base to find any similar problems 2) implement full parameter validation in that code to try to ensure it can’t happen again and 3) add exception logging code so that we will get an immediate error reported directly on the line it occurred if it does.
If I have any extra time left over, I refactor to eliminate code duplication, update/verify/add more parameter and input data validation, and increase code encapsulation.
Then I point out our crash rate has dropped by 70% in the 6 months I’ve been working on the project.
The cause was that one of the VIPER classes forced unwrapped an optional reference to its view. Force unwrapping is without the if, just assume it’s always valid and crash if it’s not.
The problem was they made this assumption because the view is never nil except for one uncommon edge case, if a network operation completed after the view was disposed. So the fix was easy, remove the force unwrap (you should never force unwrap in Swift, it’s terribly bad).
But what about our other VIPER views? It took me a couple minutes to review all 20, and Trey all had the same identical flaw. Fixing them took seconds, but individually testing the fixes took a few hours.
Should I have left those other 20 views alone so our app could mysteriously crash for some of our hundreds of thousands of users?
A more recent bug I found on my own was a memory leak caused by our main views network closures holding strong references to the ViewController. This caused code to fail that was still in old dead copies. So I fixed the closures and we no longer have old copies of the VC handing around forever. I didn’t have time to do a search at that moment, but I took notes in my dev log so I can review every view controller (50 or 60) for similar problems. Why wouldn’t I? How many known bugs and random unduplicated problems would go away if we got our Viewcontrollers memory Managemnt right? im betting at least a few, and that I’ll eliminate some future problems before they can even be found.
Software engineering is usually 80% fixing bugs. Developing rigorous standards to prevent their formation can give you much more time to build new and improved features.
One part I strongly disagree with, is this passage:
> Most programmers want to "hear" the "information" that their program component works. So when they wrote their first function for this project three years ago they wrote a unit test for it. The test has never failed. The question is: How much information is in that test? That is, if "1" is the passing of a test and "0" is the failing of a test, how much information is in this string of test results:
> 11111111111111111111111111111111
> There are several possible answers depending on which formalism you apply, but most of the answers are wrong. The naive answer is 32, but that is the bits of data, not of information.
Just because a test has passed in your Continuous Integration environment 100% of the time, doesn't mean that test is worthless. I have checked in many tests that have never failed in CI - but have failed when I was working on my code. However, since there's no shiny team-visible metrics with bar charts about how often a test failed in a local workspace, people can wrongly assume that a unit test is worthless.
> Now, how many bits of information in this string of test runs?
> 1011011000110101101000110101101
> The answer is... a lot more. Probably 32.
If I see that in our CI environment, the answer is 'this test is nearly-worthless.' The long answer is 'If it's a unit test, there's a race condition, if it's an integration test, there's a race condition, or it's flaking for reasons outside our control.'
> Another client of mine also had too many unit tests. I pointed out to them that this would decrease their velocity, because every change to a function should require a coordinated change to the test. They informed me that they had written their tests in such a way that they didn't have to change the tests when the functionality changed. That of course means that the tests weren't testing the functionality, so whatever they were testing was of little value.
... That's the whole point of blackbox testing. If observed behaviour is not expected to change, neither should the test. If a refactoring forces you to update the test, then, yes, you are testing at the wrong level of abstraction.
This piece could have saved a dozen pages if it just told us to stop testing private methods, and write more integration tests.
From quickly reading through it, for instance I found the following things directly objectionable:
> "Unit testing was a staple of the FORTRAN days".
Such an attempt at discrediting something can be applied to anything, and it comes off as disingenuous, dishonest.
> "Unit tests are unlikely to test more than one trillionth of the functionality of any given method in a reasonable testing cycle. Get over it." ... Trillion is not used rhetorically here, but is based on the different possible states given that the average object size
Who says you have to test the method? Who says you're not allowed to test functions which operate on a known, closed subset of data?
False dichotomy and plain bad math.
> If you find your testers splitting up functions to support the testing process, you’re destroying your system architecture and code comprehension along with it.
That may be right. Or it may be possibly completely backwards. It's quite impossible to tell really, without asking why they are splitting up those functions.
If people do this for gaming some sort of system about "at least 80% coverage" or whatever, then clearly you should ask why people feel the need to game the system. Gaming the system, no matter what aspect, leads to bad choices.
Me however, I'm not splitting up my functions to game the system. I'm splitting up my system to separate data-retrieval from data-processing (so I can directly test processing without depending on whatever retrieval depends on).
I'm splitting up my functions to give them name and intents, this makes the code speak much clearer about what it is doing and why. The function name should be "what". The contents should be "how".
Basically I'm splitting up my function for good reasons.
If a function is long enough to contain enough actions, intermediate variables, loops and conditionals so that you need to stop thinking about and wonder what they do, and what their role is inside that function... That functions should be split up, because you do not have a clear "what" and "how" delimiter.
And guess what? That also assists testability. You can test that the "how" correctly assesses the "what".
This is not destroying your system.
Some of the functions you end up with may even turn out to be reusable across the class/system, meaning you increases consistency and correctness as a result too.
Again: This is not destroying your system. Quite the opposite.
I could go on, but I'm just at the 4th page of 21, and my comment is already the biggest in the thread, and I'm not planning on making a blog-post length response.
Just saying this document is severely biased, contains factual errors and is not a good thing to rely on to present an argument.
Write good tests, automate what you can, and make sure you're using your chosen implementation correctly and appropriately.
I worked on projects that tested every getter and setter. Many of these getters and setters existed only for other tests! Complete waste of time testing those...
Having written an accounting system, including web/user interface and asynchronous/deferred coordination (such is HTTP/browser programming), for the last three years, I can say that functional programming is increasingly helping my team stay sane.
We do TDD always; sometimes in the form of unit tests; sometimes in the form of integration tests and we tend to write as many of our tests as random/generative tests to avoid having to write large code-bases. I've spent the last two days making a piece of our domain monoidal, having defined the three laws as property/generative tests; the rest of the time is an interactive play with the generator, to see if it can come up with counter-examples, to the code I just wrote.
I normally go about coding by writing a very high-level integration test (at the top-most layer that I still have code in the service/frontend); then I write a huge chunk of the code until I think it's correct and looks pristine and easy to maintain. Now I run the test. If it fails and the test is correct — it tests the right thing and is easy to read — then I start at the top (highest level) of the stack, and write down the assumptions I have thought about while designing the code — as unit tests. Until one passes (at some level n); at the higher level n+1, an assumption/unit test is now broken and I can divide-and-conquer until I find the line of code that doesn't work.
This, together with purification; the methodological extraction of pure functions (things that only have one output ever, for a particular input), makes it possible to avoid testing any side-effects/async things (they simply flow non-async data between pure functions).
This, together with first-level-values for control flow, aka. not using exceptions, and generative/random testing, makes it so that ALL of the input has valid output, makes it so that all functions are total functions. And this in turn makes the code uncrashable and bugfree (for the domains of bugs the above methodology removes).
The domains of bugs that the above doesn't eradicate, are primarily cross-browser bugs on the GUI-side, or UX bugs, where a feature is hard to use/understand. On the server-side we sometimes crash when our logging storage in ElasticSearch goes down and never comes up, the intermediate buffer (Logstash) fills up, and then the app buffer fills up and then the app livelocks, waiting for the logging to drain. (=> operations). The second most frequent reason we have any exceptions/errors/bugs is DNS not working.
The first year of writing the software, we still invoked libraries that threw exceptions, but now we've rewritten them all to do control flow with first-class values, so that is not an issue any longer.
I just wanted to share how we do stuff at qvitoo :), in case it helps anybody.
I can count on one hand the number of well written project test suites I’ve seen - even on code bases with 100% coverage.
We, as a group, do not seem to be good at writing them. I’m not sure why, and would love a discussion about it...
https://www.nytimes.com/2017/09/11/opinion/equifax-accountab...
In case anyone was fooled by this comment into thinking the article takes a lazy approach to testing, it argues for keeping a certain class of unit tests; for another class, preferring system tests instead; and for third class, turning them into assertions that ship with production code where possible. That's not lazy. Depending on how you currently test your code, it may actually be a harder standard to meet. It's possible you personally have nothing to learn from it, but I would wager there's something in it that will make you see your tests in a different way.