In my experience, I've seen a lot of code that was written in a way that was not unit-testable, for what was the same reason that later turned out to be poorly-maintainable.
In my experience, I've seen a lot of code that was written in a way that was not unit-testable, for what was the same reason that later turned out to be poorly-maintainable.
I've written some high-quality code and tested it to death¹, only to be told that it probably wasn't unit testable (because dependencies were hard coded instead of injected, IIRC). Making that code unit testable would have complicated it, thus lowering its maintainability.
[1] Not a single bug has ever been found after my tests.
Here’s Fowler’s take: https://martinfowler.com/bliki/UnitTest.html
Where unit tests have an advantage and where you might have been criticized is their ability to test permutations. If you write three tests for one function and three tests for a dependency of that function, you've essentially covered nine cases with six tests. In a complex application with a large dependency graph, that ability to multiply the value of each test can be difficult to overcome when testing at a higher level. Integration tests rarely cover edge cases. Additionally, unit tests are often quite simple to read and have limited setup which means that a programmer coming to the code base without prior knowledge of the code (I always pictured that being me six months in the future so that I'd be as kind as possible :-) can easily see exactly what is expected of the code.
But we shouldn't be dogmatic about the kind of testing we do and lose sight of the goal of that testing in the first place. Any test that furthers a future programmer's ability to rip the application code apart, put it back together and quickly know whether all the previous expectations of the code are still met is a good test. There's no hard and fast rule and judgment and experience are still necessary to know what to do in a specific situation. Too many people focus on the details of automated testing and lose sight of the goal.
My original point was that code that's hard to test is often hard to refactor.
I agree with that. The poster I was replying to said somewhat the opposite. He said:
> Making that code unit testable would have complicated it, thus lowering its maintainability
My basic point was that less complicated doesn't mean maintainable. There's no perfect design in the face of future changes in requirements. So the most important quality of the code we write today is the ability to refactor it and know that it still satisfies all the initial requirements that haven't changed. Unit tests give you that property, lack of complexity doesn't. Simple code can be better than complex, but it's a secondary concern to a comprehensive test suite. I'll take a full test suite over code quality any day because it allows me to go in and add the quality later without fear of breaking things.
I guess I misspoke when I said "how the application is structured" since it you've interpreted it differently than I meant it. It was a reference to the line I quoted above...that complicating code lowers it's maintainability. I just meant that complexity and tested are separate concerns and that tests more so than simplicity make code maintainable. So I think we agree more than we disagree.
> There's no perfect design in the face of future changes in requirements. So the most important quality of the code we write today is the ability to refactor it and[…]
I take issue with the idea of such protean requirements. I'm more the YAGNI type: those changes you're trying to anticipate? They likely won't happen. I'd rather keep the code simple, and add the flexibility only when needed. Simple code is easier to modify anyway.
(Of course, I'd rather not ossify an architecture if I can avoid it, but an ossified architecture is often the sign of tight couplings, which to me are a form of complexity. If we keep the interfaces between modules small and simple, we are more likely to end up with something flexible.)
Read my above point about the cartesian product of tests...if you're doing this, you're almost certainly not getting particularly good coverage of edge cases even if your test coverage percentage is high. If I have test A1, A2, A3 that use a mocked B and B1, B2, B3 that use a mocked C and C1, C2, C3, I've essentially tested 27 different cases that you'd have to write in your hard-coded dependency style while writing only 9 tests. When an application reaches even a moderate amount of complexity and that dependency graph grows beyond 3, that ability for unit tests to cover a geometrically-increasing number of ways that the software behaves is going to be hard to overcome with integration tests like you seem to have been writing. But, again, maybe your dependency graph wasn't big enough to need this. Maybe there was some other mitigating factor that made your situation different. I don't want to speculate on any individual situation since I haven't seen the code and there's no universally correct answer.
> I'm more the YAGNI type: those changes you're trying to anticipate? They likely won't happen.
The one virtually certain thing about the future is that there will be change. Of course you mostly can't predict what will change, but you can be damn near certain that something will. YAGNI is not about predicting that nothing will ever change, it's only about avoiding guessing at what will. Tests on current functionality are not guessing about what will change. The whole point of a test suite is that it avoids ossification. What can be less rigid than something that can easily be pulled apart and put back together in some other form while being sure that it still works? Simple and flexible without that property of safe reorganization is still a landmine waiting to go off in the face of future change.
I used property based tests. With relatively little effort, I ran thousands of tests. Allow me to laugh at your measly 27 from atop my high horse.
More seriously, though, I don't know how I could have mocked my ring buffer without writing another full featured ring buffer. Mocking looks a bit ridiculous in this case. Same for the queue. And the reason why this is so is not tight coupling. It's because my queue relied on the functionality provided by the ring buffer, and my reactor relied on the functionality provided by the queue. If this wasn't the case, I would have written a simpler ring buffer and queue to begin with.
Also, why mock stuff when I can test my code under real conditions?
> The whole point of a test suite is that it avoids ossification.
This feels backwards. If you just mean that test suites prevent the unintended consequences of change, sure. They're like a home made type system in that respect. But if you suggest we should bend a code base just to make it "let's mock everything" testable™, then no. Don't mock stuff for the sake of it, it's a wast of time.
I'd rather concentrate on making the code usable in the first place. With small, well defined interfaces. In my experience, such designs are naturally testable.
> Allow me to laugh at your measly 27 from atop my high horse
If you're going make snide comments, I'll stop this discussion. Suffice it to say you missed the point. 27 isn't just a number, it's a number arrived at by applying an exponent. Exponents get really big in a hurry. As the author of a crypto suite, I'd expect you to pay exponents a bit more respect. Testing 3 different cases per dependency gives you 27. Testing 10 gives you 1000. Property based testing is not a substitute for this kind of coverage. It is, however, complementary. It allows you to increase the base number which is raised to the exponent.
> I could have mocked my ring buffer without writing another full featured ring buffer. Mocking looks a bit ridiculous in this case.
It's safe to assume from this comment that you don't really understand mocking. Mocks aren't an implementation of an interface to be used in tests. Mocks verify that they were called in an expected way/order and give back the appropriate value. They don't actually do anything. They're highly coupled with each test case. This is why most mocking frameworks dynamically generate the mock per-test and let you simply verify the call(s) and tell it what to return.
> This feels backwards
It's absolutely not backwards. The main purpose of a test suite is to find a very specific type of bug...regressions. It's not to avoid other sorts of bugs other than the fact that forcing developers to write tests will force them to think about their code in ways that make them realize they missed something during implementation.
But the most important thing is to be able to rely on them to tell you when you've broken something that previously worked. Every single time I've come to a code base that was either missing tests or relied too heavily on integration tests, we've had a significant problem with regressions and, over time, developers shied away from making significant changes to the code base because they were afraid to break it. And the result is that ossification that you try not to achieve. But when you see a code base with a commitment to unit tests, you also see one where developers are free to make the necessary radical changes without fear of breaking existing functionality. They make the changes, run the suite and let it tell them everywhere they forgot to change, whether that be in application code or test code.
---
Yeah, C++. I don't know of any mocking framework there, and I'm not going to look for one —my current opinion on that stuff veers towards "worse than useless". (You just reinforced that notion by the way, see below.)
---
27 is a number you arrived by fantasizing about how mocking make your tests more thorough than they actually are —and you're not even testing under real conditions! Just an outsider's scepticism.
Your exponent is also driven by the number of dependencies, whose increasing number requires an exponentially growing number of tests. Property based testing on the other hand (search for "QuickCheck") generates as many random tests as one feels like using. (Of course, one has to test for several properties, but then each property can be thoroughly tested.)
> It's safe to assume from this comment that you don't really understand mocking.
Indeed, thanks for the explanation. But that just showed mocs are fundamentally incompatible with property based testing: I'd need to generate a moc for each random test case. Nope. I'm keeping my property based tests, they're much more powerful.
---
So you were talking about using your test suite like a type system on steroids. Good. Just keep in mind that there's a stark difference between maintaining a good test suite, and bending the code around some narrow idea of testability. I only do the former. Oh, and I tend to write tests after I write the code. Property based test don't make sense before the module is done anyway.
A quick note about unit vs integration test: in our little A->B->C example, testing C, then B (without mocs), then A (without mocs) are all unit tests in my book. A few internal dependencies aren't going to change that.
By the way, I'm talking about compile-time dependency, of the kind that don't involve runtime shared state. Shared state open a whole 'nother can of worm, and I do see the need to separate unit and integration tests for those. (I also see the need to avoid shared state as much as possible.)
Incidentally, here's two C++ mocking frameworks found from Google, in case you're more open minded than I'm giving you credit for:
https://github.com/google/googletest/tree/master/googlemock (part of https://github.com/google/googletest)
Here's a primer: you think of some property that must hold for your object. Take a ring buffer for instance: if you push N elements, it must be of size N (assuming it's capacity is N or more). Then you generate random input, and use it to test your property. One property, thousands of tests. Once that's done, you can move on to some other property (like, the ring buffer must have FIFO behaviour).
> to declare something as popular and influential as those two concepts "worse than useless" without bothering to understand them is, well, ballsy.
It wouldn't be the first time our industry demonstrated blatant inadequacy —at least in hindsight. We're still a young field, popularity doesn't mean much. Maybe I would have considered mocking if I didn't know of property based testing. But the power difference between singular unit testing an property based tests is too great to ignore.
The reason I call mocking frameworks "worse than useless" is because they turn you away from property based tests. If they were at least compatible with them… I'll allow one exception: if you can use the same mock with all the random input, or at least easily generate a moc per random input, then it wouldn't be too bad.
It wouldn't be too good either, though. I'd rather test with the real stuff. If there's a regression, it will be revealed, with or without mocking. Actually, if some dependency isn't properly tested (distracted dev or something), not mocking it gives you a second chance at detecting bugs. On the other hand, have you ever seen a regression that would have gone undetected if you didn't used mocking? That doesn't seem plausible.
The only use I see for mocks is to track down bugs once they are detected —by narrowing the search space. But I've never needed it in practice, with one possible exception: Lua.
Are you using dynamic typing? That would explain a lot.
Choosing to write unit tests instead of integration tests does not preclude the use of property based testing or any other technique that covers how you test code. As I stated before, if you were to use property based testing at the unit test level, it would be all the more powerful because you'd get that same exponential coverage just with significantly larger number of effective tests per unit of code. From the A, B, C dependency graph example, if you used property based testing to test A thousands of times, B thousands of times and C thousands of times, you'd effectively be testing thousands times thousands times thousands (aka billions) of different scenarios. And you'd still be able to use mocks because mocks are created in test code and can be as dynamic as you need them to be.
I've done testing in many languages, both static and dynamic. Dynamic languages, thanks to the lack of static type checking, often don't need a mocking framework because the language is flexible enough that you can swap in a mock-like object at runtime. But the concepts are the same. And there really is a ton of value in testing code you've written in isolation. It's the testing equivalent of separation of concerns. And just like some developers never grok the single responsibility principle, some developers never grok the need to test code in isolation. You've obviously understood how to encapsulate your concerns based on the design you implemented. And yet the testing strategy you described is the opposite...trying to test everything at once rather than focusing on each test covering one specific bit of code.
Me on the other hand could pretend something to this effect. By testing A, B, and C three times each, I effectively test B 6 times, and C 9 times, for a total of 3+6+9=18 "tests". But even then I would hesitate to pull such a number out of my ass, because the extra "free" tests are likely redundant with the real ones down the line. (And if my coverage is any good, those are guaranteed to be redundant).
Surely you don't pretend mocking lets you test more things than my approach? Surely the mocs don't behave any differently than the real thing on the inputs they are given? How come then that my "integration" tests don't supersede your unit tests?
> Choosing to write unit tests instead of integration tests does not preclude the use of property based testing
I assume at this point that your definition of unit test means mocking everything, except perhaps the standard library. You said earlier:
> Mocks verify that they were called in an expected way/order and give back the appropriate value. They don't actually do anything. They're highly coupled with each test case.
So, if tests are generated randomly, mocs must be adapted to each random test, right? Now you need some moc generator, that will generate a moc for each test case. This veers eerily close to actually implementing the interface we were trying to moc.
Now I realise I did miss something:
> Mocks verify that they were called in an expected way/order and give back the appropriate value.
Why would they do that? Oh, I see: we are mocking a mutable object that is used somewhere else. In other words, shared state. The one case for which I understand isolation could be good.
This may all be a big misunderstanding. The A->B->C dependency I was thinking of is pretty simple: C is an implementation detail of B. Each B contains and uses a C, but that C isn't used anywhere else. Thus, mocking B or C would mean testing the implementation details of A, and that is a big no-no —the kind of thing Uncle Bob would call TDD gone wrong.
> And there really is a ton of value in testing code you've written in isolation.
What value exactly? Does isolating objects in the tests make those tests catch more bugs? What kind of bugs? Does isolating objects make detected bugs faster to track down? Does isolating objects in the tests influence how the code is written in the first place? How? Would the test suite run faster?
> And yet the testing strategy you described is the opposite...trying to test everything at once rather than focusing on each test covering one specific bit of code.
Now that's a strawman. I don't test everything at once. If I did, I would merely test A. No no no, I test and debug C first. When I'm confident it is bug-free, I test and debug B. Bugs are easily tracked down, because C is bug free, so if something went wrong it must be B. I only test A last, when I'm confident B and C are both bug-free.
What's important is to have test coverage and to keep exercising it when the code changes.
The best way to achieve this depends case by case. Sometimes unit tests are better, sometimes integration test is the way to go.
I've had a lot of times where no bugs were found with huge amounts of unit testing over trivial code. It would have been easier and of more use to just do manual inspection (code review) with a bunch of people instead of writing the tests at that level, and instead writing tests at a higher level of complexity which involved more code/components.
In a perfect world you'd test all the branches and possible values but out there in the real world, if you aren't finding bugs with your testing and there are still bugs in the code, you need to be spending your testing "dollars" (time/effort/talent) more efficiently.
I wrote those tests because I was writing foundational code. The outer layers hardly have any test, which is good enough for the proof of concept we were asked for. We'll need to be a bit more rigorous to polish this into production.
But my primary proxy is size. The less code the better. But even that have to give way to other concerns, such as style constraints, and straight up performance.
While I can't show the code I was talking about above (it was proprietary), I can show my crypto library, Monocypher¹. I think it is a good example of what I mean by high-quality code. (Yes, I am boasting. I believe this is justified. Also, quality requirements for crypto code are kinda off the chart.)
And for example, if one was optimizing for code readability over performance, I expect that would look very different (although you could do much of the same with macros, but I'm not sure if that's any better).
[1]: https://github.com/LoupVaillant/Monocypher/blob/master/src/m...
I thought of showing some of my OCaml code as well, but I didn't put nearly as much effort, so I'm not so sure about its quality.
#define FOR(i, start, end) for (size_t (i) = (start); (i) < (end); (i)++)
yuckEDIT: Also, what the hell is this macro defined in the middle of a function and then only used twice:
https://github.com/LoupVaillant/Monocypher/blob/master/src/m...
Why do you need this function for "constant time comparison to 0":
https://github.com/LoupVaillant/Monocypher/blob/master/src/m...
static int neq0(u64 diff)
{ // constant time comparison to zero
// return diff != 0 ? -1 : 0
u64 half = (diff >> 32) | ((u32)diff);
return (1 & ((half - 1) >> 32)) - 1;
}
I am incredibly skeptical that you've done better than the compiler with this bit twiddling stuff. You realize that on x86_64, it probably compiles down to 1-2 instructions to write `diff != 0`, right?Now, with that being said, this strikes me as a domain issue: the OP seems well-versed in crypto and foundation / backend code, which has relatively constrained input and output behaviors and is testable using smart reasoning and fuzzing. This is somewhat different from front-end or user-facing code, which involves handling epic quantities of mutable state and requires validation of a feature end-to-end. I suspect OP writes great, "bug-free" code in the backend / systems domain, but I don't think their testing practice or anecdotes are applicable across the industry.
I'll grant OP a Magic Cloak of Error Abolishment and an enchanted ring with +4 INT, +4 WIS, and +3 VARIABLE-NAMING...
Regardless, what's the very first thing I'm gonna do when I inherit his project for support or further development? Start ripping it apart to add tests so that I can verify the behaviour of the system pre, and post, change.
If that code base was tested before, even poorly, that means extending the tests and improving their web of assertions and I will be making positive changes to the code within hours. If that code base is untested then my maintenance activities will involve unsupported refactoring and WAGs about system behaviour once the new tests come online. Days of work will hamper the start of maintenance activities. From experience, it will also likley mean finding a bunch of previously hidden issues, bugs, and suspect behaviour that has to be re-analyzed and re-tested before maintenance work can begin in earnest.
A strong feeling that something is "bug-free" does not tell me that everything works unchanged on a new platform/OS/architecture, or let me throw in a newly conceived corner case, or check some environmental oddity... If we start out with the premise that systems spend most of their lifetime in maintenance, preparing the system for maintenance activities and trying to eliminate risk in that delicate window make a lot of sense. Not for a lone developer who has a bunch of domain knowledge in their heads, but for teams where new individuals are exposed to new domains and code at the same time.
Most of those come from my test suite. Without that, I'm back to -2 INT, -3 WIS, and a Cursed Cloak of Mistakes. I did let some bugs slip through the first time around…
> what's the very first thing I'm gonna do when I inherit his project for support or further development? Start ripping it apart to add tests so that I can verify the behaviour of the system pre, and post, change.
As far as Monocypher is concerned, I already did that. Test vectors, property based test, sanitisers, code coverage analysis, the works. The public interface is testable enough that you don't need to take apart anything to thoroughly test that library (it already is).
Huge help with "making positive changes to the code within hours". This change to Blake2b for instance took me a few minutes to make and verify. https://github.com/LoupVaillant/Monocypher/commit/80ebf55307...
> A strong feeling that something is "bug-free" does not tell me that everything works unchanged on a new platform/OS/architecture, or let me throw in a newly conceived corner case, or check some environmental oddity...
It helps that Monocypher has zero dependency (not even the standard library), uses fixed size integers almost exclusively, compiles without warning as C and C++ with GCC and Clang, and that crypto code is abnormally straight line.
> If we start out with the premise that systems spend most of their lifetime in maintenance
I made sure Monocypher required next to no maintenance (I won't have the time for significant maintenance work). It's mostly a matter of size and scope.
The "used twice" macro makes clear this is the exact same code. Would be hard to make sure it is otherwise. Easier to review that way (carry code is a nightmare to test and review).
The constant time comparison is necessary to ensure timing attacks cannot happen. The arithmetic trick may be a tad slower than a conditional branch, but its timings are consistent. Note: the compiler is still allowed to just use a branch, but in practice they don't.
But will the 1-2 instruction compare be constant time?
This was foundational code, used for inter-module communication all over the place.
> unless you had actually re-written the code as unit testable how do you know it's more complicated that way or less maintainable?
Because the code was small enough to allow me to envision the necessary modifications for dependency injections. I had 3 classes, with an A->B->C dependency chain, and no reason to inject anything if it weren't for some cargo cult about unit tests. Testing the hell out of C, then B, then A, proved quite sufficient without mocking or injecting anything.
> You're responding to a study with actual data, on a generally "preference" topic, with an anecdote so yes I do feel justified in pointing this out.
Whatever evidence the study actually has is weak. Small sample size, and the failure to analyse actual outcomes (bugs, speed of development…) mean we cannot possibly get much out of it. It's a good starting point.
I don't believe I have contradicted this study's evidence. But even if I did, my personal experience gives me way more evidence than such a study, so I feel perfectly justified in contradicting it on that basis. (More solidly settled science, that's another story.)
Problem is, you do not have a privileged access to my personal experience. You only know what I just wrote. And my written report of my personal experience means little, next to that study. You'd better believe the study before you believe my report.
Then you have your own personal experience.
Also I think you may be conflating DI with testability. Items can be u nit testable and easy to di while not being built for it. I write code that is testable and it generally happens to be di-able as well but that's not the goal. From my own experience.
Anecdotes are still a valuable in that they point out where you might be interested in looking further. You can't trust anecdotes, but they may indicate where to look.
> I think you may be conflating DI with testability.
I'm conflating DI with a stupidly narrow idea of testability. I'm fully aware that my code was easily testable (and thoroughly tested!) despite hard coded dependencies.