Unit Testing Is Overrated
tyrrrz.me
tyrrrz.me
Computational code handles your business logic. This is usually in the minority in a typical codebase. What it does is quite well defined and usually benefits a lot from unit tests ("is this doing what we intended"). Happily, it changes less often than plumbing code, so unit tests tend to stay valuable and need little modification.
Plumbing code is everything else, and mainly involves moving information from place to place. This includes database access, moving data between components, conveying information from the end user, and so on. Unit tests here are next to useless because a) you'd have to mock everything out b) this type of code seems to change frequently and c) it has a less clearly defined behaviour.
What you really want to test with plumbing code is "does it work", which is handled by integration and system tests.
[0] https://en.wikipedia.org/wiki/Command%E2%80%93query_separati... [1] https://www.destroyallsoftware.com/screencasts/catalog/funct...
Operate only on your inputs. Return all of your outputs. No side effects.
You can, after all, write "functional C". (It can be hard, though.)
The idea that your business logic should be isolated from external dependencies (in his case, by making the code (pure) functional). That makes it easy to unit test the business logic, and your integration tests should be minimal (basically testing a single path to make sure everything is talking to each other).
It has advantages but it's an expensive waste of resources if you have cheap, effective integrarion tests.
Gary was coming from the land of Ruby-on-rails where a full set of integration tests could take hours. In that environment, structuring your code to enable easy testing of complex logic makes a lot of sense.
Likewise in a large enterprise environment, where integration testing across a (usually messy) set of interconnected dependencies is a pipe dream.
It's true that over-architecting is something to be wary of, but as usual, there's no one-size-fits-all answer.
It doesn't matter if the whole test suite takes hours. CI servers don't need to be supervised.
It's a really expensive way of discovering that you wrote shit code.
Abstract code solves made-up problems, while concrete code solves real ones. Normally the best way to solve a real problem is by rewriting it as a series of made-up problems, and solving those made-up ones instead.
The made-up problems don't need to be pure computational. Instead, if you restrict them to pure ones, you'll lose a lot of powerful ones. They also don't need to fit functional programming well, but there is no loss of generality on imposing that restriction.
Also, the more abstract you make that code, the less they'll need changing and the better unit tests will fit. At the extreme, once debugged they'll never change. Instead, if your needs change too much, your concrete programs will simply stop using them and use some completely different ones.
For example, let's say you want to get some users from your DB in response to an HTTP call. We rewrite this problem in terms of crafting some SQL query, taking some data from the HTTP request to create that query. We can of course easily test that the code creates the query we designed that the query contains the right information from the HTTP request etc. But, if we don't actually run the query on the actual DB with the actual users, we don't really know if our query does the right thing, even if we know our code creates the query we intended. And, if the DB changes tomorrow, our very abstract code that parametrizes a particular SQL query will still need to change, so our existing unit tests will be thrown away as well.
This is the kind of plumbing code the OP was talking about, and I don't think you can reduce the problem in any way to fix this (especially if the DB is an external entity).
A ton of dense, mathy code like hash computation, de/serialization, sin/cos computation, etc. is usually best implemented in a memory efficient C-style way but lends itself to be used in a very functional way; inputs and outputs without any retained state or side effects.
I think that subtlety is hard to articulate and gets lost.
A lot of code can be in this are where it is absolutely unit testable, but the unit tests are almost entirely useless, as the code only ever changes because the input or output types change, so the tests also need to change.
I think of this in terms of code that is 'authoritative' for its logic or not.
For example, a sorting method is authoritative - it is the ultimate definition of what sorting means. Also, a piece of code that validates some business rule defined in a document is the authority for that business rule.
But a piece of code that takes input from the user and passes it to some other piece of code is not authoritative for this transformation. The functionality of this kind of code is not defined by some spec, but by 'whatever the other piece of code wants to receive', which may be arbitrarily hard to define.
Depending on the complexity of the transformation, there may still be reasons to test parts of this code, at least to ensure that a new field here doesn't affect the way we transform that other field there, but often only small pieces of it are actually worth testing.
Unit test algorithmic code; use integration tests for everything else, i.e plumbing code.
I agree very strongly with this but a lot of people will be very unhappy with this idea.
That allows strict separation of all I/O from testable business logic.
If you can't separate pure logic from your I/O, it means you have a Russian-doll program that looks like:
readFromApi {
doSomeBusinessLogic{
writeToPersistence{
...
Instead of a pipeline like:a <- readFromApi
b <- doBusinessLogic(a)
c <- writeToPersistence(b)
If you do things this way, you can always isolate your business logic from your dependencies.
By all means, if the transformation is non-trivial, and it is captured entirely in the logic of this method, not in the shape of the API and the DB, then you should unit test it (e.g. say you are enforcing some business rules, or computing some fields based on othee fields). But if you're just passing data around, this type of testing is a waste of time (you don't have reasons to change the code if the API or DB don't change, so the tests will never fail), and brittle (changes in the API or in the DB will require changing both the code and the tests, so the tests failing doesnt help you find any errors you didn't know about).
a <- readFromApi ( Input x )
b <- doBusinessLogic(a) ( f(x) )
c <- writeToPersistence(b) ( Output y = f(x) )
You can also imagine that there are more than one lookup from the db or service calls as I/O in different parts of the pipeline (g(f(x) etc.), but it's always possible to have state pulled in explicitly and pushed down explicitly into business logic as an argument. It tends to make programs have flatter call stacks as well.
So I would argue you don't actually have business logic then. Your service is anemic, and you have a data transformation you need to do. I definitely think that you should do an integration test for that.
Moving JSON -> Postgres or whatever is something that you absolutely still can test with the output of the DML statement by your DB library. It may be a silly test, but that's because if there's no business logic, it's a silly program _shrug_.
>> Write tests. Not too many. Mostly integration.
If errors in your system result in death, and if changes must go through an expensive and time consuming process to be approved, and then an expensive and time consuming process to be applied, you should spend a lot of time ensuring your design is sound, and your implementation matches your design. A good place for formal methods.
If you're writing server side code, and deploy takes 5 minutes, you can be a cowboy for most things that won't leave a persistant mess or convince customers to leave.
If you're writing client side code that needs to go through a pre-publication review, neither cowboy or formal methods is a good choice.
I test the parts that are actually mine as best I can, but most of my debugging consists of driving it by hand.
More importantly, that your app works with the mocks doesn't give you good information about weather your app works with the actual services.
Using the old Asteroids arcade game [1] as an example: The business logic is how many lives the player has, what happens when you shoot asteroids (they break up, or disintegrate if they're small), what happens when you reach the edge of the map (you wrap around the other side), what kind of control scheme there is (there's momentum in asteroids, you don't stop on a dime) etc.
i unit test business logic since that is the core of the application and MUST work as expected.
i'm not going to unit test a link that someone clicks on goes to the page they expect.
If you are writing on one shot script to transmute data from one format to another for say an upgrade, I don't care if you have unit tests if I am confident it has been manually tested to satisfaction. No repeatability, no regression requirement. There could and likely is value in TDD so tests might still be a thing if that is how you work. No objection there.
If you are developing the plumbing code that will ensure my system adheres to financial regulations and, if it were to break, land me in jail for negligence, you can be damn sure I'm demanding a test that will be run everytime that system is built/deployed.
I wrote unit tests >10 years ago for formatting a string for postal codes that I know are still run to this day on every commit because if they get it wrong there is legal recourse for the company that owns that system.
It's also super quick to fix and failing at build is quicker and cheaper than failing in prod, even without the recourse. That test took me all of 1 minute to write. Bargain.
Unit tests and automated tests are two completely different concept.
If it's critical for your business I'd categorize that as business logic, not plumbing code, well deserving of unit test coverage.
What I've been doing is writing as many parts of the game as libraries as is possible, and then implementing the minimal possible usage of that library as a semi-automated test. For instance, our collision system is implemented as a library, and you can load up a "game" that has the simplest possible renderer, no sound, basic inputs, etc. and has a small world you can run around in that's filled with edge cases. This was vastly easier than trying to write automated tests for 3d collision code, and you get the benefit of testing the system in isolation, if not automatically. For other libraries like networking, the tests are much more automated, but they poke the library as a unit, rather than testing all the little bits and pieces individually.
I agree with this and would go even further. Divide your code into "stateless" functional code and "stateful" objects code.
Original OO was encapsulating things like device drivers that did I/O--it didn't represent data.
If you don't interleave your stateless business logic with your stateful persistence, it's easy to mock "objects" that do the plumbing, and all the meat of the program is unit tests.
Fwiw, the DI model (Guice, Spring, etc.) in modern Java/Scala shops closely hews to this, even if people don't mentally categorize it as such.
IINM you are basically referring to the difference between static and instance methods in languages like C++ and Java.
Putting code that neither reads nor writes the object state and instance methods is a common mistake made in both those languages.
That said, both stateful and stateless code are good candidates for unit testing, especially when the code under test is a state machine, rather then just a data encapsulation mechanism.
Your _program_ should have the flow of a function. At the architectural level, who-the-ef cares about about static vs instance methods in Java (I say as a person with 23 years of Java experience.) It has nothing to do with languages. You can do this in any language you want.
You want to have your inputs go through a process where you have (1) INPUT state transfer, (2) some computation F(INPUT), (3) some output and state transfer, or RESULT = F(INPUT).
If you do not have (1) or (3)--I hate to break it to you--but all your program does is burn CPU. If you don't have (2), your program does nothing at all.
The key thing with scalable systems is they manage complexity well. If you're at the level where you're worried about "static or instance methods", you're not dealing with how data changes in large systems at all. Those words are at the level of state within a language.
You need to optimize at the global systems level.
Who-the-ef should care is anyone who has to implement or maintain the code. After all, the debate at hand is what is worth unit testing, which very much concerns the programming language and the actual implementation. Don't know about you, but I both architect the system and write the code.
> If you do not have (1) or (3)--I hate to break it to you--but all your program does is burn CPU.
I haven't written production code that doesn't have (1) or (3) in my 25 years of programming, so not sure who you are talking to here.
> If you're at the level where you're worried about "static or instance methods", you're not dealing with how data changes in large systems at all. Those words are at the level of state within a language.
You have to tend to this stuff at both the generic data processing and language level. Using a given language's constructs for differentiating between stateful and stateless code is an important part of making the code document itself.
Coding style matters.
It. does. not.
If it did, PHP wouldn't be running half the world. Structure and systems matter.
OK 'brah', whatever works for you!
> If it did, PHP wouldn't be running half the world
PHP has a style guide, and there is such a thing as clean, readable PHP code.
https://www.php-fig.org/psr/psr-1/
I bet massive scale PHP based apps (like you know, Facebook) probably enforce style in their codebase.
I think of the "computational" type more as a "deterministic data transformation" type. That applies to transformations of any data whether text, images, or the state of a machine.
I think of plumbing as the movement of data without any transformation, or if a transformation occurs, it occurs at and abstracted layer that must be unit tested itself independently.
Speaking as a formerly young and arrogant programmer (now I'm simply an arrogant programmer), there's a certain progression I went through upon joining the workforce that I think is common among young, arrogant programmers:
1. Tests waste time. I know how to write code that works. Why would I compromise the design of my program for tests? Here, let me explain to you all the reasons why testing is stupid.
2. Get burned by not having tests. I've built a really complex system that breaks every time I try to update it. I can't bring on help because anyone who doesn't know this code intimately is 10x more likely to break it. I limp to the end of this project and practically burn out.
3. Go overboard on testing. It's the best thing since sliced bread. I'm never going to get burned again. My code works all the time now. TDD has changed my life. Here, let me explain to you all the reasons why you need to test religiously.
4. Programming is pedantic and no fun anymore. Simple toy projects and prototypes take forever now because I spend half of my time writing tests. Maybe I'll go into management?
5. You know what? There are some times when testing is good and some times where testing is more effort than it's worth. There's no hard-set rule for all projects and situations. I'll test where and when it makes the most sense and set expectations appropriately so I don't get burned like I did in the past.
Overall, the blog post says, unit tests take a long time to write compared to the value they bring - instead (or also) focus on more valuable automated integration tests / e2e tests because it is much easier than it was 10-20 years ago.
Your comment on the other hand, less so...
If you do make a slight tweak somewhere, the compiler will tell you there’s something broken in obscure place X that you would find out at runtime say with Ruby or Python.
THATS the winning formula. I’ve written so many tests for Python ensuring a function’s arguments are validated rather than the core logic/process of it.
If you're doing React + Typescript give Reasonml which is a syntax sugar on top of Ocaml that compiles using bucklescript a go. Ocaml has the fastest compiler out there.
Meanwhile the plugins and IDE integrations for Reason/Ocaml and F# are ready to go from the start and work pretty well.
Not so fast. For some problems it's great, for other ones it's not.
Have you tried writing numeric or machine leaning core in Haskell? You'll notice that the type system just doesn't help you enforce correctness. Have you tried writing low level IO? The logic is too complex to capture on types, if you try to use them you'll have a huge problem.
Rust's got a very Haskell-like type system, but it's a systems programming language. People are literally writing kernels in it. I think this is a pure-functional-is-a-bad-way-to-do-real-time-I/O thing, not a typing thing.
That said, I don't think it's impossible to type IO. https://lexi-lambda.github.io/blog/2020/01/19/no-dynamic-typ... isn't the same problem, but it's related.
If you try to verify the kind of state machines that low level I/O normally use with Haskell-like types, you will gain a huge amount of complexity and probably end with more bugs than without.
Let's say you're writing a /dev/console driver for an RS-232 connection. Trying to represent "ring indicator", "parity failure", "invalid UTF-8 sequence", "keyboard interrupt", "hup" and "buffer full" at the same level in the type system will fail abysmally, but that's not a sensible way of doing it.
I could definitely implement this while leveraging the power of Rust's type system – Haskell would be a stretch, but only because it's side-effect free and I/O is pretty much all side-effects.
Only half your time? You're doing testing wrong if it doesn't take 80% of the time ;-)
I have a love hate relationship with testing. Working for myself as a company of one, some of the benefits testing bring just don't apply. I have a suite of programs built in the style of your point (1). The programs were quick to market and hacked out whilst savings ran out not knowing if I would make a single sale.
Sales came, customer requests came, new features were wanted, sales were promised "if the program could just do xyz". More things was hacked on. The promise of "I will go back and do this properly and tidy up this god unholy mess of code" slowly slipped away that I stopped lying to myself I would do it.
Yes there was a phase of fix one problem add another, but I have most of that in my head now and has been a long time since that happened.
Not a single test. Developing the programs was "fun" and exciting. Getting requests for features in the morning and having the build ready by lunch kept customers happy.
Now I am redoing the apps as a web app for "reasons". This time am doing it properly, testing from the start. I know exactly what the program should do and how to do it, unlike the first time when I really had no idea. But still, I Come to a point and realise the design is wrong and I hadn't taking something into consideration. Changing the code isn't so bad, changing the tests, O.M.G.
I am so fed up of the project, I do all I can to avoid it, it is 2 years late, I wish I never started it. The codebase has excellent testing, mocks, no little hacks, engineering wise am proud of it. The tests have found little edge cases that would have been found out by customers so avoided that. But there is no fun in it. No excitement. Is just a constant drudging slog.
Am trying to avoid dismissing testing all together, as I really want to see the benefit of it in a production substantially code base. If I ever get there. At the moment, the code base is the best tested unused software ever written IMO
I've worked for myself as well and know what you mean. In my situation, I was able to save myself from testing by telling my customers "this is a prototype so expect some issues".
The thing about testing that never really gets talked about it is, what's the penalty for regressions? What's the consequences if you ship a bug so bad the whole system stops working?
Well, if you're building a thing that's doing hundreds of millions in revenue, that might be a big deal. But you? You're a team of one! You rollback that bad deploy and basically no one cares!
Your customers certainly don't care if you ship bugs. If it was something important enough where they REALLY cared, they wouldn't be using a company of one person.
So, go for it. Dismiss tests until you get to a point where you fear deploying because of the consequences. Then add the bare minimum of e2e tests you need to get rid of that fear, and keep shipping.
Human lives, customer faith in product, GDPR violations, HPPA violations, data, time/resources in space missions
https://medium.com/@ryancohane/financial-cost-of-software-bu...
I somehow doubt that comparing this 'team of one project' to the Mars Climate Orbiter leads to any useful conclusions. It's a nice bit of hyperbole though!
Anyways..this was to address the issue of a bug. I took the comment of "it's just a team of one" as a way of trying to justify not putting your engineering due diligence into delivering a product to the customer.
Yes. This is exactly what this person should do. Stop worrying about arbitrary rules and just deliver the damn product already. A hacky, shitty, unfinished product in your customer's hands that can be iterated on beats one that never got shipped at all every day of the week.
I've delivered a number of products (in the early days of my career) to clients where data loss happened and while not fun, it also didn't significantly harm the product or piss off said client. I saw my responsibility primarily to do the best I could and clearly communicate potential risks to the client.
> I took the comment of "it's just a team of one" as a way of trying to justify not putting your engineering due diligence into delivering a product to the customer.
That I do agree with, but 'due diligence' is a very vague concept. I guess honest communication about the consequence of various choices is perhaps the core aspect?
And of course 'engineering due diligence', in my opinion, includes making choices that might lead to an inferior result from a 'purely' engineering perspective.
Having said all that, I find that it's better to avoid doing some unit tests when building your own project. It can be better to do the high level tests (some integration, focused on system) to make sure the major functionality works. In many cases, for an app that's not too complicated, you can just have a rough manual test plan. Then move to automated tests later on if the app gets popular, or the manual testing becomes too cumbersome.
It's still good to have a few unit tests for some tricky functions that do complicated things so you aren't spending hours debugging a simple typo.
- Is the language you're using dynamic? Large refactors in Ruby are much harder than in Java, since the compiler can't catch dumb mistakes
- What is the likelihood that you're going to get bad/invalid inputs to your functions? Does the data come from an internal source? The outside world?
- What is the core business logic that your customers find the most value in / constantly execute? Error tolerances across a large project are not uniform, and you should focus the highest quality testing on the most critical parts of your application
- Test coverage != good testing. I can write 100% test coverage that doesn't really test anything other than physically executing the lines of code. Focus on testing for errors that may occur in the real world, edge cases, things that might break when another system is refactored, etc.
High test coverage comes from a history of writting tests there. Sadly people include feature and functional tests in the coverage.
For lexer and parser tests, I tend to focus on the EBNF grammar. Do I have lexer test coverage for each symbol in a given EBNF, accepting duplicate token coverage across different EBNF symbol tests? Do I have parser tests for each valid path through the symbol? For error handling/recovery, do I have a test for a token in a symbol being missing (one per missing symbol)?
For equation/algorithm testing, do I have a test case for each value domain. For numbers: zero, negative number, positive number, min, max, values that yield the min/max representable output (and one above/below this to overflow).
I tend to organize tests in a hierarchy, so the tests higher up only focus on the relevant details, while the ones lower down focus on the variations they can have. For example, for a lexer I will test the different cases for a given token (e.g. '1e8' and '1E8' for a double token), then for the parser I only need to test a single double token format/variant as I know that the lexer handles the different variants correctly. Then, I can do a similar thing in the processing stages, ignoring the error handling/recovery cases that yield the same parse tree as the valid cases.
A bug can be critical (literally life-threatening) or unnoticeable. And this includes the response to the bug and what it takes. When I write code for myself I tend to put a lot of checks and crash states rather than tests because if I'm running it and something unexpected happens, I can easily fix it up and run it again. That doesn't work as well for automated systems.
I keep tests together with the code, because of their documentation/specification value.
I do not write tests for functions which are compositions of library functions. I do not test pre/post-conditions (these are something different).
And I definitely do not try to have "100% test coverage".
Personally, I fast tracked through 2-4 out of sheer laziness but that's definitely my progression in regards to testing and pretty much everything related to code quality. It includes comments, abstraction, purity, etc...
More generally:
- Initially, you are victim of the Dunning–Kruger effect, standing proudly on top of Mount Stupid. You think you can do better than the pros by not wasting time on "useless stuff".
- Obviously, that's a fail. You realize the pros may have a good reason for working the way they do. So you start reading books (or whatever your favorite learning material is), and blindly follow what's written. It fixes your problems and replace them with other problems.
- After another round of failure, you start to understand the reasoning behind the things written in the books. Now, you know to apply them when they are relevant, and become a pro yourself.
The missing bit in the discussion is 1) churn, and 2) a devs ability to write fairly clean code.
Early stage and 'toy' projects may change a lot, in fundamental ways. There maybe total re-writes as you decide to change out technologies.
During this phase, it's pointless to try to 'harden' anything because you're not sure what it's entirely supposed to do, other than at a high level.
Trying Amazon Dynamo DB, only to find a couple weeks in that it's not what you need ... means it probably wouldn't make sense to run it through the gamut of tests.
Only once you've really settled on an approach, and you start to see the bits of code that look like they're not going to get tossed, does it make sense to start running tests.
Of course the caveat is that you'll need to have enough coding experience to move through the material quickly, in that, no single bit of code is a challenge, it's just 'getting it on the screen' takes some labour. The experience of 'having done it already many times' means you know it's 'roughly going to work'.
I usually try to 'get something working' before I think too hard about testing, otherwise you 3x the amount of work you have to do, most of which may be thrown out or refactored.
Maybe another way of saying it, is if a dev can code to '80% accuracy' - well, that's all you need at the start. You just want the 'main pieces to work together'. Once it starts to take shape, you've got to get much higher than that, testing is the way to do that.
When you’re starting out a project and “discovering” the structure of it, it makes very little sense to lock things in place, especially when manual testing is inexpensive.
Once you have more confidence in your structure as it grows you can start hardening it, reducing the amount of manual testing you do along the way.
People that have hard and fast rules around testing don’t appreciate the lifecycle of a project. Different times call for different approaches, and there are always trade offs. This is the art of software.
One thing I do religiously all the time is putting asserts everywhere. It's the only thing you can go crazy on. The rest is indeed always a balancing act.
So then I came along and said, "hey, why don't we have any unit testing?" and it turns out because it was pretty impossible to write unit tests with our code. So I refactored some code and gave a presentation on writing testable code - how the point of unit testing isn't just to have lots of unit tests, how it's more that it encourages writing testable code, and that the point of having testable code means that your codebase is then easier to change quickly.
I even showed a simple demonstration based off of four boolean parameters and some simple business logic, showing that if it were one function, you'd have to write 16 tests to test it exhaustively, but if you refactored and used mocking, you'd only have to write 12. That surprised people. Through that we reinforced some simple guidelines of how we'd like to separate our code, focusing on pure functions when possible, making layers mockable. We don't even have a need for a complicated dependency injection framework as long as we reduce the # of dependencies per layer.
Since that time we've separated our test suite into integration tests and unit tests, with instructions to rewrite integration tests to unit tests if possible. (Some integration tests are worthwhile, but most were just because unit tests were hard at that time.) We turned parallelism back on for the unit test suite. The unit tests aren't flaky, and now people are running the unit test suite in an infinite loop in their IDE. Over that time our codebase has gotten better structured, we have less interdependence and merge conflicts, morale has improved, velocity has gone up.
Anyway, according to this article it sounds like we've done basically the opposite of what we should have done.
And that by following those three principles, it kind of drives you to writing testable code. Because if you don't, you might have tests that are only small (simple integration tests), or only fast and reliable (testing unfactored code with lots of mocking) - and that the only way to do all three is by refactoring to write testable code that has good layer separation and therefore minimal mocking requirements.
There was stuff in there about how mutable state and concurrency leads to non-determinism and therefore unreliable tests, which is part of what justifies pushing towards pure functions that can be easily unit tested without mocking.
One of the things that distinguishes great engineers is that they make good judgment calls about how to apply technology or which direction to proceed. They understand pragmatism and balance. They understand not to get infatuated with new technologies but not to close their minds to them either. They understand not to dogmatically apply rules and best practices but not to undervalue them either. They understand the context of their decisions, for example sometimes code quality is more important and other times getting it built and shipped is more important.
As in life, good and bad decisions can be the key determiner of where you end up. You can employ a department full of skilled coders and make a few wrong decisions and your project could still end up a failure.
Some people never develop good engineering judgment. They always see questions as black and white, or they can't let go of chasing silver bullet solutions, etc.
Anyway, it's one thing to understand how to do unit tests. It's another thing to understand why you'd use them, what you can and can't get out of them, what the costs are, and take into account all that to make good decisions about how/where to use them.
Tell that to SQLite guy.
Try that with a networked application that takes user input though...
It will always break on the user's internet, because it's too diverse to predict.
Doesn't mean you can't have some networking unit tests, just that you shouldn't believe in them too much.
Edit: you said services. You thinking of server? I'm thinking of clients.
And it doesn't have to be a static mock. It's not too hard to inject a fuzzer in your mock service response, although that's probably left to a separate testing routine, and not part of your unit test setup. But if you have no mock for your network service, you can't fuzz it either.
SQLite is suitable for 100% coverage.
A lot of application code or workflow style code is hard to reach 100% coverage as they are rarely triggered.
That's an opportunity to describe, in code, what that path is supposed to do, and then make sure it does it.
Which shows 100% unit test coverage is not better than spending that time in other kinds of tests.
SELECT code execution from using SQlite - DEF CON 27 Conference
https://www.youtube.com/watch?v=HRbwkpnV1Rw
Don't get me wrong, it's a great product and I use it often, but 100% test coverage does NOT equal 100% safe.
It's often easier to just aim for 100% test coverage instead (with excluding some categories of files).
EDIT: I would not and did not start with 100% unit testing. But if there are ongoing culture wars and discussions didn't lead to a workable compromise, 100% test coverage worked for me and after some days test coverage was a non issue.
Property based testing and random values in unit tests find lots of bugs you didn't think of though.
Rather, you won't find bugs that you choose not to think of because you've let the "100%" number lull you into complacency, even though you know it's 100% of lines/branches, not 100% of inputs.
My problem in 40y of programming is still making bugs and those I make come from not thinking about edge cases or from wrong assumptions and not from being lulled into writing tests to meet a 100% number.
But personalities differ and if being lulled into security by writing towards a 100% number is a problem for you I would be careful, I totally agree here.
Perhaps I am wrong, and I would not start with a 'diktat' for obvious reasons.
As a manager, you didn't have discussions about the level of necessary code coverage? Would be interested on how you managed unit testing without 'dictat'. How it would fit into integration testing and explorative testing. What level did developers in your deparment usually find "adequat" ? If you considered it too low, how did you raise test coverage as a manager without defining a coverage level?
But in this case I think the cure might be worse than the disease. Tests for plumbing code often end up being brittle tests of methods getting called on mocks in the right order. People will notice that these require a lot of toil to keep them running as code changes while providing very little benefit in avoiding mistakes. People will rankle at being told that they must write these tests, which they can see are a waste of time.
I've done it both ways. I'm much happier with my work when I'm not trying to write tests that are tedious and don't seem to provide any value, in order to hit an arbitrary coverage metric. I suspect my teammates feel the same way, so on teams where I have input into the decision on this, I do not advocate for 100% coverage. It does make it harder to have the discussion of which tests should and shouldn't be written, but I think it's worth that cost.
Writing good testing code is harder than writing business code. Especially junior developers struggle with this, most often because many companies write not enough tests to learn writing good tests.
And if you're in an environment, where this is a non-issue I think thats great. Don't fix something that doesn't need to be fixed.
It has a side benefit that it forces devs to write testable code, which inclines them to reasonbly factored code.
I'd prefer <100% coverage plus discussions about what to test (and how) much more than working with a test suite built on the wrong incentives.
That's where the 'gaming' comes in.
The tests start just going through lines without hitting a single expect statement.
The ignore files start becoming battlegrounds in the PRs because people just exclude half the damn project.
We just have a simple rule... if you wrote code, you have to write coverage for it. If it breaks and your test doesn't catch the breakage, the bug fix goes back to you. Some people will ask "but what about what I'm working on now", you'll have to communicate that you feel your previous work was far more important.
this feels punitive, especially in the eyes of management. unless you're in a safety critical area where fully testing every code path is a hard requirement, people will eventually write bugs.
i'd rather work somewhere that recognizes defects occur and has a fast iterative process to push out new changes rather than one based on shame for having written a bug.
That is the fastest most iterative process we have found so far... as the expert on the original code, you are able to deliver the best outcome.
You're not being shamed for writing a bug, you're being shamed for not testing your code.
When my colleagues are knowledgable and open minded I would embrace every opportunity to have a good discussion.
Congratulations, now you have a war over which categories of files are excluded from the "100% test coverage" rule. ;)
What are you testing for?
This is critical because it basically gives you immediately what you should and should not test, and how. While mindless, dogmatic, metric oriented testing is a waste, testing with higher intent and purpose is extremely useful.
An example: test that something working on current vX also works on vA to vW, and when vZ is out, have the answer readily. Or that a biz feature fulfills the requirements. Or that someone not as well versed on intricate details of your piece of ownership will be confident in that piece still working after a simple fix when you’re on vacation. It can be one, some, but probably not all.
With that in mind, what to test, what doesn’t make sense to test, and what to test against becomes more clear: should I mock this? or should I run it against some staging environment? Should I perform (yikes? not!) manual testing?
The answers are highly dependent on the piece of code being tested.
Tests are here to help you answer a question, if you aren’t sure what the question is then your tests will miss the point.
I strive for working code. Sometimes I miss something in the TDD cycle and don’t have 100% and it is that which usually comes back to bite you.
I have never found 100% test coverage has bitten me, dogmatic or otherwise.
switch(type) {
case X:
...
case Y:
...
default:
throw InternalException("Unsupported type!");
}
Now if all goes well the default case will never be covered. At some point I thought "why have this code if it's not supposed to run; let's rewrite this so we can get 100% code coverage!", and I ended up with the following code: switch(type) {
case X:
...
default:
assert(type == Y);
...
}
Now we can get 100% code coverage... except the code is much worse. Instead of an easy-to-track down exception we now trigger either an assertion (debug) or weird undefined behaviour (release) when the "not supposed to happen" inevitably does happen because of e.g. new types being added that were not handled before.Is worse code worth getting 100% code coverage? In my eyes, absolutely not. I think good code + testing should be able to reach at least 90% typically, likely 95%, but 100% is often not possible without artificially forcing it and messing up your code and/or making it much harder to change your code later on.
You can be defensive to various degrees about assertions:
1. You can just use assert() to fail in Debug and do nothing in Release. 2. You can be more defensive and define your always_assert() to fail in Release as well. 3. You can double down on the UB with hints to the compiler and provide assume(), which explicitly compiles to UB when it's triggered in Release (using __builtin_unreachable() for example).
About the organization of the if statement: I agree that the former is better, I would use assert(false) though.
Throwing an exception here is basically free (just another switch case) and gives the user a semi-descriptive error message. When they then report that error message I can immediately find out what went wrong. Contrasting with a report about a segfault (with maybe a stacktrace), the former is significantly easier to debug and reason about.
assert_always would provide a similar report, of course. However, as we are writing a library, crashing is much worse than throwing an internal error. At worst an internal error means our library is no longer usable, whereas a crash means the host program goes down with it.
Better yet, omit that default case, so that in the future when you do add a new value to the enum, the compiler will warn you and force you to add a new case.
But I agree with your general thesis that it's just not worth getting to 100% coverage.
100% coverage of what exactly? Tests that go through all your lines of code without testing any of the logic, is useless. If you want to be thorough, you need to do mutation testing, which is a system that tests the quality your unit tests by mutating your logic (changing a > for >=, a + for -, etc) and then expects at least one test to fail. If no test fails, that piece of logic wasn't tested.
Without that, it's entirely possible your high code coverage doesn't actually test anything meaningful. Also, this sort of logic is exactly the kind of stuff you want to unit test. All the standard plumbing boilerplate code is not something that needs to be unit tested. The logic does.
I’d also like to add that if you contribute code to an open source project it is extremely beneficial to have iron-clad unit tests. Since there is so many devs it would be easy for someone to accidentally break something you fixed already.
But the original signature is just this:
public async Task<SolarTimes> GetSolarTimesAsync(DateTimeOffset date)
That introduces a lot of complexity:* The SolarCalculator needs to be able to work out its own location, so it needs a LocationProvider
* SolarCalculator needs to be IDiposable since it owns a LocationProvider
* The SolarCalculator will need more methods if it ever needs to calculate the times in a different location
* If fetching the location is slow, but the application needs to calculate times for multiple dates (eg to build up a table of times), then the SolarCalculator will need an method that takes in an array of dates to be efficient
But all that could be solved by making the function take all of the arguments it needs to return its value:
public SolarTimes GetSolarTimes(DateTimeOffset date, Location location)
No location provider needed, no IDiposable, just one efficient stand-alone method.Unit testing this is now just:
var calculator = new SolarCalculator();
var actual = calculator.GetSolarTimes(new Date(...), new Location(...));
var expected = new SolarTimes(...);
actual.Should().BeEquivalentTo(expected);
...so, perhaps the issue isn't that unit testing is a bad idea, but that code which is hard to use in a unit test might also be hard to use in a wider application? And perhaps the fix is to make the code easier to use?100% of the time, it was the right idea and the code became a lot better.
TL;DR: GetSolarTimes(Location, Date) is a unit-testable function.
Had some thought been put into writing with unit tests, there would be no problems with that example.
seems like you're all saying that
var actual = calculator.GetSolarTimes(new Date(...), new Location(...));
is JUST SIMPLER and better than public async Task<SolarTimes> GetSolarTimesAsync(DateTimeOffset date)
with an internal location provider as a dependency because it's easier to test.but i think that ignores the reason why DI containers were invented in the first place and assumes that the solar calculator is just a simple entrypoint-type application, rather than being a component in real application. You might have 20 layers of THING, somewhere inside which, this solar time calculator lives and is used... and you still have to get Location from SOMEWHERE to pass it into the calculator.
so what happens when Whatever uses the location provider to get the location and pass it along needs to be tested? and through how many layers of stack do you need to pass Location before you realize that every test of every intermediate layer needs to know about location, but only for the purpose of passing it along?
I think it's a more nuanced case than you're making it seem. Beyond some level of complexity in an application, it becomes simpler to co-locate dependencies where they're actually used.
If your code is broken down clearly into logic and plumbing, unit testing the logic becomes super easy. It allows you to construct software using blocks you have absolute confidence in. Unit testing plumbing is harder, and that's when integration testing shines.
The author's tests are overly complex. Instead of gleaning the actual value of this insight, which is that you're not cleanly separating your inputs and your outputs, the author concludes that unit tests are a waste of time.
Nope. Unit tests are a tool, but writing proper unit tests and understanding the value they give you is an art and a science. It requires experience and deliberate design.
It seems that most devs (me included) learn at school to write pure functions, which is great. Then they come to the industry and all of the sudden the "parseXml" function takes a ftp port as a parameter... ("be in my case the xml was on a ftp server!")
Why there is no CS course that explains this kind of stuff?
(And I am sure a bunch of other similar but differently-named concepts)
public SolarTimes GetSolarTimes(Location location, DateTimeOffset date)Refactoring is even worse. Refactoring after you've split something up into multiple parts and tested their interfaces in isolation is far more work. Any refactoring worth a damn changes the boundaries of abstractions. I frequently find myself throwing away all the unit tests after a significant refactoring; only integration tests outside the blast radius of the refactoring survive.
I find the same issue in throwing away tests when I'm writing small scale integration tests with junit. Usually I'm mocking out the DB and a few web service calls. So those tests become more volatile because their surface is exposed more. But smaller level, function and class level tests can have a really good ROI and they do push you design for testing which makes everything a bit better imo.
If you unit test all of the objects(Because their all public) then refactor the organisation of those objects then all your tests break. Since you've changed the way objects talk to each other, all your mock assumptions go out the window.
If you define a small public api of just a couple of entry points, which you unit test, you can change the organisation below the public api quite easily without breaking tests.
Where to define those public apis is a matter of skill working out what objects work well together as a cohesive unit.
They're not for catching unknown bugs, they're for safer updates.
If you change the implementation for a unit, a small piece of code, then the unit test doesn't change; it continues to test that the unit does what it's supposed to do, regardless of the implementation.
If you change what the units are, like in a major refactor, then it makes sense that you would need whole new unit tests. If you have a unit test that makes sure your sort function works and you change the implementation of your sort, your unit test will help. If you change your system so that you no longer need a sort, then that unit test is no longer useful.
I don't see why the fact that a unit test is limited in scope as to what it tests makes it useless.
I'm not arguing unit tests are useless.
Of course, you don't know ahead of time exactly which tests will catch bugs. But given finite time, if one category of test has a higher chance of catching bugs per time spent writing it, you should spend more time writing that kind of test.
Getting back to unit tests: if they frequently need to be rewritten as part of refactoring before they ever catch a bug, the expected value of that kind of test becomes a fraction of what it would be otherwise. It tips the scales in favor of a higher-level test that would catch the same bugs without needing rewrites.
That's like saying you shouldn't have installed fire alarms because you didn't wind up having a fire. Also, tests can both 1) help you write the code initially and 2) give a sense of security that the code is not failing in certain ways.
> It tips the scales in favor of a higher-level test that would catch the same bugs without needing rewrites.
Writing higher level tests that catch the same bugs as smaller, more focused tests is harder, likely super-linearly harder. In my experience, you get far more value for your time by combining unit, functional, system, and integration tests; rather than sticking to one type because you think it's best.
To go with the fire alarm analogy and exaggerate a little, it would work like this: you could attempt to install and maintain small disposable fire alarms in the refrigerator as well as every closet, drawer, and pillowcase. I'm not sure if these actually exist, but let's say they do. You then have to keep buying new ones since the internal batteries frequently run out. Or, you could deploy that type mainly in higher-value areas where they're particularly useful (near the stove), and otherwise put more time and money in complete room coverage from a few larger fire alarms that feature longer-lasting batteries. Given that you have an alarm for the bedroom as a whole, you absolutely shouldn't waste effort maintaining fire alarms in each pillowcase, and the reason is precisely that they won't ever be useful.
There are side benefits you mentioned to writing unit tests, of course, like helping you write the API initially. There are other ways to get a similar effect, though, and if those provide less benefit during refactoring but you still have to pay the cost of rewriting the tests, that also lowers their expected value.
To avoid misunderstanding, I also advocate a mixture of different types of tests. My comment is that based on the observation that unit tests depending on change-prone internal APIs tend to need more frequent rewrites, that fact should lower their expected value, and therefore affect how the mixture is allocated.
> unit tests depending on change-prone internal APIs
This in particular is worth highlighting. I tend to now write unit tests for things that are getting data from one place and passing it another, unless the code is complex enough that I'm worried it might not work or will be hard to maintain. And generally, I try to break out the testable part to a separate function (so it's get data + manipulate (testable) + pass data).
Good, this time you can get it right.
One of his examples from the article is injecting, IOC-style, the HttpClient instance into his LocationProvider class. He insists that this is a waste of time, and that the automated tests (if you have any at all), should be calling out to the remote service anyway. I can't disagree more! Hopefully you're configuring the automated tests to interact with a test/dev instance of the service and not the production instance (!). But what invariably happens is that the tests fail because the dev instance happened to be down when they ran. And they take a long time to run anyway, so everybody stops running them since they don't tell you anything useful anyway. This is even worse when the remote service is not a web service but a database: now you have to insert some rows before you run the test and then remember to delete them... and hopefully nobody else is running the same test at the same time! To be useful in any way, automated tests must be decoupled from external services, which means mocking, which means some level of IOC.
On the other hand, he also introduces the example of SolarCalculator mocking LocationProvider. I agree that that level of isolation is overkill and will unapologetically write my own SolarCalculator unit test to invoke a "real" LocationProvider with a mocked-out HttpClient, and I'll still call it a unit test. (On the other hand, the refactored designed with the ILocationProvider really is better anyway).
So I think the reason people argue about this is because they can't really agree on what constitutes a unit test. I'd rather step back from what is and isn't a unit test and focus on what I want out of a unit test: I want it to be fast, and I want it to be specific. If it fails, it failed because there's a problem with the code, and it should be very clear exactly what failed where. A bit of indirection to permit this is always worthwhile.
Can you expand more on this? I think this is where the author would disagree.
E.g., how is the code easier to reason about or refactor having introduced a location service interface that has only one one implementation?
Striving to make your code testable is almost always worth it. Someone might ask this guy to add some error handling to his code for example. :) Then he will find out that by writing code, however simple, that a "works on my machine" I.e. is proven to work in a single happy path context is painful to change. Writing code that runs in multiple contexts (composed as an app or decomposed for testing) is intrinsically more easy to work with and change.
or productivity.
I've been on projects that focused almost exclusively on unit tests and on projects that focused almost exclusively on integration tests. The latter were far better at shipping actually working code, because most of the interesting problems occur at the boundaries between components. Testing each piece with layer after layer of mocks won't address those problems. Yay, module A always produces a correct number in pounds under all conditions. Yay, module B always does the right thing given a number in kilograms. Let's put them together and assume they work! Real life examples are seldom this obvious, but they're not far off. Also note that the prevalence of these integration bugs increases as the code becomes more properly modular and especially as it becomes distributed.
I firmly believe that integration tests with fault injection are better than unit tests with mocks for validating the current code. That doesn't mean one shouldn't write unit tests, but one should limit the time/effort spent refactoring or creating mocks for the sole purpose of supporting them. Otherwise, the time saved by fixing real problems more efficiently - a real benefit, I wouldn't deny - is outweighed by the time lost chasing phantoms.
Unit tests protect you against current mistakes. They're tied to the exact implementation.
"Right now my function X should call Y on it's dependency Z before it calls A on it's dependency B. I know that my method should do this, because this is how I designed it now. Let me write a test and expect exactly that."
Integration and unit tests will tell you whether in the future your code will still work when you refactor.
"Okay, we rewrote the whole class containing the function. Does running my thing still end up writing ABC into that output file?"
Otherwise I agree with you mostly.
If unit tests are tied to an exact implementation, they''ll fail on correct behavior and that's definitely wrong. It shouldn't matter whether X calls Z:Y or B:A first, whether it calls them at all, whether it calls them multiple times, whether it calls them differently. All that matters is that it gets the correct answer and/or has the same final effect.
Unit tests should be based on a module's contract, not its implementation. This is in fact exactly what's wrong with most unit tests, that they over-specify what code (and all of its transitive dependencies) must do to pass, while by their nature leaving real problems at module interfaces out of scope.
b) Even if you have an output, it's dependent on more complex input of arbitrary types.
Assume that there's a method that returns an input based on summing the output of a method call of it's abstract dependencies.
To do dogmatically correct unit testing you'd pass those 2 mocked dependencies, and have those methods return the values when the right method is called on them.
Then you'd assert that B was called on A, that D was called on C, and that the method under test returns the sum of those returns.
As soon as you move into passing implementations of those 2 dependencies, to anyone dogmatic you're doing integration testing.
Even if the tester isn't being dogmatic, in a lot of cases these inputs are complex enough that building enough actual inputs that are consistent and realistic to cover all the cases is prohibitively costly, so they opt for mocks.
Now, suddenly you just have more code to maintain when making changes, but you feel good about yourself.
O -> int
Your unit test is concerned with narrowing the interface above to: O -> int // of specific value based on dependencies
If Os only dependencies are A and C, this can be rewritten to: A -> C -> int // of specific value
Of course if we assume both A and C, themselves, have dependencies we can recursively rewrite the above until we have a very long interface, but instead you have opted to mock (M) them: M(A) -> M(C) -> int // of specific value
You then take it a step further and mock the method calls on each to return a specific value: M(A) -> int
M(B) -> int
becomes: M(A) -> 3
M(B) -> 5
Okay. Now we can rewrite our interface to: 3 -> 5 -> int // of specific number
and our test to: 3 -> 5 -> 8
and make our assertion that the result is indeed the sum of the inputs (not to mention the ridiculous assertions that specific methods were called within the implementation). Yikes... No wonder OOP gets a bad wrap. All that for what amounts to a `sum` function.The designer of the above monstrosity could learn a lot from the phrase "imperative shell, functional core". It sounds like dogma until you are knee deep in trying to test the middle of a large object graph!
It's integration testing that validates that all your units still combine (integrate) into a working end product. That's not about testing your implementation nor your internal interfaces, that's about testing your program's inputs and outputs.
All tests protect the programmer against future mistakes. All tests are a protection against regressions.
But yes agreed, integration tests absolutely carry much more value than any unit tests might. Specifically because units tests tend to target things that are essentially implementation details.
The only time I'd say unit tests carry any value is if they're testing some especially important piece of business logic e.g. some critical computation. Otherwise, integration tests rank the highest in the teams I lead.
I don't have anything against objects, per se, but I think they tend to make unit testing much more difficult to accomplish. The closer your code resembles pure functions, the easier it is to do dependency injection and unit testing.
It's where you need to handle mutable state with objects that things get trickier.
Unfortunately, these are exactly the places where you most need tests.
I'm a big fan of constantly returning things rather than holding state in objects, for specifically this reason.
Plus the same problem arises with modules instead of objects, which traditionally are even harder to customize.
You can get pretty far with good abstractions and dependency injection. Go's io::Reader and io::Writer interfaces are a great example of this. The resulting functions aren't pure in a technical sense, but they're pretty easy to unit test none the less.
> Plus the same problem arises with modules instead of objects, which traditionally are even harder to customize.
Maybe you could elaborate. I really don't understand what you mean here.
From what I understand, modules just scope names, they don't maintain state. I don't see how they have the same problems as objects.
Which goes back to the article's point of having to write code that is unit test friendly.
Now architecture decisions have to integrate interfaces that wouldn't be needed otherwise.
> Maybe you could elaborate. I really don't understand what you mean here.
Modules keep state via global variables, module private functions and the surface control that they might expose via public API for the module.
Additionally on languages that support them, they can be made available as binary only libraries.
> Now architecture decisions have to integrate interfaces that wouldn't be needed otherwise.
You're not wrong.
But in the context of functions, that doesn't seem to me to be particularly onerous. If the worst I'm forced to do is change the type of my parameters to an interface instead of a concrete type, that seems like a pretty small price to pay for easy testability. Certainly a much smaller price than the examples in the article.
Imagine doing unit tests for a C application, where modules == translation unit/static/dynamic library, thus you can only do black box testing.
Now one needs to clutter it with function pointers everywhere, or start faking interfaces with structs, just for the benefit of unit tests.
And with static/dynamic libraries than one might need to start injecting symbols into the linker to redirect calls into mocking functions.
All just to keep QA dashboards green.
The fact that it makes unit testing easier is just icing on the cake.
Mainly due to the linking hacks and low level debugging sessions required to mock all necessary calls.
Plus that was just an example, there are plenty of languages with modules and binary libraries.
The libraries dependencies should all be indirected through whatever context struct you pass to all your calls.
Sadly not all code is great.
You probably want it rigged up to your own logger instead of just blindly writing to stdout. You probably want the library's allocations tagged somehow on the heap so you can track down memory leaks. You probably don't want it doing IO directly, because of how many different way there are to do IO.
It's all more a function of how incredibly varied c envs are, than design for testability. It just happens to be very testable as an aside.
Keep the impure code and the pure code separated.
Foo foo = new Foo(mock(Bar.class))
foo.a(x,y,z)
Not sure what the issue is there really.In my experience mock objects can be brittle. A few sprinkled in judiciously can be ok, but once the density gets high enough, it starts to feel like the test becomes decoupled from the actual code it's supposed to test.
If the only thing you inject is data, can we still call that "dependency injection"?
I suppose that's a philosophical question.
Probably the 2 most common functions I 'dependency inject' are rand() and time.now(). I feel like they count, but you might not.
I also tend to avoid HOF when I can instead pass data around explicitly.
Everyone started writing unit tests, and the code broke less. Developers became more confident in deploying, and eventually most PRs looked roughly the same: 10-20 line diff on the top, unit tests on the bottom. If there were no tests, the reviewer asked for tests. It became a fun and safe project to work on, rather than something we all feared might break at any moment.
I've since started insisting on having them as well, especially when I'm using dynamically typed languages. A lot of the tests I write in Python for example are already covered in a language like Go just by having the type system.
So we started adding unit tests. Utesting code that wasn't written for utests is painful: you often need to choose between refactoring or just patching the hell out of it. The latter is highly undesirable, since it leads to verbose tests, failures when you move a module, and the inability to do blackbox testing.
But utests encourage our new code to be clean and readable. We've found that functional programming is much easier to test than object-oriented, and is easier for engineers to grok. We just sprinkle a little dependency injection and the whole thing works nicely.
Itests have their place, but utests lead to faster feedback and more readable code.
What's a better term than "dependency injection"? What should I call an argument whose default is always used in production code, but is there to make passing a mock easy? I'm not trying to be snide -- I'm genuinely curious.
Unit tests are an easy path to fall down, because they're clearly easier to setup, to write for, require less effort to maintain, execute more quickly.
But you don't realise their significant downside until after you attempt a major refactor - you begin to see that unit tests are testing at the layer that changes the most anyway.
1) The use of unit tests as the exclusive automated test type. ie; No functional tests, integration, etc.
2) Test doubles for most or every dependency, even purely functional dependencies like math libraries.
3) Not using the appropriate kind of test double for the test at hand. (Dummies vs Fakes, vs Spies, vs Stubs, vs Mocks)
4) The overuse of mocking libraries.
Mocking libraries have their place, but in opinion, are used approximately a hundred, perhaps even a thousand times more often than they should be. I use them to create test doubles in exactly three scenarios:
1) A dependency that does not have an interface, usually a third party library. This usually happens in one place only, and is used for writing the wrapper code test.
2) A dependency that has an incredibly large interface and/or dependency graph where building a set of stubs or spies is simply not worth the effort.
3) I want to test weird edge cases that's not available any other way, such as theoretically unreachable code.
These should not be the majority of your unit tests!
It feels like the industry has blindly pushed for unit testing everything and 80% or more code coverage as the gold standard.
I’ve given up arguing about the cost/benefit of unit tests at work. I feel that the software the teams I’ve worked on over the past couple of decades still produce about as many bugs as before unit testing came along. I’m not building pace makers or aviation software, mostly LOB applications.
Unit tests provide a false sense of security (especially to management.) Yes sometimes they help catch refactoring bugs, but at what cost?
For a startup with a small team and few customers building an MVP? Unit testing is overrated.
For a company with 50 engineers in 10 teams building a product, that moved $500,000/day in revenue? Unit testing could or could not be overrated.
For a company with 1,000 engineers working in the same repo, shipping a product that moves $50M in revenue per day? Unit testing is most likely underrated - and essential.
You cannot ignore how the organization works, and the cost of a defect that a unit test could have caught. I happen to work at the third type of organization, and while unit tests might not be the most efficient type of safety net, it is a very big one. We have other types of testing layers on top of unit: integration and E2E tests as well.
Also, one more fallacy in the article: "If we look back, it’s clear that high-level testing was tough in 2000, it probably still was in 2009, but it’s 2020 outside and we are, in fact, living in the future. Advancements in technology and software design have made it a much less significant issue than it once was."
This is not true everywhere. High-level / E2E testing on native mobile applications in 2020 is just as bad as it was on the web in 2009.
You are right, but it still doesn't mean aiming for high coverage. In the big company case you'll want to cover the interfaces and dependencies and less of your team's code.
I know that part of this will fall under "integration" but definitions are sneaky.
I think unit testing definitely has its place. I use it a lot; but I have learned to moderate my reliance on unit testing.
I tend to prefer test harnesses and manual (or automation-assisted) testing. I've sometimes written my own unit-testing frameworks, because the "canned" variants didn't give me what I needed.
The term "unit testing" is quite old. It seems to mean something different, these days, from what it used to mean.
As far as I'm concerned "not testing my software" is out of the question.
For most of my projects, the testing code is vastly greater than the actual product code.
Refactoring usually changes interfaces. Things are factored differently. The clue is in the name.
The higher up the stack your test is, the more insurance it gives you for refactoring. The lower downs the stack it is, the more likely it is to be thrown away or heavily rewritten after refactoring.
No, refactoring shouldn't change public interfaces. The very definition of refactoring is rewriting code without changing interfaces.
> Things are factored differently.
internally
> In computer programming and software design, code refactoring is the process of restructuring existing computer code—changing the factoring—without changing its external behavior.
https://en.wikipedia.org/wiki/Code_refactoring
You got the definition of refactoring wrong, please get it right, it's important. If you are breaking a public API, you are not refactoring anything.
Any piece of code meant to be private shouldn't be unit tested at all, only the behavior of a public interface.
Now internally you might call a third party lib, but that third party lib is then a separate "unit" itself.
I don't like the term "unit" because it's yet another word that is easily misunderstood and lost its original meaning with time.
unit testing should really mean "public interface behavioural integrity testing" or something like that.
If users should not access it directly, you don't have to unit test it directly, you test them indirectly via the public interface.
The term "refactoring" is commonly used in a way which includes changing of public interfaces. Random article, which even cites a book, agrees: https://thoughtbot.com/blog/lets-not-misuse-refactoring
If you're developing a library, then refactoring shouldn't change public interfaces. If you're developing an application and you own all the code paths to the code, then refactoring could change public interfaces, as the external behavior here would be the UI.
If you literally agree with the "refactoring shouldn't change public interfaces" then we need a new word for "code improvement which doesn't change external behavior, which can mean UI", which is the more commonly needed term.
And then perhaps we could agree that "code improvement often changes public interfaces" and how this relates to unit tests.
They serve as a form of living documentation for the code and help increase velocity in a build under the right conditions.
For example, you do need to know a certain function does what you think it does because the rest of the system isn't even in place yet. You might have to approach this from the outside, via integration, but the speed of doing this is quite slow. Versus a unit test.
This is not to mention refactoring!
The piece seems to be more about the value of mocking and how far to go with isolation. Which is a slightly tangential issue. I agree that in particular styles of object orientated programming this becomes absurd.
This article, which is linked provides a more convincing case, on the grounds that unit testing foregrounds the system as software as opposed to the software as a useful and functional thing in the world, meeting user needs.
However, it omits the being about to _react_ to user needs is actually a central party of agility. Unit testing allows maximum reactivity to changing requirements without regressions in the code and makes this code navigable. Changing customer needs means your carefully built functional tests are going to be just as useless and rotted. This has been the case with large test suites of functional tests - say in something like Cucumber - I've seen. Better test at a lower level, which moves slightly less rapidly.
The issue here is developers don't have a sense for economics. Diminished returns, marginal utility, and opportunity cost should be studied by ALL.
Most of my experience is with React, and the majority of react devs I've seen don't unit test their code, because they don't write enough pure testable functions. As a result the community has leaned heavily on React Testing Library which is AMAZING for integration testing. Instead of checking that your function returns the right value, people will mount the component and then check to see that the right value is displayed in the rendered DOM. This obviously works, but writing the logic as a pure function and unit testing that function with a lot of different inputs gives you much more confidence.
Unit tests are very useful but we somehow landed ourselves in a place where we have lots of line coverage in our unit tests but little confidence that the code actually works.
There are at least 3 different applications at work where they unit tests are green while the code is broken or red when it's not. The cause is almost always the mocks. They either presume too much, making then fragile or they are flat out incorrect. Despite having a lot of test coverage there is little confidence for the developer that their change is correct.
In a sense the writers of the tests were "doing it wrong". The burst of articles on unit tests and their failure modes are a reaction to the prevalence of this in our industry.
I've seen unit tests as documentation cause problems a few times. Like, maybe if you've got some sort of DSL where it's actually obvious what behavior is expected. However, more often it's 10-50 lines of setup / mock code and then some number of asserts and you're left trying to decide what the point is (ops, it turns out something got misinterpreted and the test is actually nonsense).
Finally, it seems like using the type system to design code where illegal states are not representable is starting to make some headway. Additionally, we're seeing increasingly powerful type systems make it into industry acceptable languages.
* More closely encodes the behavior you want the code to have
* More likely to find weird edge cases where some input does something unexpected
* Less work to get greater coverage
Arguing against unit testing is like arguing against type-safety (and, usually, anti-unit-testing people are anti-type-safety people, too). The presumption always seems to be that if it doesn't solve every problem, it's unnecessarily slowing things down.
I already have to write unit tests so why would I bother with types they don't add any real value.
Both camps are wrong. Unit tests are an unambiguous good. Types are also an unambiguous good. Both have some rather common failure modes though and guarding against those failure modes is useful.
Unit tests of purely functional code where you only need to provide an input and validate an output provide tremendous value. The unit test can treat the code as a black box and as a result the unit test is robust and resilient to changes in the black box while verifying that the box still produces the correct answer. Unit tests of code with hidden dependencies that need to be mocked require a lot more care to construct properly. Mock scripting frameworks encourage a number of bad habits. Things like "How many times did this method get called". Or "Always provide the same answer when this method get's called." The result is a hundred reimplementations of the same interface that are at best correct for the current version of the code they are testing and at worst completely incorrect reimplementations of whatever they are mocking. They all need to be kept in sync and maintained over time.
A shared in memory Fake will in general provide more value and be less fragile over time while also ensuring that your tests are actually testing the code and not the particular script you defined for the mock.
This is the opposite of my experience. Most of the anti-unit testing people I've talked to are very much pro-type people. I wonder if anyone has done any studies that shows what the actual numbers look like.
> Arguing against unit testing is like arguing against type-safety
I disagree with this. Types (at least when they've been built on top of an actual logic) have the benefit of real costs and benefits. You can show what programs you are unable to write and you can show (mathematically) that certain failures won't happen.
Unit tests on the other hand are much more hand wavy. You can show that some refactors seem easier, but you can't prove it without a lot of data that has to be collected on a project by project basis. I'm not saying that I don't want unit tests if I'm doing a non-trivial refactor, but I am saying that that desire is more of an intuition thing. It's not like I can make any proofs around it like I can do with a type system.
There is also a corporate culture of tests that mandates bloated test frameworks, which leads to developers spending more time on writing tests than writing functional code.
If you start striving for over 100% unit test coverage, then you'll be testing if the IDE pre-generated setters and getters actually set and get the value. This adds zero value to the codebase and you'll be testing things that, if they fail, will break half the world anyway.
Unit Tests are for algorithms, stuff that does something complex and not immediately obvious. Preferably deterministic, every time X goes in, X+Y comes out type of stuff.
Most tests should be either integration tests, testing the interfaces between different parts of the software or automated tests pretending to be the user, made with Robot Framework or something similar.
- 33% unit tests (of well-defined units, such as functions that compute some subtle math logic, or read/write marshalling)
- 33% integration tests (requiring multiple components to work correctly to achieve the result)
- 33% business tests (automatically steering the entire application like a user would and testing the result like a user would see it)
The main goal of the systems I work on is to provide technical documentation of complex industrial processes. If things break, it can be pricey and/or dangerous. Having good information is a must.
However, if a user sees that something is off or just plain missing in the sometimes 30+ year old documentation, the easiest way to deal with it is to make a note of it and adapt to it for his or her work. Reporting the problem back in order to get it fixed is....difficult. There probably is a process for it, you probably don't have an account where you can log the time spent on it.
Having a quick and low-threshold way to report problems would be of enormous value in the long run.
Test coverage targets make very little sense for a statically typed language like C#, a fools errand almost. In dynamically typed languages, it's hard not to almost hit 100% if you're doing some honest TDD. Just for example.
- when to use unit-tests ?
- when are other tests more useful ?
- which bugs can unit-tests find ?
- in what situations is changing code to make it testable some-times beneficial ?
- when is it a complete waste of time to make code unit-testable ?
Things to consider when answering these questions above:
- what is the customer impact if there is bug ?
- what is the developer impact if there is bug ?
- what is the developer impact writing tests ?
- how quickly can we fix it and release a new version ?
- how quickly can we fix the hotfix in case we messed up ?
- how can we find out if there is a bug ?
[...]
I think it is great that he starts thinking about the value of unit-tests and testing in general in his particular context.
But i think it helps keeping an open mind knowing that some approaches work better than others depending on the context.
And there-in lies the problem. Remove the idea that the unit is a single method/function.
I subscribe to the idea that a unit is a unit of functionality. Nothing to do with the code implementation.
Only mock where you're reaching out outside of your codebase (filesystem, network, operating system (time, for ex.))
You can still do unit tests for individual functions when you need to work on a complicated algorithm, but those functions should have no dependencies or side effects - pass in all the data you need
At small companies we have absolutely been able to get away with 0 unit tests while maintaining agility - being able to do major reactors quickly even when working in dynamic languages, while maintaining a high level of quality, even when operating at reasonably large scale. The key is clear, well written code and strong ownership from senior engineers who have a deep understanding of the code they own. On the other hand, at large companies, extensive unit testing has been invaluable. Code bases are older, ownership changes hands frequently, new engineers join all the time, old ones move on to other projects, dozens of teams are calling each others code, refactoring is done by people who had no hand in writing the code in the first place, dependencies are higher and harder to track down completely, and it's not realistic to cover all important functionality with integration tests. Engineers often must rely on unit tests to prevent others from breaking their code, and to ensure that they are not breaking someone else's. Yes, unit tests can be highly problematic and costly to maintain, and they do add friction and time to initial development, but in these scenarios the benefits outweigh the costs considerably.
The standard reference on this is http://www.growing-object-oriented-software.com/ .
Might as well retitle this "Why I don't like the Interface Segregation Principle".
I've seen far too many cases where added complexity has killed the project's time budget because things just take so much longer to do. You need to maintain a healthy balance and actually evaluate if writing this particular system within the project in a way where we can "swap it out" later on is actually a use case.
Every project is unique and "unit testing all the things" might not be the solution for your project. Where I work we shifted our mentality when we highlighted this problem to only unit test bugs and write less decoupled code and in bigger services that are more mission critical we do integration testing instead. This works very well for us but as highlighted earlier, every project is unique and you should experiment what works for your particular team because what works for others might not work for you.
I believe unit and integration tests, apart from checking whether code does what you think it says, it serves both as a sort of executable documentation and, most importantly, it highlights development intent. If you TDD your code iteratively, reflecting not only the unit in your code, but the intent in your tests, you get a much more healthy testing base.
Code coverage is not really a good metric. Intent coverage is more interesting, but a whole lot more subjective and elusive. All in all, tests should be written not with your own self in mind, but with whoever might come later to maintain your code.
No, a unit is not a function, or a method, or a class, or a file.
A unit is a clearly separated, non-trivial software component with a minimal and stable interface to the rest of the software system.
A unit is a good if it requires little to no mocking to test and a very small number of messages to test. It should also do something non-trivial - think small library size, not class size.
The unit should probably have a README which explains the small interface and the small number of necessary dependencies. The unit tests should largely treat the unit as a black box and be based on its promises in the README.
If module A is very expensive to fix and has a high probability of failing in production, well of course you want to have it rock solid and should be thoroughly tested.
You could build a priority index, with something like :
priority for testing module A = cost of fixing A x probability A fails
The other problem is getting management onboard and demonstrating a ROI justifying the time spent building these tests and using them. I personally failed at that and still am trying to figure out how to get them to understand the benefits. I've progressed, but it's one heck of an uphill battle for me.
* a 'journal' - at the date this test was written, this is how we expect the system to behave
* an ELI5 - if I'm trying to use this method, why am I passing all these complex objects?
Unit tests declare expected behaviour, and should make the developer think about their methods.
For example, why pass complex objects to just to print a string or a count or similar? And why pass those objects? could a method be generalised to take a lambda, or an interface instead? Could the method be pure? And so on.
Unit testing isn't overrated, just a bit misunderstood.
Heh I feel like this is the crux of the issue. Because there's no standardized definition of what a unit is, people sometimes tend to choose the wrong unit to test.
I personally believe it is usually wrong to test a single function or method in a class. I tend to test the behavior of a whole class at the same time. Testing each individual method is too white-box, and makes your testing code too coupled to the implementation. Basically, don't test internal details (unless the internals are very complicated), just test the externally visible behavior, which is usually presented as a whole class.
I also don't agree with always mocking out dependencies. If your "dependency" is just an instance of a different class, then just use the real deal. Sure, now your definition of a unit now encompasses not just your code, but your dependency's code, but if your dependency's code affects the externally visible behavior of your code, it remains your responsibility if your dependency changes things and breaks your code. That's what abstraction means: you present an interface, your client doesn't need to care about how it's implemented, and what dependency your code needs.
The second time, less often, when tests are useful is when making large changes with broad implications, and quickly verifying that everything still works and that you haven't overlooked some subsystem.
The least frequently useful application of tests is regression tests. 99 out of 100 regressions only happen once. If you add a regression test, it won't happen again (unless you overlooked something), but it was unlikely to happen again anyway.
Generally I write the first category of tests when I feel that it would be useful to solving the problem I'm working on. Then, when I finish the code, I commit the test, because why not? It's written. It'll never fail again, but hey. This creates a reasonably managable collection of tests which, more often than not, test the more complex (and therefore more fragile) parts of the codebase, and are a decent representative sample of all of the subsystems. This provides sufficient test coverage to support the second case. The third case is so rare that it can be addressed on a case-by-case basis.
The most stable software is software which doesn't change. Keep your scope small and your complexity low, and don't be afraid to mark a finish line. This is more effective than exhaustive automated testing.
Someone is pretty inexperienced by making that claim. Unit tests help to isolate the expectations and verification to very small levels.
> While these changes may seem as an improvement to some, it’s important to point out that the interfaces we’ve defined serve no practical purpose other than making unit testing possible.
No it simplified the responsibility of the class. It also simplified your tests as well. Now you don't have to have tests that test many different scenarios.
> Note that although ILocationProvider exposes two different methods, from the contract perspective we have no way of knowing which one actually gets called.
Tests don't verify how you use or call other classes/methods. You can do verification if you want via mocks.
> Unit tests have a limited purpose
Yes, as does integration tests, functional tests, system tests. Etc. You shouldn't be trying to do unit level testing via an integration test.
> For example, does it make sense to unit test a method that calculates solar times using a long and complicated mathematical algorithm? Most likely, yes.
It does if you want to verify that the functionality is setting up the request correctly.
> Does it make sense to unit test a method that sends a request to a REST API to get geographical coordinates? Most likely, not.
That's not a unit test. That's an integration test(if you use a mock) or a functional test if you want to end a live end point.
> Unit tests lead to more complicated design
It highlights that the original design was complicated or the methods had side effects (which you should avoid). In his own example he seperated the resources from the functionality.
> Unit tests are expensive
No they're not. Mocks do not belong in a unit test. That's for integration tests. If you have them there you have issues. They're cheap to write and quick to run.
> Unit tests rely on implementation details
It's all about how you write you code, if you're trying to lump everything together, that's what you get.
> Unit tests don’t exercise user behavior
Correct, unit tests don't. Functional/feature or above do.
I stopped reading after this. It feels like the guy is just trying to argue that he doesn't like testing.
I find it to be some mixture of importance and complexity, while always balancing against the single responsibility principle. A simple `average` function might be trivial but if it's important to your business logic you probably want to test that separately, even if it is nested within a "unit".
I find that following some iteration of "functional core imperative shell" helps here as it helps keep your core business logic to being data transformations and transformations on data are easy to test and their concerns are easy to reason about.
This then helps me reason about what is "implementation" and what is a contract/interface which should be tested robustly.
I guess really the art of writing just enough unit tests is to identify the seams and boundaries of your abstractions in your codebase, and potentially accepting that business seams in your codebase may be different from the seams of your actual software domain -- the latter being possibly more granular.
That said, the moment the abstraction is in a clean state and consistently passing all tests and has been integrated well with existing logic, the unit tests are deprecated as far as I am concerned. I wont ever explicitly delete them, but I recognize that the unit tests are potentially just as flawed as what they are testing (the same developer wrote them after all), and coming back into that abstraction after 6+ months elapses means I'd probably just have to rewrite tests from scratch as a mental exercise and in order to restore my own sanity.
I can recall at least one occasion where I wasted 2 days chasing down a failing unit test only to find out the test itself was flawed - in the worst way, random pass/fail based on a race condition between multiple threads that were part of the testing code. I think that's the biggest danger with unit testing. Other developers assuming the tests you wrote are foolproof and sending them on pointless errands.
Now I'm not saying that they are unnecessary and i definitely believe they are needed but i do think they are overrated.
I've worked on many buggy systems that had very good unit test coverage. It was only with sufficient integration testing that were able to prevent constant regressions.
Of course if your integration tests are lacking you'll have problems but this does not tell you anything about the value of unit tests, just that integration tests are obviously valuable.
I see that as an useful feature: This shows you the cost of refactoring and the fact that this cost includes re-testing everything.
Very long article that leads to a trivial conclusion that unit testing is not a substitute to higher level testing, and carries the overhead effort.
It seems to me the author is addressing cases where unit test coverage is blindly used as dev project metric. PMs are often not familiar with code internals, but they need some assurances that project is on-track. Thus trying to collect insights from whatever output provided by automated tools.
Unit tests can have 100% coverage, but still testing the wrong thing. A common misuse is testing to the actual code, instead of testing the expected behavior.
Unit tests is a developer's tool, not PMs metric. On the other hand Acceptance tests should be the common ground, which may be the Integration tests or a whole separate suite altogether tailored to user requirements.
When coding a function, one needs to test preconditions and assumptions prior to following the main logic. Unit tests serve the very same purpose by enforcing assumptions on Unit level at the granularity that makes sense. Noone needs to test trivial getters and setters, but one needs to ensure that objects remain in valid/known states. Thus Unit tests untie developer's hands to recraft the "unit" without fear that the pearls of correctness would get lost.
Paradoxycally, Unit testing is a productivity tool, just as diagraming, or ... sketching prior to coding. Some developers can maintain a perfect mental picture of their code without any need for such tools. Some devs see design and implementations right away. If I were to inherit their codebase, I'd rather see their assumptions validated, so I don't need to do the coverage by eyeballing the code.
Tests that validate the interface of the big black box and assert that it does the right thing (integration tests). I.e assert the behavior customers depend on.
That black box is made up of many smaller black boxes that talk to each other. So you test the boundaries of each of those smaller black boxes. (Unit tests)
The ROI for testing the big black box is much higher. But when something fails it’s hard to know what exactly caused it and how to fix it. If you have a decent number of tests for the smaller black boxes then you know what box needs to be fixed.
But all good tests are black box tests. I.e they don’t test the internals but the interface of using that box.
A well designed system is made up of small boxes that do one thing very well and work with other boxes. Once you’ve tested them you can forget about them and work on higher abstractions. They will continue doing what they promised.
The argument that it's more productive to test modules through the application only applies to private, dedicated modules that will only ever be invoked in support of the use cases arising in the application.
Basically, it boils down to whether or not the piece of code is an independent product (where "product" could refer to something with "customers" internal to the organization).
In engineering, all building blocks that are separate products to be integrated into other products are rigorously tested on their own, whether they are integrated circuits, or steel cables or whatever.
https://github.com/boothby/dissert
I like assertions. They're a really good alternative to unit tests because they can be used in a real environment without the need for maintenance-intensive mocking etc. But assertions have a significant performance cost, especially when they involve consistency checks on large datastructures.
So, a pattern that I've found useful is heavyweight asserts that can be enabled / disabled through external means, in conjunction with high-coverage integration testing. This is really easy in some languages (c/c++, for example, using the preprocessor) and can be more fragile in languages like python (hence the plug above -- which isn't the best way, but is a way, to achieve this 'best of both worlds' testing).
I've been saying this for years.
The whole point of testing is make sure you aren't breaking something when you add a feature/refactor/delete old code. Its purpose is to speed up development. Excessive unit testing just creates a brittle test suite, and adds more work without much benefit. It slows you down.
As a Rails dude focused on startups, iterating rapidly and what the user sees is what you care about. Therefore, I focus on integration tests that run the whole stack. That lets me mess with the implementation code without re-writing the test suite. At the same time, it provides regression protection and a good place to start troubleshooting. Plus, Rails already has tests for the "plumbing".
Tests should serve the developer, and speed up the iterative process, not add work to the project b/c of a dogmatic adherence to TDD or unit test all the things design.
Just my (unpopular) opinion.
I do a lot of dev in Ruby and testing there is super easy and powerful. Say what you will about monkey patching, open classes, and reflection, but it does make it very easy to write great testing libraries. I'd argue that testing libraries in Ruby lead to better design (e.g. using more methods, making things single purpose, thinking about the interface first can lead to better naming etc)
That said, I don't test as much as I used to. Unit testing is really good for testing algorithms (e.g. transform this complex JSON; rebalance load). Outside of that, some light end-to-end testing will catch most other things, and you can use staging and gradual rollouts to derisk bigger changes
On the other hand, unit testing does present some benefits. Consider a function that cannot be easily unit tested. This tells you that it's probably too complex and would benefit by refactoring into more manageable parts. Also, business logic unit tests are beneficial.
Unit testing should be as easy as writing a few lines of code, just to ensure that an interface behaves as expected, but it's often not the case because well many languages make it incredibly difficult, for various reasons, or that you end up writing walls of code (thus behavior ironically) with mocks, stubs, fakes and what not, just to test a single method...
"Testing experience" should absolutely be a core concern for any modern language. More generally the notion of "developer experience" should be a core design aspect in any new language and not be delegated to third party tools.
What is the language with the best testing experience?
I support an existing reporting system with several hundred tables and thousands of interdependent stored procedures. This system has a fairly often-used "global report" which touches nearly 90% of the tables. We run this user report after all changes to the database with a few swaps in parameters (history vs current etc). If this report and a few others "pass" (results are the same as previous reports), we know that the changes are safe. When the reports show differences, we confirm that the changes were intended. If not, something broke.
It was Kent Becks brain child. The first NUnit framework was SUnit.
What was a unit? Kent was never absolute about this, I always felt because as a consultant he wanted to peddle the theory far and wide. But the early examples all had a pretty strong trend towards "a unit is an object." Nowaday "what is an object" is pretty loose; I just finished some Dart tutorials where an object is a way of "organizing our code into smaller reusable pieces" which ironically, I did in Fortran77 with common blocks and well factored files. But in the Smalltalk world, which was "objects all the way down", an object was small amounts of imperative behavior bound to data, where computational results were achieved via an approximation of the way cellular biology solves problems: lots of little glumps of data that achieve a larger result by sending messages to each other.
This process of turning behavior and algorithms into things, or reification, was sometimes easy and sometimes hard. An object for Point, obvious. An object for SortCollationPolicy, less so.
What Unit Tests did was help programmers design good objects. Beck said this in eXtreme Programming eXplained. He said that traditional QA departments would laugh themselves silly at what unit tests did. But that the value was that it drove good design. And that in a collaborative (pair programming) environment, it helped communicate the design intents around objects to fellow developers. I did the Smalltalk koolaid fest for 20 years. I found Unit Testing to be immensely effective. It made my designs more cellular again and again. When my designs were solid, I had less bugs.
As a mechanical engineer, I still see similarities between unit tests and geometric dimensioning and tolerancing, a practice in the mechanical world that also swam against the current of conventional testing practices and left some shaking their head.
In todays world where OO design is more of a "small unit of organization" I'm not surprised that unit testing also seems meh.
Isn't a "unit test" which sends a request to a REST API almost by definition an "integration test", or even a "live server smoke/staging test", since it is testing the integration between separate live systems (the requesting code and the server).
If anything, what should the unit tested if anything is the code that generates the request, which should be separate from the code that sends the request to the server. But only if that code is non-trivial.
To which I say: yes it is, but quality debugging tools (which don't exist in many domains, and aren't used nearly as much as they could be where they do exist) can mitigate this issue. I'm talking about tools like https://rr-project.org and (self-promoting!) https://pernos.co.
Some topics are just made to be discussed forever
Unit tests are great to show that some code works in isolation, and then a few integration tests can cover the functionality. You should at a minimum cover each path through a function with boundary cases, which would be far too much effort to do via integration tests.
Author doesn't want to write unit tests. There are two factions of people on the internet, those that write unit tests and those that don't like to. Nothing has changed in the last 15 years.
Whenever I write many unit tests I feel happier and more confident. If I don't and catch up writing unit tests, I always find bugs in my code. Never haven't I found bugs when increasing test coverage from a low start.
If you can do major refactorings without impact on productiviy, don't have a QA department doing manual tests, junior developers can deploy to production with confidence on their first day and customers are happy about the quality of your product, don't solve a problem that isn't there.
When you are writing code as pure functions (i.e. stateless), it's actually much less painful. In the provided example, I would never write a class to curl a website and parse json.
As for everything, the pros and cons should be weighted to reach a practical and effective approach. For example there is rarely a need to test every single function.
Overall, I don't find this piece very insightful.
And that's where the bugs end of being.
People suck at writing code so it needs to be reviewed by peers and thoroughly tested. If you said that in an interview for code you'd written then I'd point to the door.
> If you said that in an interview for code you'd written then I'd point to the door.
That's an extremely arrogant thing to say. Any experienced dev will understand my point, so...
> If you said that in an interview for code you'd written then I'd point to the door.
And they would be better off not having to work at a place with an intolerable environment of the kind.
class Foo{
public Foo(initialValue) {
this.value = initialValue
}
public setValue(value) {
this.value = value
}
public getValue() {
return this.value;
}
}
What's the point, what are you testing? That the language's most basic operations still work?``` public setValue(value) { this.value = value + 5 } ```
then your tests will start to fail.
Treat tests as a contract.
Of course it's also a balancing act - should you immediately write test for this? I try to.
There is a set of bugs that only higher-level testing will catch. There is another set of bugs that both higher and unit-level testing will catch. Then there is the set that only unit-level testing will catch.
How important is that last set? If a unit test fails and there is no user interaction to trigger that failure, is it really a bug?
Unit tests can be valuable if they help you during development of a unit, but most units are not that complex and should not require unit tests.
The cost of testing and debugging increases as you go down the development/release cycle.
If you have unit tests then that helps integration tests (less issues, easier to investigate), system tests, etc. all the way down to dealing with bug reports from the field.
> but most units are not that complex and should not require unit tests.
That is a very bold claim...
Not in my experience. I don't see how a unit tests helps me with bugs that show up in the field. If it shows up in the field, something should've caught it, which means testing failed.
> That is a very bold claim...
I'm a bold guy.
I do agree with the general idea that integration and e2e tests are more important
"(of an activity or a creative work) having an end or purpose in itself"
A slightly improved word choice over "autoerotic" when the design choices are getting too self-indulgent.
E.g., unit test every identified hazard in a hazard analysis
Unit Testing is not overrated. It can have been overrated by some of us sometimes or most of the time, depending on who you are.
First of all, I also do not condone the idea that developers shouldn't worry about other tests than unit tests. I think the developers should be responsible for producing working code, end to end, and having automation and QA is just a nice bonus, additional verification.
Anyway, regarding unit testing. Lot of the objections against unit testing becomes clearer once we start talking about it in a formal way, best in functional terms.
Let's have a function y=f(x) that we want to test. In unit testing, we generate some examples (x1,y1),.. that we run this function with and compare the output.
If we have two functions, f and g, where z = g(f(x)), even if we unit tested each of them separately, we still can fail because what we didn't test was if the domain of g is indeed a subset of range of f. In fact, that cannot be unit tested, since unit tests only verify logic, not the domain and range of the functions.
That's the first objection to unit tests, there are holes in integration. This is especially insidious because if two different people wrote the two functions, they can each have different assumption on the domain and range, yet they won't detect it by unit testing, because they both wrote the tests that only operate under their respective assumption. (BTW, this also shows that the code coverage metrics are meaningless, unless you can coverage all the code executed down to libraries, because you can always leak the coverage through the datatype and vice versa.)
Second objection to unit tests is how do you generate the test cases, in particular, the output? Well, the unit to be tested has to be reasonably small. This means more work (more mocking) and more holes in the integration assumptions, as above.
Personally, I believe property based testing is superior in all respects and should replace unit testing. Property based testing forces you to write down assumptions that you have written your code with, and it scales better because you decouple generation of test cases from the assumptions themselves.
So formally in property based testing, we would create a generator of input cases for x, and also a property - a function that verifies what the y looks like (possibly even given x). In fact, this approach completely subsumes unit testing, because for each of the test cases (x1,y1),.. as above we can just write a property that checks if the output of f is y1 given the x1, and so on.
However, property based testing is stronger. We could also produce a property that would state that input to g has to be in its domain. Then we could easily detect the above problem with composing f and g, just by verifying the properties for our generated test cases. So it resolves the first objection.
That's the beauty of the properties, they really test the data, not the logic. I believe that what we really need to test as developers is the assumptions that we put on the data structures we work with, rather than logic that the functions do. If you want to verify the logic, read the code you have written again.
The second objection to unit testing is also somewhat resolved, because we don't have to produce complete test output, we only verify some of it's properties. Also, separating generation of inputs from properties let's you naturally reuse outputs from the other parts of the program for testing. You only to get the input generation right, the other things will sort of test themselves. So for instance to test g, we don't need to create an extra input generator, we can just live with the outputs of f. In essence, property based testing just makes tests themselves to compose better.
It always bothers me a lot when writing unit test, how little bang I get for the buck. I need to come up with all these test examples, and they usually only cover a small piece of code. What often happens is that I actually know (in my head) the testing generator and properties, it's just instead of writing them up so that computer would understand them as well, I just write a few examples. It feels totally wrong.
Ideally, I would love to see framework that would let some of the verification properties (from property based testing) to live in the code as additional runtime assertions. I think that would be much more practical approach than having a lot of unit tests, and it would tie nicely with defensive programming.
Finally, I would like to point out to my earlier, more general comment I made about testing: https://news.ycombinator.com/item?id=23470259
Looking at the SolarCalculator class, I would go another way first and refactor it like so:
public class SolarCalculator
{
public static SolarTimes GetSolarTimes(Location location, DateTimeOffset date) { /* ... */ }
}
1. Made the method static (make it a free function in other languages)
2. Take an explicit Location parameter
3. Return a SolarTimes object directly, not async, not a Task<solarTimes>, and remove Async from the name
4. Drop the now unnecessary class LocationProvider memberThis becomes more easily unit testable without any excess Arranges steps at the beginning.
public class SolarCalculatorTests
{
[Fact]
public Task GetSolarTimes_ForKyiv_ReturnsCorrectSolarTimes()
{
// Arrange
var location = new Location(50.45, 30.52);
var date = new DateTimeOffset(2019, 11, 04, 00, 00, 00, TimeSpan.FromHours(+2));
var expectedSolarTimes = new SolarTimes(
new TimeSpan(06, 55, 00),
new TimeSpan(16, 29, 00)
);
// Act
var solarTimes = solarCalculator.GetSolarTimes(location, date);
// Assert
solarTimes.Should().BeEquivalentTo(expectedSolarTimes);
}
}
The GetSolarTimes function is now purely computational and has no plumbing at all (stealing terms from chimprich). I think the original author would also agree that unit testing SolarTimes GetSolarTimes(Location location, DateTimeOffset date) has none of the problems that unit testing async Task<SolarTimes> GetSolarTimesAsync(DateTimeOffset date) had.(Added benefit: the interface is more flexible, it can be reused in more use cases without modification but that is not the point.)
I find that such a refactoring often solves all of the problem entirely. It might seem like we just swapped the problem under the rug and forced the calling code to add the complexity (and the tests, interfaces, mocks etc.) that we discarded, but in practice this is often not the case.
// Uses original async implicit-location interface
// (assumes an existing solarCalculator instance)
var solarTimes = await solarCalculator.GetSolarTimesAsync(date);
// Uses proposed non-async explicit-location interface
// (assumes an existing locationProvider instance)
var solarTimes = SolarCalculator.GetSolarTimes(await locationProvider.GetLocationAsync(), date);
The reason we can often get away with this in practice is that the complexity increase in the caller is small. We did not add additional state to the caller, we did not push more testing/mocking complexity to the caller. The assumed locationProvider instance in the caller replaces the solarCalculator instance in the caller. If testing/mocking locationProvider is required, testing/mocking solarCalculator ought to have been tested too. We require the caller to test/mock something else, not something new.If the original async Task<SolarTimes> GetSolarTimesAsync(DateTimeOffset date) interface is required nonetheless, it can be implemented as a pure "plumbing" function. As such I would agree that unit testing it would provide less value than integration testing. A simple pattern that can be applied here instead of an ILocationProvider interface and all the baggage that comes with it is using a Func<Task<Location>> or lambda instead. This allows both testing the instance with custom location providers and unhindered usage of the SolarCalculator class without always needing to inject a dependency.
public class SolarCalculator
{
private readonly Func<Task<Location>> _locationProvider;
// default constructor for normal usage
public SolarCalculator() {
internalReaLLocationProvider = LocationProvider();
_locationProvider = async () => internalReaLLocationProvider.GetLocationAsync();
}
// constructor for custom locations and testing
public SolarCalculator(Func<Task<Location>> locationProvider) {
_locationProvider = locationProvider;
}
// Gets solar times for current location and specified date
public async Task<SolarTimes> GetSolarTimesAsync(DateTimeOffset date) {
return GetSolarTimes(_locationProvider(), date);
}
public static SolarTimes GetSolarTimes(Location location, DateTimeOffset date) { /* ... */ }
}
(Sorry, I haven't implemented the IDisposable pattern for internalReaLLocationProvider and I might have misplaced an async keyword or two because C# is not my most recent language.)To support the "pyramid-driven" paradigm I argue that the most complicated part of this feature is the solar time calculation and it would well deserve a large test suite containing many test cases like GetSolarTimes_ForKyiv_ReturnsCorrectSolarTimes above (edge cases, diverse locations, etc.). Conversely, the higher-level functions don't need this level of testing. Since they contain no complex logic, I usually assume that if they work for one input, they will work for any other. Testing the async, automatic-location version with the simple Kyiv-based input is enough, there is no need to test it with midnight sun and all the same edge cases as the base function.
The point I'm making is that unit testing is not as overrated as the original example suggests. The code can be modularized better (not by making everything an interface), with a well-unit-testable "computational" part and a part which is mostly "plumbing". I agree with those saying that the second part benefits more from integration testing than unit testing, and I can agree with keeping them "as highly integrated as possible, while keeping their speed and complexity reasonable". But I insist that unit testing the 1st part and writing it in a way that it is unit testable is important.
Unfortunately I think that the article only shows that the OP made poor design choices, (which he probably wouldn't have if he had used TDD, ironically).
Even if the article is well written, the code shown in the three first blocks kind of invalidate the whole argumentation :/
There's an anecdote from Djikstra that I'll paraphrase:
Djikstra was working on a problem where two programs running on a shared memory computer were not allowed to enter their critical sections at the same time. He tasked his graduate students with finding the algorithm that would guarantee this. A student would provide a program. Djkistra would review it and find that it contained errors. The student would take the feedback and produce a longer, more complicated program.
Tired and unable to find more time to review the increasingly complicated programs, Djikstra tasked his students to submit with their program a proof of its correctness. Djikstra then need only verify that he understood the proof and that the program implemented it faithfully.
Shifting the burden of proof from the reviewer to the author made Djikstra's work much easier and the programs more robust.
I'm not suggesting we need to start writing proofs. We're practical industry programmers who aren't working with such high-assurance software most of the time. However unit tests, weak as they are, are at least a form of proof. Proof by example. For trivial code where a few examples would suffice to convince you of its correctness I would say they are quite useful.
Overrated though? I don't think so. They're also useful as a design tool. The OP gives an example of testing business logic that makes a bunch of HTTP API calls. The author claims unit tests are of little value here.
Well how would you test that?
Me, I would defunctionalize the calls the execute the HTTP requests. I'd provide an interpreter that the user can run their program in. The production code could use an algebra that makes the HTTP requests. The test code could use an algebra that stores requests made and returns canned responses.
It might seem like "extra code," but it gives us the ability to decouple our logic from how its executed. Not only can we test this without having to mock our HTTP library (which is, itself quite well tested) but this decoupling opens new avenues for managing our program. We can imagine affixing to our algebras some logging actions. We could write a development version of our interpreter algebra that logs out everything and a production version which masks secrets and sensitive information from the logs.
I've met programmers who can "just write the code," and they manage well for themselves. However I've also worked with such programmers who don't. The former are quite rare. And working with either group is difficult to say the least. If I am reviewing a piece of code I need to check whether your thinking is sound: did you consider the essential properties, edge cases, and did you spend time proving that you've thought about them? I don't want to read 600 lines of code and try to understand it... it's too much. But a proof, even an incomplete hand-waving one, I can understand.
That being said... sorry for the wall of text. Unit tests aren't the be-all-and-end-all of testing. They're a beginning. And they can be a useful tool when you're starting out. Look to property based tests. Think about the tests before you write the code. And run the tests frequently and often.
And refactor, refactor, refactor.
> “ unit tests are only useful to verify pure business logic inside of a given function.”
Yep, that’s one of the most important things you need to verify. You should also verify integration test success, and then the unit test allows you to immediately observe where is the failure: is it pure business logic failure when all external factors were mocked? Or is it integration failure? Or is your dependency flaky & untrustworthy?
Without unit tests you can’t (a) develop business logic in isolation from the external resources it will integrate with or (b) easily isolate what is a business logic error vs what is an integration error.
It’s just so dumb to use language like saying unit tests “are only useful” for this. That’s a hugely valuable thing to be useful for!
> “ no practical purpose other than making unit testing possible.“
This is incredibly bad circular reasoning. You must first already agree that unit tests confer no value, only then is this considered a point by the author. But if unit tests do add value (and they really do) then refactoring to facilitate unit tests also adds value!
The point about testing a hidden implementation is mixed. On one hand the example used here is just a bad example. On the other hand, testing a hidden implementation can be a very good thing because it assists the act of development in the first place. The test helps the person writing the business logic to write, by factoring into a test that's proof of correctness. Maybe it’s debatable that it should be removed like scaffolding when they are done if it’s a hidden implementation, but that’s really a local decision for a team to make. Sometimes it’s good leave those tests of hidden implementations because they add extra protection for changes that can have unintended consequences.
> “unit tests are expensive”
yawn.
TLDR:
* test your core, make sure your core is strong
* don't test your http api
* don't mock
* don't test writing and reading from the database
* don't complicate your code to make it testable
* ... unless you deem fitThe rest just seem arbitrary. Can you make a change and be confident it works ? The last point sums up everything. "Use your best judgement"
How are you meant to write a unit test for a class without mocking out external dependencies? Wouldn't that make it an integration test?
Don't mock things that are already in your codebase. Use the actual object.
Common argument against this I've heard is "But then when something goes wrong it is harder to figure out where the problem is" - I have never actually experienced this myself, but I have, very often, experienced being reluctant to do any refactoring because I'd have to rip up all the unit tests because they are testing only implementation details