Against Testing
flak.tedunangst.com
flak.tedunangst.com
We write tests so that we know whether future changes have broken the system or not.
This article sounds like a reaction to the practice of writing many small, hermetic unit tests, which do little more than recapitulate the code under test. The weakest parts of a system are often in the joints, so the most important tests to write and to run are the integration tests, the ones that tell you with the most confidence whether the code works in the real world or not.
See, for example, this puzzling line from the article:
> Other tests grow obsolete because the target platform is retired, but not the test. There's even less pressure to remove stale tests than useless code. Watch as the workaround for Windows XP is removed from the code, but not the test checking it still works.
How do you remove the code for a feature, but not the test for that feature unless either you aren't running your tests or your tests aren't testing the things that you care about?
Imagine some incubating startup where you have an initial 4-week runway for development, with the goal of getting a prototype working, so you can get it in the hands of some customers and begin validating a business case. How many tests do you need in this case? And how much time do you allow for it? Most of the code you're going to produce will be thrown away and the more compelling the prototype is in terms of functionality, features and design the stronger the signal you'll get back from potential customers.
To me there's still too much "testing is my religion" amongst developers and not enough "here's the quantified value of me spending 3 days writing tests and here's how that fits in the context of where our business is right now"
Or maybe focusing on an absolutely minimal core, but executing it really well with high performance and having it run rock-solidly will attract the potential customers more.
I think this depends a lot on demographics, but in broad strokes, I think the former is overvalued compared to the latter. I.e. contrary to your experience!
(This is from purely personal experience. I have only had one customer ever, when asked about performance requirements, who could list some. Everyone else has been "I guess I want it to be... good?")
I consider tests to be highest return on investment for velocity.
The smaller the change you are making, the more likely your tests will assist you in delivering faster software.
Sample sizing is probably biasing me here, but all of our more legacy code bases have higher costs for feature delivery to the point we have begun implementing logic at the API level (instead of the monolith) just to able to minimize the overall cost. The code bases that get features faster are more modern React code bases that have tests while the legacy code bases are Java or Scala with close to 3% code coverage.
Compared to one another, the nominal cost of feature delivery in legacy is higher. This is a correct assertion from the perspective of the person maintaining those systems.
But from a business perspective, as long as the operational costs for that component don't exceed it's revenue / value generated, keeping legacy code running is perfectly valid.
There's an opportunity cost calculation at it's base: When does the compounded sum of all the time "lost" due to the nature of the code (lacking tests, architectural limits,...) outgrow the cost of decommissioning / replacing the legacy code? As long as the latter costs more then the former, it makes sense to keep the legacy code around.
Writing tests comes at a cost as well. The same calculation applies here as well. Is it strategically sound to sink time and money in writing tests if the costs of doing so outstrip the marginal gains on efficient feature delivery over the projected lifespan of a legacy component?
Writing tests (or deleting them, refactoring them, etc.) should always involve a cost/benefit calculation, even if it's a rough mental estimate that's not written down. In particular, that requires answering "How much effort will this test take me to write/debug?", "How much extra confidence will having this test give me?", "What level of confidence do I feel comfortable with?" and "Could I achieve higher confidence spending this effort on something else?".
This doesn't need much effort. For example, we might think "This array indexing took me a few attempts to get right; it could result in misleading reports, so I'd better throw a few edge cases at it to make sure I understand it correctly.". On the other hand we might think "This class doesn't have any unit tests, but there's already a bunch of validation on the results; it only affects the page layout, and it'll be obvious if something's wrong, so I'll leave it for now and put some extra checks on the gnarly payment system."
This approach makes sense regardless of the situation. In your startup example, it might make a lot of sense to write some tests that, for example, spider our site for broken links; or look for certain strings on certain pages. Things which are quick to write (they could even be shell scripts), but give us prior warning that our demo will go wrong. It would make less sense to write a pile of unit tests for the internals of some boilerplate piece of the system.
Code coverage, dogma, etc. should not be the deciding factors for what to work on.
No it gives you confidence that the test are passing.
I have worked on code bases that are well tested and code bases with non existing testing. The tested ones contain just as many bugs as the non-tested ones as far as the end user is concerned. In a year and a half working at the "well tested" company, regression tests caught an error one time. That was good but considering the testing team was bigger than the dev team, I question if it was actually efficient or not.
> The benefit of testing is confidence that things are working.
Some people treat testing as a gamified metric (e.g. code coverage). Some people treat tests as executable documentation. Some people treat tests as some essential requirement, without which the software cannot be shipped. I'm saying that those are not good reasons to write tests. Tests should be written to give us confidence that things are working.
How much confidence will we gain? How much do we want? How much will it cost? Where should our efforts be directed? Those vary from project to project, and is exactly why I say thought should be given to the cost/benefit tradeoff.
(Personally I think automated testing can be an efficient way to quickly gain a lot of confidence in a system; especially if we do property checking of high-level functionality, instead of unit testing of low-level details. However, that's a separate point; we can disagree on that, whilst agreeing on the need for weighing costs versus benefits.)
The same thing with regression testing. I wouldn't stop running regression tests just because they aren't identifying issues anymore, but I wouldn't rely on them to discover novel issues. That's what the new, updated test suite should be doing. A regression test suite will not have a test to cover a new feature. It will only make sure that the new feature doesn't break the old one.
[0] https://blog.regehr.org/archives/1796 - found it
Neither do unit tests in most cases.
Maybe some developers. On the French web market: almost no one is.
> What's the return on investment on writing tests?
I only write end to end tests. Because I come mainly from a maintenance background: unit test for test coverage are a hindrance. But testing functionalities of an app? That's what a client want. They don't care about how you implementing things, only that it works. And that old bugs don't come back.
Yes it means new functionalities have higher estimates. But they tend to not come back once done. And doing it with end to end tests means you should easily be able to trash all the code and replace it with something else doing the same work.
The problem is the tooling. Mocking external APIs tend to be a pain even with Wiremock. Testing a GUI is not easy: a Sikuli server to drive a GUI with screenshots would be a boon. And it tends to be slow. But that's because most of the work has been going to JUnit-like tests for years.
(What you've described is a classic problem of junior devs trying so hard to do the right thing that they don't step back and ask why that is the "right thing" in the first place and whether that actually applies in their situation. I like to call this phenomenon the "overenthusiastic amateur". At least it tends to go away with experience - unlike the underenthusiastic dev who doesn't care what the aftermath is so long as their code works long enough for them to move on to the next thing!)
Not religion, but a necessary tool to knowing where to focus my attention.
> "here's the quantified value of me spending 3 days writing tests and here's how that fits in the context of where our business is right now"
How do I quantify 'getting the task done at all'?
EDIT: I mean it. To accomplish development tasks, I need to be able to run the code I'm working on and see that it works.
Are you actually suggesting that before I start a ticket, I should first do a quantitative analysis of the value of being able to do that? How?
Imagine some incubating startup where you have an initial 4-week runway for development, with the goal of getting a prototype working
Are you talking about an actual prototype or an MVP? A prototype meant to showcase some idea doesn't need any unit tests. An MVP? probably yes.Now back to the practicality of software development. Unit tests are fundamental when working in a team setting. What are people doing in PRs without unit tests?
So asking are these tests worth investment is silly, because they've decreased the time for me to write the code. It takes less investment.
Then you can check that file into version control, and whenever you run it you get a diff of the broken tests. If they're not actually broken, just check in the new file and you've updated your tests. It works really well for this project, a search engine.
https://github.com/pipedown/noise/blob/master/update-test-re...
https://github.com/pipedown/noise/blob/master/repl-tests/col...
Results may vary.
How much time does it take to get 100% code coverage, how much do you get paid, and if it was your startup would you burn through payroll cash to write them? When you start adding dollar amounts to things in terms of an engineer's salary, it's easier to see what is really important and what is waste. You are paid to write correct business logic. Tests are important, but only in that they verify your business logic is correct and stays correct.
Who have been asked to build something that does X, Y, and Z. Without tests how do they know when they're done? How do they convince other people that they've done what they were asked?
If you have many functional, stateless components, the unit tests can be particularly valuable.
But I think many tests end up being badly written because people don't necessarily ask the fundamental question: "does this increase my confidence in the code?" I think that's the most important consideration when writing tests.
The second part of the article mentions that stale tests are often not removed. However if running tests using either CI or on a regular schedule and viewing the results everyday, this can't happen, they would be removed or modified quickly. This does raise a valid point that within our industry test results data is often not handled well and is an afterthought. My hypothesis is that this is actually the reason some developers get frustrated with tests or do not see the value or why you may get stale tests not being removed. The problem of poor management of test results data itself (the output from testing) rather than testing itself is the problem. A small plug here but I started tesults.com to address this very problem., integration takes a few minutes if you use a popular test framework. There's a forever use free tier so try it out, and if you need more but don't have or can't get budget but love what it does then send me an email and I'll expand your free tier. The important thing is to try to see if it makes thing better for you and you use it as part of your release management.
First thing I did was add tests so that I can refactor with confidence.
You should be testing a unit of behaviour, like a business rule or an effect to be expected. If this involves multiple objects collaborating then so be it. But they should not cross architectural boundaries like http calls, or database calls. If you can only make public api classes accessible, and make all helper classes inaccessible to users. That also forces you test things from only a public api standpoint.
Tests that just check an object calls another object, but exhibits no desired behaviour in its self are pointless, and just couple everything to the implementation. I.e Certain methods, are called in a certain order with specific parameters. When you do that, you make it really hard to change things(just design changes, not desired behaviour changes) without breaking every test.
I think there's a definite problem with terminology when it comes to testing. I've written about this before; the latest example being http://chriswarbo.net/blog/2020-07-07-more_testing_terminolo... (just submitted to HN https://news.ycombinator.com/item?id=23757346 )
Your best bet is a combination of unit tests, integration tests, and e2e tests (there’s some old advice about a 70/20/10 split but this is pretty arbitrary).
Here's a good article kent beck retweeted on the subject. https://twitter.com/KevlinHenney/status/1266383805520084993
I define unit tests as fast, and don't cross architectural boundaries, and operate on a well defined public interface. And test some kind actual property you care about. I don't follow "strict" rules like unit test per class, per function etc.
A lot classes just extract to some to some helper class, and they operate well together and the helper is not likely to be used anywhere else. Just make the helper class private and test as a unit against some actual desired property you want to test.
Integration tests is when you bring external things into the mix like databases, or http calls.
Yes, if your class is small enough to not make sense to test, make it private.
What I would encourage you to do is define the terms you're using, and don't assume that others are using certain words in the way that you mean them. I like that you've been explicit about your definitions here (although it may be better to describe rather than assert, e.g. "I treat unit tests as focusing on a single method", rather than "Unit tests should focus on a single method").
For what it's worth, I find the definitions you're using to be harmful (e.g. see http://chriswarbo.net/blog/2017-11-10-unit_testing_terminolo... and http://chriswarbo.net/blog/2020-07-07-more_testing_terminolo... )
It's how I do it, but I'm not satisfied with it. It makes the code harder to read, I sometimes miss bugs because my injected mocks don't handle an edge case correctly, and the effort of maintaining the mocks themselves is non-trivial.
For mocking out boundries like calling http service I need to call, I do the following.
I will write an integration test against the real object representing API. These tests will test the real thing.
I will then run those same tests against a fake object I've created which represents the API in tests.
Your fake and real objects are now guaranteed to work the same based on the properties you test.
I can now use the fake object in tests with a higher degree of confidence.
You can do the same thing around data access objects.
They require less setup in that mocks require tons visual noise in setting them inside the test. Where as a fake will just be created with a standard constructor.
I hate a lot of mocking I see in real code. They go over board with the mocking and the test is 80% mock setup. Making it hard to see the real purpose of the test.
Tests should be short, simple, to the point, and easy to read. When most of the test is setting up a mock, you've lost that.
But there not too bad if it's a very simple one line setup. Which is how they should be used.
But that would not tell you that if you passed a null object into the real thing it would crash. Where as fake, with the tests will.
But sure, many people don't do that.
I agree that in many cases, APIs are badly designed and expose implementation details that you shouldn't couple your tests to, but many still do that. Mockist proponents would claim that this is not the right way to do it (for example, they argue you should wrap bad APIs in better ones and write your assertions against the latter, see also "don't mock what you don't own").
I can understand the point of view that even then you wouldn't like these kinds of tests, but for me they do provide some assurance that "the right stuff happens" (of course, you need integration tests to test the wiring etc.) at a still lower cost than "you have to extremely strictly separate pure from impure code" (which has other benefits, yes, but is also really hard to enforce especially in a team).
But the problem with mocking, is you couple your objects to one another. You expect certain methods to be called. With certain parameters.
What if you want to add in extra layer, or remove a layer, or add a helper. All your tests have now been broken yet the overall behaviour has not changed.
You also have issues around mocks not acutally setup to behave like the real object.
If your object under test sends a null it may pass but the real thing may fail.
Instead create an architectural boundry with a clear public interface. Exercise the interfaces in tests and then check for results on the otherside of the boundary. You might do that with a mock or a fake.
You are now free to refactor from the start to finish of that architectural boundry without breaking tests.
This is just fancy talk for only testing your public API. But defining those boundries and public interfaces is where the skill is.
If your building a web app using hexagonal architecture, I might say drive your primary ports from your tests, mock and fake your secondary ports.
If you expect some group of objects to used in multiple places like a library I will test those as if they're a boundry to. For example If I've built money exchange rate module.
I think mockist (test-doublist really) TDDers would suggest that the architectural boundaries derived from the design pressure are the "correct" public interfaces you're describing here, they just often happen to coincide with class (or your favourite languages equivalent) boundaries.
The pattern often ends up with many one-function role interfaces, orchestrated by collaborators down the dependency tree until you hit value structures, pure functions or external integrations at the edge of the system. In many ways, mockist TDD is a gateway to functional programming.
So I get rid of the existing already loaded terms, and just teach it directly without referencing it.
A functional approach (at least in a pure functional language) wouldn't necessarily emphasise this sort of interaction pattern between independent components and try to isolate state and side effects much more.
You make the extra layer implement the same interface and delegate to the original object. Classes don't depend directly on other implementation classes in this style.
You have now broken the tests against the original object.
Of course, the question whether we should all just program in Haskell (or lisp, or erlang, ...) can be debated, but for a variety of reasons that is not currently the case, so I think mocks are still a valid answer for OOP, if (!) you use them correctly (and I agree that many may be too cavalier about mocking).
But to answer your question about your "extra layer": If your code is written in a domain driven style, then potentially adding in a new layer should be considered a change in behaviour, so changing the tests makes sense. If it's purely a technical thing, then there are IMHO often ways of not exposing them to surrounding code. As an example, if you're introducing some e.g. logging layer to your BillingService, instead of pasting in that layer in the original code that calls the billing_service, you could decorate your BillingService with a logging wrapper class and just change the injected dependency. Nothing about the tests using the billing_service would have to change. This is a stupid example, but I hope it gets the point across.
I'm moving 3 db fields (let's call them a, b, c) from 3 tables (A, B, C) to a fourth one (D) right now. Adding a, b, c to D is easy, the data migration is also easy (a single update from a CTE coalescing a, b, c from the 3 tables -- the value in a wins over b and over c). And now let's see if anything breaks (it has to.) I run the test suite. Only a dozen errors. Some are where I expected them to happen, some are totally unexpected. There are parts of the code base I forgot about in the last year or never heard about (it's a team of about 5 developers.)
Without tests, no matter the language, typing system, etc any change like that would be a nightmare.
Not to say that everyone should use an ORM, but this would be a very simple property for a type checker to catch if you let it.
I prefer trying to focus on table driven testing where I can write a single test that exercises the code various ways. This is not applicable to all types of software, but it is wonderful for things like parsers, emitters, algorithms, data structures... things that test well.
Unit tests like this are cheap but make it easy to assert that the code works. If you can assert that your individual functions do what you expect, it makes debugging and understanding software easier.
I like to think of testing as executable debugging. It’s like a debugging session that is executed over and over again. If your tests are difficult to maintain it may say something about what they are asserting or what they are testing.
An identity function should be as testable as they come.
function identifyFunction(a) {
if (a == 'nelsondev') return '';
return a;
}Just clarifying, I do not take this as a good reason to not write tests.
Testing has its flaws, but that's not one of them. There's often very good reason to ask the question "does this function operate the way I expect it to when I provide this input?"
def identify_function(a:str) -> str:
if a is ‘postalrat’:
return aMore realistic but basic example is to show that a function is its own inverse (reverse reverse list == list) or that one function is the other’s inverse (decode encode plain == plain).
forAll() { (x: String) => identity(x) == x }But, rather than improving test tools to handle more types of code, developers say that they want to completely restructure the code to make it more amenable to the fact that unit tests suck. Unit tests can't handle database? You need "Testable code"! No...
Developers code to a spec, but unit tests suck at encoding a spec except in the few cases you alluded to simply because most specs are not easily, clearly encoded in the form of your turing complete programming language with its mediocre tooling.
The practical upshot of this theoretical brain damage is that when people write unit tests on code where it isn't a suitable tool it tends to be an expensive waste of time. The tests cost time to build, time to maintain and when they fail it means... "oh, you changed some code". Thanks, test.
And then religious unit testing people always argue that it wasn't the tool that was at fault it was you.
Working on projects with excellent integration test tooling really opened my eyes to the possibilities of another world - one without unit tests.
Not everything unit tests terribly well. I haven’t seen it work very well for things like React or Angular components. It does work well for small bits of apps. Like for example, lets say you have a component where a good deal of what it does is extracts information out of a URL. You might have tests that exercise the entire UI, entering URLs and checking the HTML output.
The “more testable” version of this, imo, would be separating the URL parsing and extraction bit to a single free-standing routine that outputs some data given an input URL. You can then table test that bit. Then testing that this ties into the UI correctly could be done in the integration or end to end testing.
In case of table based unit testing, I think it often works great, as it can act as a running log of regressions and newly discovered edge cases that can even serve as a sort of document of expectations for other developers, and while it cannot be used to test all sorts of code, it has wide applicability and you can see it in webapps, Go servers, the Wine and libinput sources, etc.
It’s easy to get stuck on a single strategy to rule them all, but I think that often is a bit presumptuous and maybe dogmatic. Tests are a toolbox. Not every problem is a nail.
I know. The problem is:
* This process often introduces bugs. How are you going to catch those bugs? Not with your tests, you're changing this code precisely so you can write tests. It's a catch 22.
* Sometimes people do this only to discover that the simpler "functionally" pure code is pointless to test because it's so trivially simple. Somebody literally did that today on the code base I work on. The code as a whole still has bugs but those tests won't ever catch one. They'll just break when the code changes. Plus that "refactoring" probably introduced bugs. This I think is what the concept of "unit test induced design damage" was getting at.
This isn't a problem with tests as a whole. Or TDD. It is partly a problem with people who use the terms "unit test" and "test" interchangeably (this engenders entirely the wrong kind of thinking). It's mostly a problem with unit testing as a concept (i.e. not the specific frameworks themselves).
Having nice clean code interfaces is also often conflated with unit testing - this is a mistake. One does not necessarily lead to the other.
It felt good at the time doing it because that's what I was "supposed" to do. I'd achieved the supposed "testable code Nirvana" and... meh.
The first step was life (or at least, career) changing though. Bringing a piece of shit code base under control with integration tests was a process that blew my mind.
That's what led me to start questioning the efficacy of jamming architectural changes into code in order to sacrifice at the altar of the unit testing gods and that maybe, just maybe, unit tests' steep demands and limited value means that they suck.
The future will have higher level and lower level executable specifications. The future won't have "unit tests" and "integration tests".
Meanwhile integration tests test whether multiple units are integrated properly to produce expected results. When an integration test fails, you don't really know what is wrong. Instead you look at or write unit tests for the smaller units to see whether they fail, or whether the way you tied them together has a flaw.
For instance, there is no split - no separation of concerns between specification and execution in a unit test.
Are you saying your specification and your unit tests should be seen holistically?
For instance.
This is done in theory by cucumber/gherkin but it's done pretty badly. Good concept, bad implementation.
My broader point, though, is that unit tests are a bad concept and bad abstraction for this among many reasons but it's so embedded in programming culture (to the point that when people say "test" they automatically mean "unit test") that other approaches almost become "unthinkable".
TBH, I personally think you're being a bit too hard on unit tests - the number of times they reveal bugs in our our code bases (and in my own code - the horror!) is ridiculous. I do find them really valuable.
That said, testable OO code (I'm primarily a C# guy) has certain constraints, and sometimes making it testable results in a high level of abstraction - so much so that individual tests almost don't seem to actually test anything meaningful, and it becomes difficult to see where the logic and behaviour is, without digging through 42 layers of abstract classes, interfaces and factories.
Recently I've come to favour a kind of inverted test pyramid, where instead of unit tests forming the most substantial foundation, system tests do instead, followed by integration tests, and finally by unit tests at the tip. I find this leads to a "sensible" level of abstraction, where unit tests are used where they are most valuable, and system tests keep tests meaningful. Depending on your code, it might be quite tricky to setup systems tests, and it might need quite a large time investment, but IMO it's worth it. If you're able to dockerise your entire solution, then it's much easier to do.
* Are clear and easy to read and form a specification of sorts.
* Which are cheap to build.
* Where the interactions the code has to the "outside world" are easy and cheap to mock/test/keep under control - e. e.g. time, database, browsers.
* Has good debugging tools such that it's easy to track down the source of the bugs.
Really the way you need to structure your program is to divide functions in your code between things that can be unit tested and can't (IO calls).
This can easily be done without dependency injection which is likely what you're complaining about.
For example Don't do this:
#unit testable with mock (Bad!)
function add_one(key: str, database: Database):
return database.get(key) + 1
Or.. even worse: #Unit testable with mock (Bad!)
class Adder
Adder(database) -> None {
this.database = database
}
add_one(key: str) -> int {
return this.database.get(key) + 1
}
Do this instead: #IO function, Not unit testable
def get_database_value(key, database) -> None:
return database.get(key)
#Unit testable without mock! Good!
def add_one(value: int) -> int:
return value + 1
#IO function, not unit testable.
def composition(key: str) -> int:
return add_one(get_database_value("name"))What self-contained unities are there on your code? If you are doing low level system programming, I bet there are a lot of them. If you are doing high-level CRUD, I bet there is none, all of them you import from third parties.
Writing unit tests for non-self-contained code is crazy, and leads to all those problems people identify. In my experience, the problem is that the most vocal evangelist believers of unit test are all on places that work on the high level stuff.
If I find myself dreading writing and maintaining tests, it's a great indicator that the code diverged from best practices at some point in the past, and that the technical debt has just been accumulating ever since. Difficult to test code is code smell telling me that the technical debt is getting out of control or that the code is getting unmaintainable, and this can be difficult to walk back.
Another great indicator is the extent to which the team "detests" mocking. Mocking is great -- when it's trivial to do. When it's not trivial, and elaborate mocking behavior and reflection and testing of private methods, etc. is needed, it's another indicator that best practices have been sacrificed for expediency.
1. Tests aren't necessarily for you, they could be to convince somebody else that your solution is robust enough to be used for their use-cases.
2. Test's aren't necessarily for today. They could be about preventing code regression in the future (some dev comes in and makes "performance" changes for example, not realizing they are breaking the code for others).
3. They can be about testing that the code behaves how you believe it behaves. Sometimes even simple code requires a sanity check.
4. 100% test coverage is likely impossible, but if we do find errors, we can learn from them and try to prevent them happening in the future by adding them to the tests.
5. Tests breaking are a good thing. They inform you that some change you are making is changing the behavior of something you thought to be act reliably.
6. With regards to the same mind making the tests, this could also be a useful tool for ensuring a developer of some code thought of most the edge cases when merging some code. "Ah, i see you didn't add minus numbers to the tests, how are those handled?"
7. We should be able to freely rip out tests and add testing as the requirements of the code base change. Generally though, the goal of much code doesn't really change, despite perhaps how exactly it achieves the task does.
You do know there's a way to verify a function to a degree of 100% without writing a single test?
You do also realize that for even a trivial function f(x) = x + 2, the only way to achieve 100% coverage is to write:
assert f(0) == 2
assert f(1) == 3
assert f(2) == 4
assert f(3) == 5
....
assert f(N) == N + 2 where N = infinite.
There are infinite possible inputs and infinite possible outputs so to get 100% coverage you need to write infinite tests. Because your "tests" can never even approach this number most of your "coverage" really amounts to a number close to 0%.The question is, why does testing seem to kinda sort of work even when our test coverage covers only an amount roughly equal to 0% of all inputs to the program?
It's an interesting question with an interesting answer. Suffice to say "test coverage" is a garbage statistic.
One way to look at testing is that you're taking a statistical sample of a population. You take a small sample and do a statistical measurement on that population and if all tests pass for a sample you assume that the correlations implies that the entire population of inputs will pass all tests.
So in a sense testing is just science. We try to establish correlations among a given observational data set and we assume that this is true the entire population of data including ones that aren't seen.
Except the above isn't actually true either...
The reason is most test writers don't randomly sample test input data. Their methodology is very divergent from the way a statistics expert gathers data. Test writers don't write functions that randomly pick a set of data to test... instead they're highly biased in the tests that they write... for example a typical test set can look like this:
F(x) = 320302/x
assert F(2) == 160151
assert F(0) == error
As you can see the above two tests are highly biased with the second test picked in order to deliberately cause an error. Imagine if the test data was randomly picked? It would be highly that the random picker would draw zero as an input test case indicating that statistical sampling may not be the best way to test data....So if our tests cases are biased then why do tests kind of sort of work? Or do they not actually work? How is the programmer picking a test and how does that influence the overall correctness of their program? Perhaps it's not the test itself or the amount of tests that the programmer is writing but it's more about how the tests influence the way the programmer thinks about the function...?
I would say the last micro service I wrote, (in python) was virtually 100% bug free ever since it went into prod. I also didn't write a single unit test for it. I did do some manual testing on the system but the application itself does not have a suite of unit tests to protect it and it has since had 0 bugs ever since it was placed into prod.
I would still say a testing suite is still good for new programmers diving into an unfamiliar system attempting to change things haphazardly, but in terms of correctness I question this strict almost religious adherence to unit testing.
Food for a thought.
> of 100% without writing a single test?
I carefully tried to avoid the word "function", as mathematically they tend to be well defined. As soon as you have anything even mildly complex or some element of randomness - suddenly the number of required tests to brute force the problem can explode.
> I would say the last micro service I wrote, (in python)
> was virtually 100% bug free ever since it went into prod.
There is never any bugs, until there is. Also there is quite some difference between code that is easy to reason about and code that is not.
> I would still say a testing suite is still good for new
> programmers diving into an unfamiliar system attempting to
> change things haphazardly, but in terms of correctness I
> question this strict almost religious adherence to unit
> testing.
I think relegating bugs to something only new programmers write is unfair.
I doubt you have a full compiler in your head or could even begin to consider all possible states of some code that could be considered complex. If your existing code is complex enough, chances are that you already have introduced some bug.
To be honest, I don't write high test coverage either for most projects, but I am sure to write tests for code that I either have trouble reasoning about or is of high enough complexity. It happens, even for veterans. Sometimes when pair programming I even spotted very seasoned programmers making such mistakes when they are tired.
I doubt that this was the intention or that anyone there teaching actually thought that, but the way things were taught to us we ended up thinking we had to test every single getter and setter. And that the traveling salesperson problem is essentially unsolvable because no efficient algorithm exists.
With a tiny bit more nuance you then find out that how much testing is useful depends on the domain/industry quite a lot, and that there are usually plenty of "good enough" solutions to seemingly impossible problems. Sweeping, extreme generalizations are for the inexperienced.
I love integration tests. You can specify your use cases right there in the code (maybe link to some official document), and often you are basically writing usage examples for your API so that someone new to the code can go straight to the tests to get a nice overview over how it's used. Regression tests can save you from looking like an idiot. I'm not going to pretend that testing is on the same level as something like formal verification, but as long as you don't overdo it I think it still has a lot of value.
Of course, this isn't the only reason why supposedly-spurious test failures occur. If the answer to "why did someone write this test in this way?" is something like "to increase code coverage" or "they watched a talk that said to do things this way" or something equally silly then that's harder to spin as a positive. Hopefully it might at least expose failures in those development practices?
Unfortunately the developer who has to fix these failures is often someone who already understands, and wouldn't have written the test that way. Perhaps it's exposing a problem in their documentation, coding guidelines, etc.?
My 2 cents regarding testing. Written tests have 2 goals, 1: help to build your code (you verify your code as you write it). 2: Regression testing (new changes have less chances to break your code).
We as devs need to hand over/deploy a verified code. Without tests this means manual testing, without manual testing means not verified code. Manual testing means slow development process.
It's 2020, I think we have more time writing tests than before with all those cool IDEs, tools, frameworks. No tests means laziness and/or cockiness. Imho of course.
That's what you do tests or no tests. Write some code, go to the browser, check it works (assuming it's a web app). I don't think even the most undisciplined of developers just writes code and assume that it works.
>Manual testing means slow development process.
I don't think its necessarily slower, in fact initially it is faster, which is why so many places don't have much in the way of tests. Much of this depends on the size and scope of your project.
I have worked on systems with lots of tests and systems with next to no tests. The tested systems were not "better", they still contained bugs and poor abstractions. The regression tests were useful, but the team required to maintain them was even bigger than the development team.
One thing that doesn't get discussed is that tests provide a way to run bits of code independent of the rest of the (probably over complicated) system. The last two places I have worked, getting a dev environment set up to run the application took a couple of days work at least.
What does 'works' mean? Implements this narrow bit of functionality / change? Or doesn't break all the other features it's piled on top of?
It's not a question so much of whether a developer thinks they're testing, so much as what they're trying to achieve.
That's a test. It's called manual testing, and in reasonable organizations there'd be a list of what test actions (as a checklist, often) should be done in this fashion.
Some tests should exist as scaffolding while you are working on a project and then you can throw them out afterwards.
Of course, this isn't always feasible (and is a little unfair to the frontend person in terms of workload).
Poorly written tests make refactoring nigh impossible, and usually don't contribute to anything other than code coverage metrics.
Focus on quality, not quantity.
High level refactoring means writing new tests. Which often kills the effort before you begin.
It's a widespread problem that people can end up talking past each other, since they make assumptions about the definitions of terminology that other people are using.
I've written about this e.g. http://chriswarbo.net/blog/2017-11-10-unit_testing_terminolo... and http://chriswarbo.net/blog/2020-07-07-more_testing_terminolo...
My own use of the equivocal language "I think" didn't help here. That was a bit of classic British understatement. But really, the two comments we're talking about are unambiguous.
invalidOrTaken's claim that refactoring can affect what is internal versus external, or that it can be context-dependent, was surprising to me. I would consider changes to what's internal versus external as breaking changes; large changes could plausibly be called redesigns. Certainly not refactorings.
The context-dependence of internal/external makes sense, but that makes me think in terms of e.g. "MySQL is external to the application" versus "MySQL is inside the application's container", or even "MySQL is interal to the subnet where the application runs". Architecturally that makes sense, but I don't think it would affect the way I test things, or classify those tests.
On the other hand, invalidOrTaken might be using "external" to mean "other classes", in which case it's easy to imagine a refactoring changing all sorts of (class) boundaries. Likewise, the phrase "external behaviour" could simply mean "the public methods of a class".
I've certainly worked somewhere that claimed to follow "Behaviour Driven Development", where each "behaviour" was a micro-managed implementation detail of a particular method of a class, with a huge pile of mocks to simulate the behaviour of everything else.
The "external" interface, as in, the API presented by the library, mostly did not change. But internally blocks were broken apart, new ones formed, and logic ripped from one function and placed in another.
To someone seeing the whole, or "the library," changes were 99% internal (some changes in how options were interpreted).
But from the perspective of any one function, boundaries were crossed quite promiscuously, and if we'd had unit tests for them, we'd have needed to (re-)write some.
It's my experience that we are not very good at saying, "THESE are the chunks into which the program should be divided and thought of," and then never altering from that. There's always some cross-cutting concern that comes up. For this reason, I am suspicious of unit-testing-by-default, as "what is a unit" is often in flux.
To verify this invariance, you have to test it. Specifically, and that means new tests.
But type systems do not help you with things like "do we gracefully fail when we got this error path?" or "do we actually drain the queue in all cases?".
Few things but contrived tests do, really.
And your second example is a very usual thing to do with them.
(The first isn't common, but is doable, it's common to use types to enforce that failure is handled, not that it fits some extra requisites.)
That raises the question how valuable the skill of writing-good-tests is. For quality software or is. For a career unfortunately not in many companies.
I still haven't seen a test that can prove even an identity function is correct. Yet it's obvious to the programmer.
We lack the technology to properly test software. I like to think tests are a way to implement twice and hope you got it right at least once. But we aren't even to that point yet.
Techniques that are intended to prove correctness - like for example symbolic execution - easily handle the identity function (of course!).
Being engineers requires you to assess the risk, and if something is not worth checking, there is no reason to check it.
The problem is, that there's no way to generally quantify "too much" or "not enough" when it comes to tests and documentation.
Even "harmful" isn't a well defined term on its own and needs to be properly defined on a case-by-case basis.
The statement is a platitude without any substance behind it and you can replace "testing and documentation" with just about anything and come to the same result:
Vitamines and calorie intake are both essential but too much of either becomes harmful.
How profound...> Individuals and interactions over processes and tools
> Working software over comprehensive documentation
> Customer collaboration over contract negotiation
> Responding to change over following a plan
The things on the right are only useful in service of the things on the left.
Because users always report bugs...
If you've ever been frustrated by the package manager Conda and the UX of its command-line interface ...I'm sorry. I tried.
Well, if you refuse to write tests and you are scared to make new changes without them... then you can't push new fixes at all. Which is apparently fine for some organizations.
As with any hypothesis (empirical) testing, software testing is not about proving code is correct, but about striving to falsify that claim by poking at ways it would be likely to fail if it wasn't.
(There are formal analytical methods for proving code correct as opposed to empirical methods of falsifying it, as well, but those aren't tests.)
Tests exist to help prevent you from changing behavior unknowingly.
Tests can be a tool for documenting expected behaviour, intended use, and edge cases. In contrast to written documentation, they have to be up-to-date by their very nature, which is an advantage over code comments or external docs.
Also, I agree that's exactly what testing is, in fact I often now test code like that, get someone else to do one simple implementation in Python, a second high-performance implementation in C++/Rust, and compare them.
Then, as long as the same bug doesn't occur in both versions (and that seems very unlikely, as they are implemented with different algorithms), we find all the bugs in both.
This to me is the biggest thing that makes writing good tests a "grind". For a simple, well-defined function, you can toss up a bunch of inputs and outputs and feel satisfied that all the corner cases work. For any moderately complex behavior, however, so far the only effective way I've found to test the behavior is to effectively duplicate the logic a second time in the testing code, and check to see that I got the same result; which is pretty discouraging when you're doing it.
It's certainly nice once it's done, though, to be able to make a change, re-run your test suite, and have a pretty reasonable confidence that nothing broke.
"Sorting" should "preserve length" in {
forAll() { (l: List[Int]) => assert(sort(l).length == l.length) }
}
"Sorting" should "be idempotent" in {
forAll() { (l: List[Int]) => assert(sort(l) == sort(sort(l))) }
"Sorting" should "put elements in order" in {
forAll() { (l: List[Int]) =>
whenever(l.nonEmpty) {
sort(l).foldLeft((true, l.min))({
case ((result, max), elem) => (result && max <= elem, elem)
}) match {
case (result, max) => assert(result && max == l.max)
}
}
}
}
"Sorting" should "not change elements" in {
forAll() { (l: List[Int) => {
val sorted = sort(l)
assert(l.all(elem => sorted.contains(elem)))
assert(sorted.all(elem => l.contains(elem)))
}
}
Each of these tests can focus on one aspect of sorting and ignore everything else. Each on its own is not enough to give us confidence in the 'sort' function (e.g. the identity function would pass the idempotence test, the same-length test and the same-elements test; returning an empty list would pass the elements-in-order test; etc.), however together they are pretty good.In contrast, we can't implement the 'sort' function in such a piecemeal way. It needs to take every requirement into account, in case a step which satisfies one requirement invalidates the others. That's also why writing a whole new implementation to test against is best seen as a last resort.
That is not 'implement twice and hope you get it right at least once'. It isn't a perfect proof of correctness either, but it gives much more guarantees than a simple double implementation.
Notably, when my code to solve a linear system using LU decomposition failed, I new I didn't need to worry about my decomposition, because those tests still passed.
IBM team did 40% fewer defects than non-TDD team, Microsoft team: 60–90% (fewer defects than non-TDD fellows).
You can read more here: https://medium.com/crowdbotics/tdd-roi-is-test-driven-develo...
The usual wisdom is that the later a problem is detected, the more it costs. 15-33% extra time for removing ~half defects very early is a great trade-off!
I think the debate is not whether testing in general reduces defects (which is rather obvious), but if adhering to a stringent Test-Driven Development (TDD) process is better than regular testing (typically adding the tests after writing the code, instead of before).
> The experiment was conducted on “live” projects that the teams were developing at that moment.
There are plenty of experiments on those conditions. If they had consistent results, we could trust those, but they are all over the place, what means that confounding factors are more relevant than the thing being studied.
In web development it also makes back-end development faster. Switching to a browser, reloading the page and inspecting the console/properties/page, is so slow.
I wasn't really into testing until a certain point, but having the confidence that your business logic still works after doing change requests is a bliss.
No testing it in the browser, no testing it with postman, no checking if mails are send in your dummy smtp server interface etc.
When you don't write software that changes very often or has many use cases, go ahead and leave them out, but the more you take on your shoulders, the better automating the tedious work.
Test may serve as some kind of "persistent REPL" to make sure you didn't miss some edge case. But, at least to me, the real value of tests is confidence for refactoring.
Otherwise it is so easy to "simplify" a function but break code due to some implicit type cast, or making it work only for positive values, etc.
Given a complete test suite and stable interfaces, you can make substantial changes to existing code and if you break anything you'll know. It's often the difference between making changes that need to be made, and avoiding those changes because there's no safety net.
I'd be curious how the author approaches refactoring.
No, you can't. That's why when the bug occurs, you can add that test and then that bug never happens again.
A few years ago, I implemented tests in a Django website I manage on the side, and am around 90%+ "coverage". Since then, I've been able to upgrade major/minor versions, add new features, and have deployed around 100 updates. Automated testing before deploying has caught numerous small typos and errors before anything went live.
On the other hand, I'm also writing some games right now, and those don't have any tests. I can't see as clear of a reason to implement them.
Would others agree that certain areas of dev are more suited for testing than others?
It seems like testing in games is a huge problem, in that game QA testing is a big industry. I bet puzzle games and rpg games could gain a huge competitive advantage by testing their core domain. If they're long lasting games with lots of changes coming up.
I wouldn't agree at all with that. Code are rules written in way that can be translated by programs into a machine executable format.
Saying that code for some areas is more suitable for testing than code for other areas would imply that the rules are different somehow.
I don't see why that would be the case. I'd rather suggest that certain development processes and -methodologies (or lack thereof) may discourage systematic testing.
The rules that govern games are often expressed on a different level and by different people than the implementation. It's therefore hard to come up with reasonable assumptions to test in the first place.
Add to that the goals are often quite different (e.g. correct code vs the software doesn't crash and runs at an acceptable performance) and you may get the impression that tests just aren't well suited for say games.
In reality, the rules of the game, playability and fun need to be tested anyway. So from a cost-perspective it can make sense to skip out on automated code tests as errors and crashes would surface during gameplay tests anyway and overall correctness isn't a requirement for games anyhow.
When even minor differences can completely alter text flow you have to basically invert the testing process. You run the program, use your human brain to decide if it looks good, then take that output as the 'good' output for automated testing. That's in contrast to a lot of testing where you can set up the expected results as you write the tests, before the code.
So for me, that definitely falls under the domain of being less suitable for testing. A lot of the usual tricks don't work, and the benefits are less certain.
Since it seems like it'd be very similar in practice, I'm curious how browser devs test the visual aspect (not the parsing or DOM) of browser/CSS rendering - how do they decide the graphical result shown is 'good'?
It's comparing output rendered to bitmap images and verified on a per-pixel basis.
My other impression is that a lot of the unit testing advocates are coming from dynamic languages and need to runtime execute their code to find basic errors like typos. I know myself from writing in JS and Lua how much of a pain this is.
The last part is that gameplay code changes a lot, its very much experimentally driven. In my experience of following strict processes like TDD it becomes laborious, hurts iteration time and is error prone because rewriting tests often introduces bugs just by itself.
That said integration tests for bits and pieces of the game engine are gold.
We almost always write tests so that they pass - we should write tests so that if someone makes some change in the future that they fail easily.
Tests should be an aid, they should give you the heads up that something happened. They should be flags to help you not a game of achievement or failure.
We should call them traps. We make traps in code to trap bugs. Psychologically you can view them as games but avoid the negative connotations of failure.
Writing tests for the "happy path" is like confirmation bias: we only go looking for things which reinforce what we already think. Good experiments should challenge our assumptions.
One good way to write tests is when we're debugging: the bug itself falsifies our theory, so we can use that as a test (AKA a regression test). This requires we can reliably reproduce the bug, but that's usually an important first step when debugging.
Once we've got our regression test, we might wonder how this situation arose. This is a helpful way to tease out the assumptions we're making about the code, and turn them into tests. For example, the buggy result might be calculated from some intermediate values 'foo' and 'bar', but the bug makes no sense because 'foo' and 'bar' always satisfy certain properties that would prevent the bug. Well there's two new tests we can write! We can keep working back like this until we find the cause of the bug and fix it. I like this method because we end up with tests that correspond to those features of the code that we found ourselves doubting. That's usually a good sign that those tests are worth having.
So testing will definitely happen _at some point_. The debate is about where it happens: on the developer's computer, on a tester's computer, on a CI agent, or on the user's computer.
I like to cover as much code as possible before the software reaches the user, but that's just me.
I find that two many unit tests actually result in worse architected code as refactoring takes (at least) twice as long as you have to refactor the tests too. This can be mitigated a lot by seeing your tests also as code that needs to be architected well (shared code refactored out) but it is still a problem.
The flip side though is when you are writing code that is difficult or time consuming to get running. I am the lead for a telecomms system, the whole setup to get the code running requires hardware, signal generators and spectrum analysers. We have all that available on remote access but it can take time and fiddling for each code iteration. If I instead write automated tests at a couple of levels I can write a lot of code with confidence and then integrate with (i.e. test on) the hardware at the end with low risk of finding issues.
I had very similar experience when writing tests after.
And it turned around completely when I started writing tests first.
Now the tests are not aiming at specific properties of the code, nor are they duplicating the thought patterns in the code. They are simply examples of what the code should do, and as such are documentation and formal/informal specification. Formal in that code is a formal system, informal in that they aren't actually trying to be a complete specification.
So if you're suffering the same symptoms as the author, I suggest you try test first, it's a world of difference.
(And of course all the comments about safety for refactoring, programming over time etc.)
This matches my professional experience with writing and maintaining unit tests. They are often coded to the impl and are basically worthless IMO. I think you generally can get more bang for your buck with higher level functional tests, i.e., does this change preserve the functionality of the critical paths in an e2e flow?
First I declare a "foo.src.js" file to hold a module's logic, and then a test file as "foo.js" and have the test file import the logic and then export whatever is to be public. Client code imports the test file and not the source file, so the tests sit in between the logic and the client code, as it were. No more balancing testability against encapsulation.
The second is indirection. I create a mockability decorator. Then have a function (I call it 'mock') which can be used to temporarily redefine any function declared mockable, and then switch back to the real thing when the current test is through. So I get:
const readFile = mockable(fs.promises.readFile)
...
mock(readFile)(async () => 'mock file contents')
Combining these approaches doesn't so much prevent rework on tests when code changes as it makes it much less trouble to implement tests for 98% of cases.it amazes me we're doing the opposite of the best practices by piling everything unto the mythical full stack dev, while the rest of the world moves toward specialization
this focus on forcing devs on writing tests and learning to test better is highly inefficient for all parties involved, both in term of output quality and time commitment.
and the first thing that gets cut when under time pressure is, in fact, testing. and a dedicated tester is resistant to that.
the real thing is.. almost nobody want to pay top money for a specialist so nobody wants to specialize in the profession.
When that’s not in place, people often create selfish code where only their own needs are addressed.
They can still end up platform (or more generally, environment) specific, though.
I think he is missing the point. In tests I write I capture what users do so I can gain some trust before releasing new code, that old still behaves the same way. I also often debug my new code by writing tests, especially on systems where I cannot just run my code against.
If any of the following is true then you're writing brittle tests:
- Your test code looks like the code under test
- Your test suite includes examples that tweak one parameter and assert a result
- Your tests pass when you delete an arbitrary line of code from the implementation
I can't imagine shipping any significant project without any tests. How would I know that what I wrote implements my specifications faithfully? Hand waving and trust?
Unit tests are proof by example. They're trivial and don't prove the absence of errors. So I use them sparingly for simple, pure code where a few assumptions are enough.
Property tests are where I spend more of my time and focus. I generate the tests from a specification of a property I want to ensure will hold. It's not a formal proof that there are no errors but if my code survives 500 generated test cases and I have a good distribution over expected and hostile input I can be satisfied that my code is correct.
I don't spend time writing unit tests for effectful code. If I have business logic that is tied up calling a lot of APIs and digging into databases I defunctionalize the effect handling code and write interpreters for the data structures instead so that testing is still pure and easy and the effect handling code is constrained and found in one place. I test the pure version to makes sure the interpreter receives the correct sequence of data.
Sometimes I will use gold tests for serialization code. If I need to make sure that a contract I have with an external system isn't broken by my changes to the code I make sure to run the tests that will check what gets serialized out matches up with the golden examples.
If I need to ensure certain temporal properties hold like resource usage and performance... well I'll need regression and load tests.
These are all a part of the development process. You can ship code with no tests but good luck refactoring, maintaining, fixing, and understanding that code a year from now. It's possible, don't get me wrong, but it takes extreme discipline and even then that sometimes fails. People get tired, leave, get bored, etc. Having the computer check for you that everything is still working as expected is essential in my books.
Writing testable code is not about dependency injection etc. It's about having the behaviour of a unit clearly defined. This can be done explicitly in text, but also implicitly by using the component in a certain way. If anything beyond that behaviour is tested, you start to defining the behavior in the testcase. This is fine, but often happens unintentionally leading to a very weird and very specific definition of the behaviour.
For example, I work on a component that generates a configuration. I know that the order of lines of the output doesn't matter and should not matter. But the initial author wrote tests that fail if the order changes, because they just check the output against a string.
I'm curious since I think there is a lot of cross-talk when it comes to the terminology around testing.
There is definitely a lot of cross-talk. Can you make an example where my previous comment would be misguided?
When I read your final paragraph I realised that I've hit that same problem many times myself, e.g. checking if two things are equal, when in fact the order doesn't matter!
If you are implementing interfaces that plug into some other program (e.g. an IDE, or even different teams in a project), tests are invaluable to check that your code works as expected, and does not break when a new version of the program is released.
Yes, writing tests takes longer. Yes, writing tests makes it harder to change the code in the future. However, they have helped me identify cases in my code that I had not considered (when integrating my code with an IDE), and to prevent issues when refactoring and extending the code.
There are ways to mitigate the issues of changing the code, such as creating scaffolding to bridge the old and new code via adaptors, etc.
I pick the last sentence "If you like writing tests and are very good at it, you may continue to do so." This basically means, well "I don't like it, I'm not very good at it so I won't do it".
What I want to answer to that is that if you are not good at it, spend time to get better. Everyone can learn what is a unit test in 5 mins, writing good tests and good testable code takes years. Do tests get deprecated and useless, are they hard to maintain? Yes that's right, this happens. So don't aim for any coverage percentage, if you aim for 80% coverage but the last 30% are a pain to write, a pain to maintain and don't bring a lot a value then don't write them. Focus on the rest, make sure the core functionalities are well covered.
Unit Tests - This is usually per method and on a disconnected basis (you don't have a direct call to a resource on this)
Integrations - this is a combination of other unit tests. May use stubs and or mocks to isolate the external resources involved.
Functional/feature- This is where you get embedded services involved
Product acceptance test - This tends to test the individual application from deployment to running. This tends to verify outward features with live resources behind it. (Pipeline - it's one piece of the pipeline running live)
System test - this is where you can get live test environments setup and prepped for it. This tends to test the system as a whole (Pipeline- it's the entire pipeline in a test environment) Usual definitions are:
Smoke test - a briefer test of the System test.. but meant to prove out things work. Doesn't verify
When I am on a ruby console and I test manually, I should be able to create a test that will replicate the one done by hand. I should not have to learn a test framework.
For all for reasons I decided to only test real edge cases. Also I have now to create that magical gem that will generate automatly test.
In my experience, the "error detecting" value that a test sometimes gives you only occurs when you revisit code. Preventing regressions is very nice, but this is still not the main value of tests.
The real benefit of writing tests comes when you write them before you write any production code. That way the tests give you feedback about your software design. Code that is difficult to test is also difficult to use correctly.
They also complain about the friction you feel when you want to change behavior, but tests require the existing behavior. In this case the test should be changed or deleted. People are often too hesitant to delete tests.
Imagine that you're Microsoft, except you're against testing. You decide to refactor some Windows XP source a bit. What now? Just compile it and ship the result without even trying to boot it once?
Clearly that's unacceptable, so at least some testing must be okay in at least some scenarios. The question is how much and when. One or two black box tests of your program can get you a long way. Add more as needed to gain whatever degree of confidence is necessary. Remove tests that are redundant, fragile, or just exist reflexively because "code needs tests."
Take a nontrivial (but by no means enormous nor “messy”) system like the Roslyn C# compiler. It’s developed to a formal spec which is extremely rare, but it’s likeky almost impossible to add e.g a syntax change and foresee how that affects all scenarios. I can guarantee that any nontrivial change will break some tests which will force the author to iterate and solve the issues.
Most systems you encounter will be older, many will be much larger, messier, and not have the luxury of a spec that Roslyn has.
Of course archiving the tests can be a pain if nothing is prepared for it. Automating the tests is yet another pain. Still I worked on a library once when I added quite a few automated tests, because the infra was basically there and just waiting for it.
I prefer test harnesses and “monkey testing.”
I write about that, here: https://medium.com/chrismarshallny/testing-harness-vs-unit-4...
But I am REAL BIG on testing. Yeah, it ain’t fun, but I’ve been writing shipping software for most of my adult life, and, in my experience, about 60% of ship is boring stuff.
I also really did become a better programmer from them too. I used to think I knew what all the functions did and their requirements, but tests would often show me how a case a failed.
And of course if anything changes the test reveals this too.
I would love to have a decent solutions for unit tests on embedded systems with "bare-metal" firmware. There are some approaches, but the situation is certainly not optimal.
Also sums up my feeling of reading this author.
I only read the first two paragraphs and determined this article needs unit and integration tests.
Assert sentences:
Have noun, verb, and subject
Are not run-ons
Do not contain typos that render them unintelligibleBy testable code I mean code that's either easy to test or write tests for. I see a lot of tests retrofitted assuming that the code must not change. So my thumb rule is if it's hard to test then re-design it.
Not testing because it can't be perfect is a fallacy.
A good unit test should be written with the mindset of a hacker - how can I break it - rather than just a few assertions. It's not easy to do, and that's why there's so much bad stuff.
The article should be renamed to "Against Bad Testing".
Good tests do exist ... but they are hard to write (but easy to maintain). They require a lot more software design competence that writing non-testable production-code-that-works-today-and-maybe-tomorrow. An example of good tests: have a look at the Ninja codebase.
Writing such good tests for a legacy codebase (not written with testability in mind) is even harder (and I'm not sure that it's always possible in this case. It might. But I don't have any example).
Forcing everyone to write tests when they don't have the competence yet is a recipe for disaster and false conclusions ("this testing thing doesn't work / is harmful").
PS: now that I think of it, do we have a list of well-tested public-source codebases? this could be a good pedagogical tool.
I've found that I get better tests, and that I can write tests faster. It's also an antidote to Unangst's point that the developer who failed to imagine an test case while programming won't necessarily have an epiphany when they go to write tests.
But it isn't always a good fit with one's domain.
It seems like this should be googleable, but all I get when I try to find this is unit test frameworks and blog posts about how to unit test.
If you’re gonna be against testing, then be against it
For instance, if you are writing an application, then you probably don't need to be writing unit tests for internal APIs, since they might as well be considered private entities. Just write tests for application behavior. This stance seems rare, in my experience. Every team I've worked with insists on testing effectively private logic as well as testing the application at a high level, which takes a lot of time.
A lot of debugging time can be saved by guarding against unusual circumstances and providing useful error messages. Unfortunately, most people treat tests as if they are the documentation for how parts of the application are supposed to behave, which I think is generally the wrong way to look at it since tests can be difficult to decipher when you dive back in to them.
The greatest sin of testing, in my opinion, is the idea that if you write enough tests that you can avoid errors. The only way in which this works in some capacity is when you write your tests while you code(aka TDD). But the reality is that you simply can't avoid bugs no matter how meticulous you are in writing your tests. Every application I've worked on has ass loads of tests, and yet there are regressions every week. This is not the fault of anyone in particular, but the nature of the beast. In which case, you've got to decide whether it's worth testing all the minutiae of your application, or spend more time on the stuff that you really care about. Every test you add contributes to wasted developer time, especially when it gets to the point that it's no longer practical for programmers to run your entire test suite locally.
EDIT: To expand upon this, while I think integration tests are better than unit tests, generally speaking, I don't think they a good substitute for high level application tests when writing an application(not an API, framework, or library).
Your best bet for testing the truly desired behavior, getting an understanding of how performant your application is, and measuring your app's complexity, is application tests. If your application test requires a ton of ridiculous special cases to be set up, or faking test data becomes difficult, that's a sign that your app is too complicated and that you should stop and address that before anything else. If your application tests are becoming slow, that means that the application will become slower for your users. Integration and unit tests are unlikely to capture how your app is actually going to behave or perform. If you TDD your application while writing application tests, you will quickly know whether your work is improving the app or making it worse.
Appliction tests, by design, are much slower than low-level tests. This is a good thing, because you'd better make good choices or else your test are going to take forever to run. Things like unit tests sweep performance problems under the rug.
The problem with good application tests is that adding them to an existing app that performs poorly takes a tremendous amount of effort. If your app is well established, has tons of lower-level tests, but runs like crap and has a bunch of over-complicated inner workings, then getting new application tests in there will be painful and might not even be worth the effort to the business. Sadly, so many apps I've worked on faced a great deal of rot, which I think is largely a result of flawed testing philosophy, that made it very difficult to improve.
The reason to write a unit test _is to try to find bugs._
If your most effective way to find bugs is to test an internal API, then you should do that.
Based on the declining quality of everyday software it's not altogether surprising, but it's disturbing to see it written in black and white so brazenly.
He found issues in code you most likely rely on, like the Linux kernel : https://linux.kernel.narkive.com/Cw4EtIti/coverity-untrusted..., and also the pretty important CVE-2010-5298 in OpenSSL.
If only it had a unit test...