Multiple assertions are fine in a unit test
stackoverflow.blog
stackoverflow.blog
How is that even possible in the first place?
The entire job of an assertion is to wave a flag saying "here! condition failed!". In programming languages and test frameworks I worked with, this typically includes providing at minimum the expression put in the assertion, verbatim, and precise coordinates of the assertion - i.e. name of the source file + line number.
I've never seen a case where it would be hard to tell which assertion failed. On the contrary, the most common problem I see is knowing which assertion failed, but not how the code got there, because someone helpfully stuffed it into a helper function that gets called by other helper functions in the test suite, and the testing framework doesn't report the call stack. But it's not that big of a deal anyway; the main problem I have with it is that I can't gleam the exact source of failure from CI logs, and have to run the thing myself.
expected_positive = [
'abc',
'def',
...]
for text in expected_positive:
self.assertTrue(matcher(text), f"Failed: {text}")
Before I added the assertion error message, `f"Failed: {text}"`, it was quite difficult to tell WHICH example failed.But thanks for bringing it up, it seems it also has "subtest" support, which might be easier than interpolating the error message in some cases:
https://docs.python.org/3/library/unittest.html#distinguishi...
Unless you need the library style in order to drive it, switch to pytest. Seriously.
- assert rewriting is stellar, so much more comfortable than having to find the right assert* method, and tell people they're using the wrong one in reviews
- runner is a lot more practical and flexible: nodeids, marks, -k, --lf, --sw, ...
- extensions further add flexibility e.g. timeouts, maxfail, xdist (though it has a few drawbacks)
- no need for classes (you can have them, mind, but if functions are sufficient, you can use functions)
- fixtures > setup/teardown
- parameterized tests
Since you already have the object right there, why not do the latter approach?
All the other points you make as negatives are all positives for me. Biggest thing is, if you're making this change that alters things so drastically, is that really a good approach.
Also, fixtures aren't magic. If you can't scope the fixture to module or session, that means by default it runs in function scope, which would be the same thing as having expensive setup/teardown. And untangling fixtures can be a bgger PITA than untangling unexpected circular imports
I don't think this is a big problem; trying to focus on multiple examples at once is difficult.
It might be a problem if tests are slow and you are forced to work on all of them at once.
But in that case I'd try to make the tests faster (getting rid of network requests, disk/DB access by faking them away or hoisting to the caller).
I usually even comment out all the failing tests but the first one, after translating a bunch of specifications into tests, so I see the "green" when an example starts working.
Maybe it would be easier to grok multiple regex examples than algorithmic ones, but at least for myself, I am skeptical, and I prefer taking them one at a time.
Of course, ymmv depending on how good your initial intuition is, and how tricky the problem is.
With pytest you can use the -x flag to stop after the first test failure.
Even better you can use that in combination with -lf to only run the last failed test.
> Even better you can use that in combination with -lf to only run the last failed test.
Fwiw `--sw` is much better for that specific use-case.
`--lf` is more useful to run the entire test suite, then re-run just the failed tests (of the entire suite). IIRC it can have some odd interactions with `-x` or `--maxfail`, because the strange things happen to the cached "selected set".
Though it may also be because I use xdist a fair bit, and the interaction of xdist with early interruptions (x, maxfail, ...) seems less than perfect.
Another option is to use a custom mark on the test you want to run, and then do something like "pytest -v -m onlyrunthis"
for example,
expected_results = {...}
actual_obj = some_intance.method_call(...)
for key, val in expected_results.items():
assert getattr(actual_obj, key) == val, f"Mismatch for {key} attribute"
You could shift this off to a parametrized test, but that means you're making N more calls to the method being tested, which can have its own issues with cost of test setup and teardown. With this method, you see which key breaks, and re-run after fixing.also,
>your example stops on the first one that fails, and does not evaluate the others after that.
I would argue this is desirable behavior. there are soft checks, ie, https://pypi.org/project/pytest-check/, that basically replace assertions as raised exceptions and do your approach. But I do want my tests to raise errors at the point of failure when a change occurs. If there's alot of changes occurring, that raises larger questions of "why" and "is the way we're executing this change a good one"?
When you get many tests each emitting multiple failures because one basic thing broke, the output gets hard to sort through. It's easier when the failures are all eager.
They're used for code like this:
assert_eq(list.len(), 1)
expect_eq(list[0].username, "jdoe")
expect_eq(list[0].uid, 1000)
The idea being that if multiple properties are incorrect, then all of them will be printed out to the test log.If the list is empty, then the log will contain one error about the length being zero. If the list has one item but it has the wrong properties, the log will contain two errors.
See that's so completely unclear I utterly missed that there were two different calls there. Doesn't exactly help that the functions are the exact same length, and significantly overlap in naming.
That said, NUnit has much better syntax for this, where you put parallel/multiple assertions like this in an explicit block together: https://docs.nunit.org/articles/nunit/writing-tests/assertio...
Much less boilerplate duplication in the actual test framework, too.
It is.
> and you were reading code in context rather than on HN you'd have noticed.
(X) doubt
I envy you for never having seen tests atrocious enough where this is not only possible, but the common case.
Depending on language, framework and obviously usage, assertions might not be as informative as providing the basic functionality of failing the test - and that's it.
Now imagine this barebones use of assertions in tests which are entirely too long, not isolating the test cases properly, or even completely irrelevant to what's (supposedly) being tested!
If that's not enough, imagine this nightmare failing not after it has been written, but, let's say 18 months later, while being part of a massive test suite running for a while. All you have is a the name of the test that failed, you look into it to find a 630 lines long test "case" with 22 nondescript assertions along the way. You might know which line failed the test, but not always. And of course debugging the test function line by line doesn't work because the test depends on intricate timing for some reason. The person who wrote this might not be around and now this is your dragon to slay.
I think I should stop here before triggering myself any further. Therapy is expensive.
As for multiple asserts, that is really meaningless. The test case should test one thing. If it requires several asserts that's okay. But having a very long test function with a lot of assertions, is strongly indicating that you're testing more than one thing, and when the test fails it will be harder to know what actually happened.
The test failures I see that are actually hard to debug are ones where the failures are difficult to reproduce due to random input, tests running in parallel and sharing the same filesystem etc. I don't think I've ever not known what assert was failing (although I guess in theory you could make that happen by catching AssertionError).
There's nothing wrong with integration tests, but they're not unit tests. It's fine to have both, but the requirements for a good unit test and those for a good integration test diverge. The title of this post, at least, was specific to unit tests.
The longer I program the more I am convinced that the larger your unit the better. The unit tests is a statement that you will never refactor across this line, and that eliminates a lot of flexibility that I want.
It turns out that debugging failed integration tests is easy,the bug is in the last thing you changed. Sure the test covers hundreds of lines, but you only changed one.
I certainly don't see it as that. I see it as "this is the smallest thing I _can_ test usefully". Mind you, those do tend to correlate, but they're not the same thing.
Then you're testing useless things.
Usefulness is when different parts of a program work together as a coherent whole. Testing DB access layer and service layer separately (as units are often defined) has no meaning (but is often enforced).
Queue in memes about "unit tests with 100% code coverage, no integration tests" https://mobile.twitter.com/thepracticaldev/status/6876720861...
> Then you're testing useless things.
We'll have to agree to disagree then.
> Testing DB access layer and service layer separately (as units are often defined)
Not at all. For me, a unit is a small part of a layer; one method. Testing the various parts in one system/layer is another type of test. Testing that different systems work together is yet another.
I tend to think in terms of the following
- Unit test = my code works
- Functional test = my design works
- Integration test = my code is using your 3rd party stuff correctly (databases, etc)
- Factory Acceptance Test = my system works
- Site Acceptance Test = your code sucks, this totally isn't what I asked for!?!
The "my code works" part is the smallest piece possible. Think "the sorting function" of a library that can return it's results sorted in a specific order.
If those fail, it means that neither your design nor your code works.
The absolute vast majority of unit tests are meaningless because you just repeat them again in the higher level tests.
Why would they? Do these edge cases not appear when the caller is invoked? Do you not test these edge cases and the behavior when the caller is invoked?
As an example: you tested that your db layer doesn't fail when getting certain data and returns response X (or throws exception Y). But your service layer has no idea what to do with this, and so simply fails or falls back to some generic handler.
Does this represent how the app should behave? No. You have to write a functional or an integration test for that exact same data to test that the response is correct. So why write the same thing twice (or more)?
You can see this with Twitter: the backend always returns a proper error description for any situation (e.g. "File too large", or "Video aspect ratio is incorrect"). However, all you see is "Something went wrong, try again later".
> It's much simpler to thoroughly test each unit. Then test how they work together.
Me, telling you: test how they work together, unit tests are usually useless
You: no, this increases the number of tests. Instead, you have to... write at least double the amount of tests: first for the units, and then test the exact same scenarios for the combination of units.
----
Edit: what I'm writing is especially true for typical microservices. It's harder for monoliths, GUI apps etc. But even there: if you write a test for a unit, but then need to write the exact same test for the exact same scenarios to test a combination of units, then those unit tests are useless.
Unit two - calls unit one - test that, if unit one returns an error, it is treated appropriately. One test, covers all error conditions because they're all returned the same way from Unit one.
Unit three - same idea as unit one
If you were to test the behavior of unit one _through_ units 2 and 3, you'd need 2*N tests. If you were to test the behavior of unit one separately, you'd need N+2 tests.
You're missing the point that you don't need to test "the exact same scenarios for the combination of units", because the partitions of <inputs to outputs> is not the same as the partitions for <outputs>. And for each unit, you only need to test how it handles the partitions of <outputs> for the items, it calls; not that of <inputs to outputs>.
There are only two possible responses to that:
1. No, there are not 2*N tests because unit 3 does not cover, or need, all of the behavior and cases that flow through those units. Then unit testing unneeded behaviors is unnecessary.
2. Unit 3 actually goes through all those 2*N cases. So, by not testing them at the unit 3 level you have no idea that the system behaves as needed. Literally this https://twitter.com/ThePracticalDev/status/68767208615275315...
> You're missing the point that you don't need to test "the exact same scenarios for the combination of units", because the partitions of <inputs to outputs>
This makes no sense at all. Yes, you've tested those "inputs/outputs" in isolation. Now, what tests the flow of data? That unit 1 outputs data required by unit 2? That unit 3 outputs data that is correctly propagated by unit 2 back to unit 1?
Once you start testing the actual flow... all your unit tests are immediately entirely unnecessary because you need to test all the same cases, and edge cases to ensure that everything fits together correctly.
So, where I would write a single functional test (and/or, hopefully, an integration test) that shows me how my system actually behaves, you will have multiple tests for each unit, and on top of that you will still need a functional test, at least, for the same scenarios.
You don't, but it's clear that I am unable to explain why to you. I apologize for not being better able to express what I mean.
If you don't, then you you have no idea if your units fit together properly :)
I've been bitten by this when developing microservices. And as I said in an edit above, it becomes less clear what to test in more monolithic apps and in GUIs, but in general the idea still holds.
Imagine a typical simple microservice. It will have many units working together:
- the controller that accepts an HTTP request
- the service layer that orchestrates data retrieved from various sources
- the wrappers for various external services that let you get data with a single method call
- a db wrapper that also lets you get necessary data with one method call
So you write extensive unit tests for your DB wrapper. You think of and test every single edge case you can think of: invalid calls, incomplete data etc.
Then you write extensive unit tests for your service layer. You think of and test every single edge case you can think of: invalid calls, external services returning invalid data etc.
Then you write extensive unit tests for your controller. Repeat above.
So now you have three layers of extensive tests, and that's just unit tests.
You'll find that most (if not all) of those are unnecessary for one simple reason: you never tested how they actually behave. That is, when the microservice is actually invoked with an actual HTTP request.
And this is where it turns out that:
- those edge cases you so thoroughly tested for the DB layer? Unnecessary because invalid and incomplete data is actually handled at the controller layer, or service layer
- or that errors raised or returned by service wrappers, or the db layer either don't get propagated up, or are handled by a generic catch all so that the call returns a nonsensical stuff like `HTTP 200: {error: "Server error"}`
- or that those edge cases actually exist, but since you tested them in isolation, and you didn't test the whole flow, the service just fails with a HTTP 500 error on invalid invocation
Or, instead, you can just write a single suite of functional tests that test all of that for the actual controller<->service<->wrappers flow covering the exact same scenarios.
After much googling and buying or ancient text books I hit a dead end. At this point I think "unit" is just noise that confuses people into making distinctions that don't exist.
Which is obviously not what people really mean these days, but the phrase stuck. The early Xp'ers even found it an issue back then.
For a while people tried to push the term "micro tests", but that didn't really take off.
I agree with Gerard Mezaros and Martin Fowler and typically follow their (very mainstream) definitions on this stuff. Integration and functional testing have their own ambiguities too, it's definitely a frustrating situation to not have solidly defined foundational terms.
The Web API has a well-defined and documented interface. Is it a “unit”?
Think in terms of business rules, not the code structure: What's one thing your API does?
This is roughly my definition of unit test: "tests run in a single process"
That’s not true.
A correct change might expose an existing bug which hadn’t been tested or expose flaky behavior which existed but hadn’t been exercised. In both cases the solution is not to revert the correct change, but to fix the buggy behavior.
REQUIRES(v1::foo(0) == v2::foo(0));
REQUIRES(v1::foo(1) == v2::foo(1));
And the second assert fails the error message will tell me exactly that, the line, and the value of both function calls if they are printable. What more do you want to know "what actually happened"?This should not ever be possible in any semi-sane test environment.
One could in theory write a single test function with thousands of asserts for all kinds of conditions and it still should be 100% obvious which one failed when something fails. Not that I'd suggest going to that extreme either, but it illustrates that it'll work fine.
Yeah, and if you write one assertion at a time, it will be harder to write the tests. Decreasing #assertions/test decreases the speed of test debugging while increasing the time spent writing non-production code. It's a tradeoff. Declaring that the optimal number of assertions per test is 1 completely ignores the reality of this tradeoff.
As for me, I tend to write reasonable tests and cover several cases that guard the intended behavior of each function (if someone decides the function should behave differently in the future, a test should fail). One emerging pattern is that sometimes during testing I realize I need to refactor something, which might have been lost on me if I skimmed on tests. It's both a sanity check and a guardrail for future readers.
Patterns are context specific advice about a solution and its trade-offs, not hard rules for every situation.
Notice how pattern books will often state the relevant Context, the Problem itself, the Forces influencing it, and a Solution.
(This formulaic approach is also why pattern books tend to be a dry read).
You're falling prey to slippery slope fallacy, which at best is specious reasoning.
The rationale is easy to understand. Running 100 assertions in a single test renders tests unusable. Running 10 assertions suffers from the same problem. Test sets are user-friendly if they dump a single specific error message for a single specific failed assertion, thus allowing developers to quickly pinpoint root causes by simply glancing through the test logs.
Arguing whether two or three or five assertions should be banned misses the whole point and completely ignores the root cause that led to this guideline.
As if this actually happens in practice, regardless of multiple or single asserts. Anything that isn't non-trivial will at most tell you what doesn't work, but it won't tell you why it doesn't work. Maybe allowing an educated guess when multiple tests fail to function.
You want test sets to be user friendly? Start at taking down all this dogmatism and listening to the people as to why they dislike writing tests. We're pushing 'guidelines' (really more like rules) while individuals think to themselves 'F this, Jake's going to complain about something trivial again, and we know these tests do jack-all because our code is a mess and doing anything beyond this simple algorithm is a hell in a handbasket".
These discussions are beyond useless when all people do is talk while doing zero to actually tackle the issues of the majority not willing to write tests. "Laziness" is a cop-out.
If that's the case, the test framework itself is severely flawed and needs fixing even more than the tests do.
There's no excuse to have an assert function that doesn't print out the location of the failure.
Granted, this is not the way to do things. But it happens anyway.
The comment thread is about “Assertion Roulette” — having so many assertions you don’t know which went off. Which really seems like a test framework issue more than a test issue.
Couldn't not read that in Peter Sellers' voice https://m.youtube.com/watch?v=2yfXgu37iyI&t=2m36s
> ... to find a 630 lines long test "case" with 22 nondescript assertions along the way.
This is where tech team managers are abrogating their responsibility and job.
It's the job of the organization to set policy standards to outlaw things like this.
It's the job of the developer to cut as many corners of those policies as possible to ship code ASAP.
And it's the job of a tech team manager to set up a detailed but efficient process (code review sign offs!) that paper over the gap between the two in a sane way.
... none of which helps immediately with a legacy codebase that's @$&@'d, though.
I can't tell if this is supposed to be humor, or if you actually believe it. It's certainly not my job as a developer to ship worse code so that I can release it ASAP. Rather, it's my job to push back against ASAP where it conflicts with writing better code.
And furthermore, you are not the developer most non-tech companies want.
Those sorts of companies want to lock the door to the development section, occasionally slide policy from memos under the door, and get software projects delivered on time, without wasting any more thought on how the sausage gets made.
Well, avoiding ecosystems where people act dumb is a sure way to improve one's life. For a start, you won't need to do stupid things in reaction of your tools.
Yes, it's not always possible. But the practices you create for surviving it are part of the dumb ecosystem survival kit, not part of any best practices BOK.
In Go, in general, tests fail and continue, rather than causing the test to stop early, so you can tell which of those 22 checks failed. Other languages may have the option to do something similar.
it('fails', async () => {
expect(await somePromise()).toBe(undefined)
expect(await someOtherPromise()).toBeDefined()
})
No chance figuring out where it failed, it's likely just gonna run into a test suite timeout with no line reference or anything.So essentially you could write
Assert(()=>width>0 && x + width < screenWidth)
And you would get: Assertion failed:
x is 1500
width is 600
screenWidth is 1920
It used Expression<T> to do the magic. Amazing debug messages. No moralizing required.This was a huge boon for us as it was a legacy codebase and we ran tens of thousands of automated tests and it was really difficult to figure out why they failed.
It's incredibly useful once you know how, and encourages you to stop using reflection, passing magic strings as arguments, or having to use nameof().
"GetProductPage().GetProductPrice().Should().Be().LessThan(...)"
If you just want the assignments then it's simpler:
Add to the Evaluate method a test for MemberExpression and then:
The variable name is:
((MemberExpression)expr).Member.Name
The value is:
Expression.Lambda(expr).Compile().DynamicInvoke()
I like to use Jest’s toMatchObject to combine multiple assertions in a single assertion. If the assertion fails, the full object on both sides is shown in logs. You can easily debug tests that way.
The only way to make it even possible is to do some eval magic or to use a pre-processor like babel or a typescript compiler plugin.
But if you find something, Lemme know.
You can inspect the source, but functions don’t run in isolation so you cannot extract the source and run it somewhere else.
Simple example:
var x = 0;
var y = function() { x++; };
You can’t reliably execute function y with just y’s function body.Maybe if you replace the function with another by injecting code into a copy of the source, but I’m not sure if that’s even possible.
expect(width).toBeGreaterThan(0); expect(x+width).toBeLessThan(screenWidth);
When an assertion fails it will tell you: "Expected: X; Received: Y"For example:
public void CheckIsTrue(bool value, [CallerArgumentExpression("value")] string? expression = null)
{
if (!value)
{
Debug.WriteLine($"Failed: '{expression}'");
}
}
So if you call like this: CheckIsTrue(foo != bar && baz == true), when the value is false it prints "Failed: 'foo != bar && baz == true'".[1] https://learn.microsoft.com/en-us/dotnet/csharp/language-ref... [2] https://learn.microsoft.com/en-us/dotnet/csharp/language-ref...
test <@ width > 0 && x + width < screenWidth @>
And part of the output is: width > 0 && x + width < screenWidth
500 > 0 && x + width < screenWidth
true && x + width < screenWidth
true && 1500 + 500 < 1920
true && 2000 < 1920
true && false
false
[0]: https://github.com/SwensenSoftware/unquoteMany times I searched for 'how do I do something with xUnit' and found a github issue with people struggling with the same thing, and the author flat out refusing to incorporate the feature as it was against his principles.
Other times I found that I needed to do was override some core xUnit class so it would do the thing I wanted it to do - sounds complex, all right lets see the docs. Oh there are none, 'just read the source' according to the author.
Another thing that bit us in the ass is they refused to support .NET Standard, a common subset of .NET Framework and Core, making migration hell.
"We can't land this tiny fix because we have a release planned within three months" sort of thing.
One of the key differences, is there are no test cases generated, nUnit tests just wont run, while xUnit will throw.
Since it's completely legal and sensible for a certain kind of test to have no entries in an xml, we needed to hack around this quirk. When I (and countless others) have mentioned this on the xunit github, the author berated us that how dare we request this.
So nUnit might be buggy, but xUnit is fundamentally unfixable.
The problem is, since it's a legacy codebase, there's many fields which are only tested incidentally by this behaviour, by tests that actually aren't intending to test that functionality.
Then, I packaged it into a Google Test matcher, and from now on, the problem you describe is gone. I write:
EXPECT_THAT(someObject, IsEqAsJSON(someBlobFromADifferentFile));
and if it fails, I get output like this: Expected someObject to be structurally equivalent to someBlobFromADifferentFile; it is not;
- #/object/key - missing in expected, found in actual
- #/object/key2 - expected string, actual is integer
- #/object/key3/array1 - array lengths differ; expected: 3, actual: 42
- #/object/key4/array1/0/key3 - expected "foo" [string], actual "bar" [string]
Etc.It was a rather simple exercise, and the payoff is immense. I think it's really important for programmers to learn to help themselves. If there's something that annoys you repeatedly, you owe it to yourself and others to fix it.
It's a cultural problem. _I_ can do that, but my colleagues will just continue to write minimum effort tests against huge json files or database dumps where you have no idea why something failed and why there are a bunch of assertions against undocumented magic numbers in the first place. It's like you're fighting against a hurricane with a leaf blower. A single person can only do so much. I end up looking bad in the daily standup because I take longer to work on my tickets but the code quality doesn't even improve in a measurable way.
Assuming that’s not the case and you’re interested in the state of two object graphs then you just compare those, not the json string they deserialize to.
https://github.com/skyscreamer/JSONassert seems decent.
but it can be done from scratch in a few hours (I'd recommend this if you have 'standardized' fields which you may want to ignore):
Move to a matcher library for assertions (Hamcrest is decent), and abstract `toJSON` into the a matcher, rather on the input.
This would change the assertion from:
`assertEquals(toJson(someObject), giantJsonBlobFromADifferentFile)`
to:
`assertThat(someObject, jsonEqual(giantJsonBlobFromADifferentFile))`
The difference here is subtle: it allows `jsonEqual` to control the formatting of the test failure output, so on a failure you can:
* convert both of the strings back to JSON
* perform a diff, and provide the diff in the test output.
Decent blog post on the topic: https://veskoiliev.com/use-custom-hamcrest-matchers-to-level...
I'd actually rather people just assert on the fields they need, but it's a larger team than I can push that on.
Hamcrest is a drop-in replacement for `assertEquals`, and provides obvious benefits. Politically, it's easy to convince developers onboard once you show them:
* You just need to change the syntax of an assertion - no thought required
* You (Macha) will take responsibility for improving the formatting of the output, and developers have someone to reach out to to improve their assertions.
From this: you'll get a very small subset of missionaries who will understand the direction that you're pushing the test code in, and will support your efforts (by writing their own matchers and evangelising).
The larger subset of the developer population won't particularly care, but will see the improved output from what you're proposing, and will realise that it's a single line of code to change to reap the benefits.
EDIT: I've added a lint rule into a codebase to guide developers away from `assertEquals()`. Obviously this could backfire, and don't burn your political capital on this issue.
There are a set of unit testing frameworks that do everything they can to hide test output (junit), or vomit multiple screens of binary control code emoji soup to stdout (ginkgo), or just hide the actual stdout behind an authwall in a uuid named s3 object (code build).
Sadly, the people with the strongest opinions about using a "proper" unit test framework with lots of third party tooling integrations flock to such systems, then stack them.
I once saw a dozen-person team's productivity drop to zero for a quarter because junit broke backwards compatibility.
Instead of porting ~ 100,000 legacy (spaghetti) tests, I suggested forking + recompiling the old version for the new jdk. This was apparently heresey.
Basically it's nearly impossible to fully parse it correctly.
If you have to pick one or the other, then you're breaking the common flow (human debugging code before pushing) so that management can have better reports.
The right solution would be to add a environment variable or CLI parameter that told tap to produce machine readable output, preferably with a separate tool that could convert the machine readable junk to whatever TAP currently writes to stdout/stderr.
But unlike TAP, it's fairly Perl-specific as opposed to just being an output format. I imagine you could adapt the ideas in it to Node but it'd be more complex than simply implement TAP in JS.
And yes, I think the idea of having different output formats makes sense. With Test2, the test _harness_ produces TAP from the underlying machine-readable format, rather than having the test code itself directly product TAP. The harness is a separate program that executes the tests.
Nothing should have to be parsed. Write test results to sqlite, done. You can generate reports directly off those test databases using anything of your choice.
your-program test-re.sqlite output.htmlhttps://web.archive.org/web/20110114031716/https://browserto...
- Make it possible to disable timeouts. Otherwise, people will need a different runner for integration, long running (e.g., find slow leaks), and benchmark tests. At that point, your runner is automatically just tech debt.
- It is probably possible to nest before and afters, and to have more than one nesting per process, either from multiple suites, or due to class inheritance, etc. Now, you have a tree of hooks. Document whether it is walked in breadth first or depth first order, then never change the decision (or disallow having trees of hooks, either by detecting them at runtime, or by picking a hook registration mechanism that makes them inexpressible).
On the one hand, running tests in any order should produce the same result, and would in any decent test suite.
On the other hand, if the order is random or nondeterministic, it's really annoying when 2% of PRs randomly fail CI, not because of any change in the code, but because CI happened to run unrelated tests in an unexpected order.
Therefore the tool should run the tests in random order, to flush out the non-decent tests. IMHO.
I sometimes run into issues not so much due to order dependencies specifically, but due to tests running in parallel sometimes causing failures due to races. It's almost always been way more work to convert a fully serial test suite into a parallel one than it is to just write it that way from the start, so I think there's some merit in having test frameworks default to non-deterministic ordering (or parallel execution if that's feasible) with the ability to disable that and run things serially. I'm not dogmatic enough to think that fully parallel/random order tests are the right choice for every possible use case, but I think there's value in having people first run into the ordering/race issues they're introducing before deciding to run things fully serially so that they hopefully will consider the potential future work needed if they ever decide to reverse that decision.
The solution is to use random ordering and print the ordering seed with each run so it can be repeated when it triggers an error. Immediately halt all new work until randomly run tests don't have problems.
This isn't as bad as it sounds, generally it's a few classes of things that cause the interference which each will fix many tests. It's unlikely that the code actually has a 2%+ density of global-variable use, for example.
This way the framework can kill the test if it doesn't see one of these messages in a short amount of time.
This fixed our timeout issues. We had tests that took too long, specially in debug builds and we'd end up having to set too large a timeout. Now though, we can keep the timeout for the heartbeat really short and our timeout issues have mostly gone away.
The problems isn’t with breaking rules. The problem is with promising yourself or others that you will fix it “later” and then breaking that promise.
I was a TL on a project and I had two "eng" on the project that would make test with a single method and then 120 lines of Tasmanian Devil test cases. One of those people liked to write 600 line cron jobs to do critical business functions.
This scarred me.
I was a long-time maintainer of Debian's cron, a fork of Vixie cron (all cron implementations I'm aware of are forks of Vixie cron, or its successor, ISC cron).
There are a ton of reasons why I wouldn't do this, the primary one being is that cron really just executes jobs, period. It doesn't serialize them, it doesn't check for load, logging is really rudimentary, etc.
A few years ago somebody noticed that the cron daemon could be DoS'ed by a user submitting a huge crontab. I implemented a 1000-line limit to crontabs thinking "nobody would ever have 1000-line crontabs". I was wrong, quickly received bug reports.
I then increased it to 10K lines, but as far as I recall, users were hitting even that limit. Crazy.
There indeed exist a few non-Vixie-cron-derivative implementations but as far as I'm aware, all major Linux and BSD distributions use a Vixie cron derivative.
Edit: I see now where I caused confusion. In my original post, I should have said all default cron implementations.
It used to be the major distribution. Funny how times change.
The test runner in VS2019 does this, too and it's incredibly frustrating. I get to see debug output about DLLs loading and unloading (almost never useful), but not the test's stdout and stderr (always useful). Brilliant. At least their command line tool does it right.
Assertions throw an exception and the test runner catches them along with any exceptions thrown by the code in test, marks the test as a failure, and reports the given error message and a full stack trace.
When every test calls ->run_base_tests() before running it's own assertion sometimes things fail before you get to the root cause assertion.
The other problem of stacking assertions is that you'll see the first failure only. There may be more failures that give you a better picture of what's happening.
Having each assertion fail separately gives you a clearer picture of what's going wrong.
Fwiw, the book doesn't suggest what the reader is saying. It says what I've said above more or less.
You don't need to always stick to the rule but it generally does improve things to the point I now roll my eyes when I come across tests with stacked assertions and lots of test harness code that runs it's own assertions, I just know I'm in for a fun time.
Aside from that, you could do things like
int failures = 0;
failures += someAssertion ? 0 : 1;
failures += anotherAssertion ? 0 : 1;
failures += yetAnotherAssertion ? 0 : 1;
assertZero(failures);
With the appropriate logging of each assertion in there.Consider the situation of "I've got an object and I want to make sure it comes out in JSON correctly"
The "one assertion" way of doing it is to assertEqual the entire json blob to some predefined string. The test fails, you know it broke, but you don't know where.
The multiple assertions approach would tell you where in there it broke and the test fails.
The point is much more one of "test one thing in a test" but testing one thing can have multiple assertions or components to it.
You don't need to have testFirstNameCorrect() and testLastNameCorrect() and so on. You can do testJSONCorrect() and test one thing that has multiple parts to verify its correctness. This becomes easier when you've got the frameworks that support it such as the assertAll("message", () -> assertSomething, () -> assertSomethingElse(), ...)
Secretly, they would have the reader write 32 functions to separately test every bit of a uint32 calculation, only refraining from that advice due to the nagging suspicion that it might be loudly ridiculed.
That doesn't mean you have to duplicate code, you can deal with it in other ways. In Junit I like to use @TestFctory [1] where I'll write most of the test in the Factory and then each assertion will be a Test the factory creates, and since they're lambdas they have access to the TestFactoy closure.
[1] https://junit.org/junit5/docs/current/user-guide/#writing-te...
(I wrote Foundations of Programming for any 2000s .NET developer out there!)
To hear this is still a fight people are having...It really makes me appreciate the value of having deep experience in multiple languages/communities/frameworks. Some people are really stuck in the same year of their 10 (or 20, or 30) years of experience.
Holy moly! Think I still have your book somewhere. So thank you for that.
In my last +10 years of .net development I haven't heard anything about single-assertion.
> Some people are really stuck in the same year of their 10 (or 20, or 30) years of experience.
I think this has manifested even more with the transition into .net core and now .net 5 and beyond. There are so many things changing all the time (not that I complain), which can make it difficult to pick up what's the current mantra for the language and framework.
And to the author: Your bubble is significantly different from mine. Pretty much every competent developer I've worked with would laugh at you for the idea that the second test case would not be perfectly fine. (But that first iteration would never pass code review either because it does nothing and thus is a waste of effort.)
expect { foo.call() }.to change { bar.value }.by(2)
That is, regardless of the absolute value of bar.value, I expect foo.call() to increment it by 2.
The point of the 1 assertion per test guideline is to end up with tests that are more focused. Giving that you did not seem to think of the above technique, I'd say that this guideline might just have helped you discover a way to write better specs ;-)
Guidelines (that is, not rules) are of course allowed to be broken if you have a good reason to do so. But not knowing about common idioms is not a good reason.
You might argue that the above code is just sugar for 2 assertions, but thats beside the point: The test is more focused, there -appears- to be only one assertion, and thats what matters.
Any rule that says there should be only 1 assertion ever is stupid.
The only reason I can really see to have more than one assertion would be to avoid having to run the setup/teardown multiple times. However, its usually a desirable goal to write code that require little setup/teardown to test anyways because that comes with other benefits. Again, it might not be practical or even possible, but that goes of almost all programming "rules"..
So you have one test that indicates that a log error is outut. then another that tests that the property X in the return from the error is what you expect. then another test to determine that propery Y in return is what you expect?
that to me is wasteful, unclear, bloated. About the only useful result I can see that is it allows bragging about how many tests a project has.
That's one way to get dogma-driven assertion roulette, as you will not know which particular error occurred.
assertThat(fooReturningOptional(), isPresentAndIs(4))
Or
assertThat(shape, hasAreaEqualTo(10))
Or
AssertThat(polygonList, hasNoIntersections())
Custom matchers can go off the deep end really easily. One of those cases of learn the principle, then learn when it does not apply
Can always refactor to 2 tests if the test setup changes and the assertions begin to differ or become too complex.
That being said, there can of course me other tradeoffs e.g. performance and even cases where simple test setups are downright impossible.
I find that reducing assertions per spec where I can a good guideline. E.g. combining expect(foo['a']).to eq(1) and expect(foo['b']).to eq(2) into expect(foo).to include('a' => 1, 'b' => 2) yields better error messages.
Invariants would either have to be publically available and thus easily testable with similar methods, or, one would have to use assertions in the implemention.
I try to avoid the latter, as it mixes implemations and 'test/invariants'. Granted, there are situations (usually in code that implements something very 'algorithm'-ish) where inline assertions are so useful that it would be silly to avoid them. (But implementing algos from scratch is rare in commercial code)
foo.call() might have a return value.
Also, the whole story invocation shouldn't throw an exception, if your language has them. This assertion is often implied (and that's fine), but it's still there.
Finally the test case is a little bit stupid, because very seldom code doesn't have any input that changes the behavior/result. So your assertion would usually involve that input.
If you follow that though consequently, you end up with property-based tests very soon. But property-based tests should have as many assertions as possible for a single point of data. Say you test addition. When writing property-based tests you would end up with three specifications: one for one number, testing the identity element and the relationship to increments. Another one for two numbers, testing commutativity and inversion via subtraction, and one for three numbers, testing associativity. In every case it would be very weird to not have all n-ary assertions for the addition operation in the same spot.
test "pressing d key makes mario move 2 pixels right" {
expect { keyboard.d() }.to change { mario.x }.by(2)
}
I could test the value of the d() function, but I dont because I don't care what it returns.
Didnt understand the "whole story invocation" and exception part, am I missing some context?
Sure property-based testing can be invaluable in many situations. Only downside is if the tests become so complex to reason about that bugs become more likely in the tests than the implemenation.
I've sometimes made tests with a manual list of inputs and a list of expected outputs for each. I'd still call that 1 assertion tests (just run multiple times), so my definition of 1 assertion might too broad..
Near as I can tell, many people are made uncomfortable by this in practice because these tests feel childish and dare I say demeaning. So they try to do something “sophisticated” instead which is a slow and lingering death where tests are co corned.
Lacking self consciousness, you can whack out hundreds of unit tests in a couple of days, and rewrite ten of someone else’s for a feature or a bug fix. That’s fine and good.
But when your test looks like an integration test, rewriting it misses boundary conditions because the test is t clear about what it’s doing. And then you have silent regressions in code with high coverage. What a mess.
Testing (especially unit) is an area of tech weirdly with a lot of dogmatism. I think Uncle Bob is the source of some of it.
That first iteration would not be subject to code review. The author is using TDD.
If you painstakingly craft a scenario where you create a rectangle of a specific expected size why wouldn’t it be acceptable to assert both the width and height of the the rectangle after you have created it?
assert_equal(20, w, …
assert_equal(10, h, …
A dogmatic rule would just lead to an objectively worse test where you assert an expression containing both width and height in a single assert?
assert_true(w == 20 && h == 10,…)
So I can only assume the rule also prohibits any compound/Boolean expressions in the asserts then? Otherwise you can just combine any number of asserts into one (including mutating state within the expression itself to emulate multiple asserts with mutation between)!
That's what's bound to happen under that rule. People just start writing their complex tests in helper functions, and then write
assert_true(myTotallySimpleAndFocusedTestOn(result))The part that is glossed over is that the test suite takes several hours to run on your machine, so you delegate it to a CI pipeline and then fork out for parallel execution (pun intended) and complex layers of caching so your suite takes 15 minutes rather than 2 and a half hours. It’s monumentally wasteful and the tests aren’t any easier to follow because of it.
The suite doesn’t have to be that slow, but it’s inevitable when every single assertion requires application state to be rebuilt from scratch, even when no state is expected to change between assertions, especially when you’re just doing assertions like ‘assert http status is 201’ and ‘assert response body is someJson’.
I came up in ruby, heard this, and quickly decided it was stupid.
Enginnering decisions have tradeoffs. When the testsuite becomes too slow, it might be time to reconsider those tradeoffs.
Usually though, I find that to road to fast tests is to reduce/remove slow things (almost always some form of IO) not to combine 10 small tests into one big.
It just so happens that your tests become IO bound because those small tests in aggregate hammer your DB and the network purely to set up state. So if you only do it once by being more deliberate with your tests, you’re in a better place.
I can't speak for Ruby, but what I would call 'clean' and happily dogmatise is that assertions should come at the end, after setup and exercise.
I don't care how many there are, but they come last. I really hate tests that look like:
setup()
exercise(but_not_for_the_last_time)
assert state.currently = blah
state.do_other()
something(state)
assert state.now = blegh
And so on. It stinks of an integration test forced into a unit testing framework.I like them to look like:
foo = FooFactory(prop="whatever")
result = do(foo)
assert result == "bar"
I.e. some setup, something clearly under test, and then the assertion(s) checking the result.There’s no avoiding it though when you want something end-to-end, or a synthetic test. You’re piling up a whole succession of stateful actions and if you tested them in isolation you would fail to capture bugs that depend on state. In that sense, better to run a ‘signup, authenticate and onboard’ flow in one test instead of breaking it down.
setup()
exercise(but_not_for_the_last_time)
was_blah = (state.currently == blah)
state.do_other()
something(state)
assert was_blah && state.now == blegh
In fact, this last version is worse, because if do_other() can fail if state wasn't blah, then what you'll get is the exception from that failure interrupting the test before the assert would have been reported.Applying any rule dogmatically is often bad, and this is no exception. The is that we don’t like lacking rules. Especially it goes to hell when people start adding code analysis to enforce it, and then developers start writing poor code that passes the analysis.
One assert imo isn’t even a good starting point that might need occasional exceptions.
I disagree, it can be a good starting point for most cases. You should be able to condense your test in 3 steps (arrange, act, assert) each one a single line, and even if you do not do it because setup is too complicated, it's not worth it for that single test, you assert more things etc., I think the mental exercise to try to think: "can this be made in a 3 line test" is invaluable in writing maintainable tests.
> is that then one assert logically for any N? At what point does it stop?
This is one of the hard things about good tests, they are a little bit of art too. Maybe you can apply the single responsibility like I said before: the test should change for one reason only. By one reason I mean one "person/role": it should change if the CFO of our clients want something differet, or if Mark from IT wants some change.
I am not stressing or enforcing single asserts too much, I feel like tests allow a little bit of leeway in many ways, as long as the decision enhances expressiveness. If the extra lines are not making the test clearer, if a single assert would be clearer for the story that the test is telling, then it should go into a single assert. If I can break the story into multiple stories that still make sense, then I do that, such that each story has it's own strong storyline.
So instead of:
- error W doesn't equal 20
Fix that
Run test again
- error H doesn't equal 10
Fix that
Run test again
It's - Error Width doesn't equal 20
- Error Height doesn't equal 10
Fix both
Run test
I think the time savings are negligible though. And it makes testing even more tedious, as if people needed any additional reasons to avoid writing tests.
To that end, I think a better guideline than "only have one assertion per test" would be "only test one behavior per test". So if you're writing a test for appending an element to a vector, it's probably fine to assert that the size increased by one AND assert that the last element in the vector is now the element that you inserted.
The thing I see people do that's more problematic is to basically pile up assertions in a single test, so that the inputs and outputs for the behavior become unclear, and you have to keep track of the intermediate state of the object being tested in your head (assuming they're testing an object). For instance, they might use the same vector, which starts out empty, test that it's empty; then add an element, then test that its size is one; then remove the element, test that the size is 0 again; then resize it, etc. I think that's the kind of testing that the "one assertion per test" rule was designed to target.
With a vector it's easy enough to track what's going on, but it's much harder to see what the discrete behaviors being tested are. With a more complex object, tracking the internal state as the tests go along can be way more difficult. It's a lot better IMO to have a bunch of different tests with clear names for what they're testing that properly set up the state in a way that's explicit. It's then easier to satisfy the above two requirements I listed.
I want to be able to look at a test and know exactly what I don't mind if a little bit of code is repeated - you can make functions if you need to help with the test set up and tear down.
How would a test suite with one assertion per test work? Do you have all the test logic in a shared fixture and then dozens of single-assertion tests? And does that rule completely rule out the common testing pattern of a "golden checkpoint"?
I tried googling for that rule and just came up with page after page of people arguing against it. Who is for it?
This "rule" is known mostly because it is featured in the "Clean Code" book by Robert C. Martin (Uncle Bob). You should have heard of it ;)
I’ve come to learn to completely disregard any non-specific criticism of that book (and its author). There is apparently a large group of people who hate everything he does and also, seemingly, him personally. Everywhere he (or any of his books) is mentioned, the haters come out, with their vague “it’s all bad” and the old standard “I don’t know where to begin”. Serious criticism can be found (if you look for it), and the author himself welcomes it, but the enormous hate parade is scary to see.
Looking at the Amazon listing and a third-party summary[0] it seems to be the sort of wool-brained code astrology that was popular twenty years ago when people were trying to push "extreme programming" and TDD.
[0] https://gist.github.com/wojteklu/73c6914cc446146b8b533c0988c...
That paragraph is followed by 'Single Concept per Test' where he starts with: 'Perhaps a better rule is that we want to test a single concept per test.'
So, technically he doesn't say it.
When I first wrote tests years ago, I would try to test everything in one test function. I think juniors have a tendency to do that in functions overall - it's par to see 30-100+ line functions that might be doing just a little too much on their own, and test functions are no different.
It would be easier if it were all terrible advice.
https://docs.rubocop.org/rubocop-minitest/cops_minitest.html...
But even this rule/plugin has a default value of 3 which is a more sane value.
- "But it tests more things!"
Well ok, but those are integration tests, not unit tests... It is unacceptable that a unit tests can fail because of external system...
yes, we perform "somewhat real tests" - tests start the app, fake db, and call HTTP APIs
it's really decent
IMO there's no point in checking that you got a response in 1 test, and then checking the content/result of that response in another test. The useful portion of that test is the response bit.
E.g. `assertEqual(actual_return_code, 200, "bad status code")` should lead to output like `FAILED: test_when_delete_user_then_ok (bad status code, expected 200 got 404)`
FAILED: test_when_delete_user_then_ok
Assertion failed: `actual_return_code' expected `200', got `400'.
Note it mentions the actual expression put in the assert. Which makes it almost always uniquely identifiable within the test.That's the bare minimum I'd expect of a testing framework - if it can't do that, then what's the point of having it? It's probably better to just write your own executable and throw exceptions in conditionals.
What I expect from a testing framework is at least this:
FAILED: test_when_delete_user_then_ok
Assertion failed: `actual_return_code' expected `200', got `400'.
In file: '/src/blah/bleh/blop/RestApiTests.cs:212'.
I.e. to also identify the file and the line containing the failing assertion.If your testing framework doesn't do that, then again, what's even the point of using it? Throwing an exception or calling language's built-in assert() on a conditional will likely provide at least the file+line.
If I understood this part correctly, you are making the dangerous assumption that your tests will run in a particular order.
This also wasted time, because you always had to look at the tests, and eventually realized that they all failed from the same root cause. And sure, you can use test dependencies if your framework has that and do all manner of things... or you just put the asserts in the same test with a good message.
If an assertion says “expected count to be 5 but got 4” you wouldn’t be looking at the not null check assertion getting confused why it’s not null…
Still better than investigating why the whole system failed, though.
I believe writing test code is its own skill. Hence, like a coder learning SRP and dogmatically applying it, so does a person that is forced to write unit tests without deep understanding. (And of course, bad abstractions are worse than code duplication)
I think it's very possible to have a developer with 10 yes experience but effectively only 2 years experience building automated test suites. (Particularly if they came from a time before automated testing, or if the testing and automated tests were someone else's job)
IIUC, the guideline is so that when a test fails you know what the issue is. Therefore, if you are testing more than one condition (missing parameter, invalid value, negative number, etc.) it is harder to tell which of those conditions is failing, whereas if you have each condition as a separate test it is clear which is causing the failure.
Separate tests also means that the other tests run, so you don't have any hidden failures. You'll also get hidden failures if using multiple assertions for the condition, so will need to re-run the tests multiple times to pick up and fix all the failures. If you are happy with that (e.g. your build-test cycle is fast) then having multiple assertions if fine.
Ultimately, structure your tests in a way that best conveys what is being tested and the conditions it is being tested under (e.g. partition class, failure condition, or logic/input variant).
None of this is the failure of code and tests alone; but both can be indicative of the structural health and resilience of the wider situation.
I also use Apple's XCTest, which does a lot more than simple assertions.
If an assertion is thrown, I seldom take the assertion's word for it. I debug trace, and figure out what happened. The assertion is just a flag, to tell me where to look.
A good unit test has the phases Arrange, Act, Assert, end of unit test.
You can use multiple assert statements in the "assert" phase, to check the specifics of the single logical outcome of the test.
In fact, once I see the same group of asserts used 3 or more times, I usually extract a helper method, e.g. "AssertCacheIsPopulated" or "AssertHttpResponseIsSuccessContainingOrder" these might have method bodies that contain multiple assert statements, but the question of is this a "single assert" or not, is a matter of perspective and not all that important.
The thing to look out for is - does the test both assert that e.g. the response is an order, and that the cache is populated? Those should likely be separate tests as they are logically distinct outcomes.
The test ends after the asserts - You do not follow up the asserts with a second action. That should be a different test.
* The primary goal of test automation is to prevent regression. A secondary goal can be performance tuning your product.
* Tests are tech debt, so don’t waste time with any kind of testing that doesn’t immediately save you time in the near term.
* Don’t waste your energy testing code unless you have an extremely good reason. Test the product and let me the product prove the quality of your code. Code is better tested with various forms of static analysis.
* The speed with which an entire test campaign executes determines, more than all other factors combined, when and who executes the tests. If the test campaign takes hours nobody will touch it. Too painful. If it takes 10-30 minutes only your QA will touch it. When it takes less than 30 seconds to execute against a major percentage of all your business cases everybody will execute it several times a day.
This is the No True Scotsman issue with testing. When it fails, you just disregard the failure as "bad tests". But any company that has anything that resembles a testing culture will have a good amount of those "bad tests". And this amount is way higher than people are willing to admit.
> you cannot refactor with any confidence
Anecdotally, I've had way more cases where I wouldn't refactor because too many "bad tests" were breaking, not because I lacked confidence due to lack of tests.
There are many things beyond tests that allow you refactor with confidence: simple interfaces, clear dependency hierarchy, modular design, etc. They are way more important than tests.
Tests are often a last resort when all of the above is a disaster. When you're at a place where you need tests to keep your software stable you are probably already fucked, you're just not willing to recognize it.
You shouldn't have zero tests, but tests should be treated as debt. The fewer tests you need to keep your software stable, the better your architecture is. Huge number of tests in a codebase is typically a signal of shitty architecture that crumbles without those crutches.
Tests are one of the ways you have to ensure your code is correct. Consequently, they are business-oriented code that exist to support your program usage, and subject to its requirements. How much assurance you need is completely defined by those requirements. (But how you achieve that assurance isn't, and tests are only one of the possible tools for that.)
It doesn't mean it's a debt worth taking, though IME most companies are either taking way too much or way to little. Not treating tests as debt typically leads to over-testing, and it is way worse than under-testing.
Also, what you're talking about (business-oriented requirements) is more akin to higher level tests (integration/e2e), not unit tests.
I've never seen a team where bad tests were simply deleted. They were always "fixed" instead.
Your tests being a nightmare doesn't imply my tests are a nightmare.
If your mocks and tests are in the way when introducing features or refactoring, they are likely not on the right level. Too much unit testing of moving internals, rather than public apis, usually being one of the culprits.
The issue is when you use multiple assertions for multiple logic statements: do > assert > do > assert... In that example imagine that you were also checking that the reservation was successful. That would be considered bad, you should create a different test that checks for that (testCreate + testDelete) and just have the precondition that the delete test has a valid thing to delete (usually added to the database on the setup).
Migrate up -> assert ok -> rollback 1 -> assert ok -> rollback 2 -> assert ok
I don’t see much benefit to breaking it up, and you’re testing state changes between each transition, so the entire test is useful and simpler, shorter, and clearer than the alternative.
Imagine an object which sees numbers, tries to pair up identical numbers, and reports the set of unpaired numbers. A good test would be:
Create object
Assert it has no unpaired numbers
Show it 1, 2, and 3
Assert it has 1, 2, and 3 as the unpaired numbers
Show it 1 and 3
Assert it has 2 as the unpaired number
This test directly illustrates how the state changes over time. You could split it into three tests, but then someone reading the tests would have to read all three and infer what is going on. I consider that strictly worse.
This is also the reason why tests should be run in arbitrary order, to avoid unexpected interactions due to order.
Flow tests can be useful in some situations, but they should never replace individual feature tests.
Another failure mode is when test scaffolding builds up. Imagine that migrate up part becoming multiple schemas, or services. It then fails, now finding exactly where to fix the test scaffolding becomes a multi-hit exercise.
I'm not saying the example is bad, but it can put you on a path where if you constantly build on top of it, it can bad (eg, developers that don't care for tests nor test code quality, or just want to go home, and they just add a few assertions, add some scaffolding, copy-paste it all and mutate some assertions for a different table & rinse-wash-repeat across 4 people, 40 hours a week for 3 years...)
I wouldn’t actually write tests for migrations themselves.
On my company we developers usually create white-box unitary/feature tests (we know how it was implemented, so we check components knowing that). But then we have an independent QA team that creates and run black-box flow tests (they don't know how it was implemented, only what it should do and interact)
var flagtests = []struct {
in string
out string
}{
{"%a", "[%a]"},
{"%-a", "[%-a]"},
{"%+a", "[%+a]"},
// additional cases elided
{"%-1.2abc", "[%-1.2a]bc"},
}
func TestFlagParser(t *testing.T) {
var flagprinter flagPrinter
for _, tt := range flagtests {
t.Run(tt.in, func(t *testing.T) {
s := Sprintf(tt.in, &flagprinter)
if s != tt.out {
t.Errorf("got %q, want %q", s, tt.out)
}
})
}
}
Sometimes there will be an additional field in the test cases to give it a name or description, in which case the assertion will look something like: t.Fatalf("%s: expected: %v, got: %v", tt.name, tt.out, got)
Another evolving practice is to use a map instead of a slice, with the map key being the name or description of the test case. This is nice because in Go, order is not specified in iterating over a map, so each time the test runs the cases will run in a different order, which can reveal any order-dependency in the tests.1 https://dave.cheney.net/2019/05/07/prefer-table-driven-tests
There are no hard and fast rules.
I agree with keeping the number of assertions low, but it isn't the number that matters. Keeping the number of assertions low helps prevent the 'testItWorks()' syndrome.
Oh, it broke, I guess 'it does not work' time to read a 2000 line test written 5 years ago.
I think it was a guidance, like the SRP where you should be testing one thing in each test case. I also think a growing number of assertions might be a sign your unit under test is wearing too many responsibilities.
Maybe it’s better to say “few assertions, all related to testing one thing”
It's practically impossible to thoroughly test such code with one assertion per test; it would mean having dozens of tests just for one object method. Correspondingly, the fixtures/factories/setup for tests would balloon in number and complexity as well to be able to setup the exact circumstance being tested.
But the example in TFA is, imo, bad because it is testing two entirely different (and unrelated) layers at once. It is testing that a business logic delete of a thing works correctly, and that a communication level response is correct. Those could be two separate tests, separating the concerns, and resulting in simpler code to reason about and less maintenance effort in the future.
We want to know if the DeleteAsync(address) behaves correctly. Actually we want to know if DeleteReservation() works, irrespective of the async requirement. Testing whether AnythingAsync() works is something that is already done at the library or framework level, and we probably don't need to prove that it works again.
Write a test for DeletReservation() which tests if a valid reservation gets deleted. Write another related test to ensure that a non-existent or invalid reservation does not get deleted, but rather returns some appropriate error value. That's two, probably quite simple tests.
Now somewhere higher up, write the REST API tests. ApiDelete()... a few tests to establish that if an API delete is called, and the business logic function it calls internally returns a successful result, then does the ApiDelete() return an appropriate response? Likewise if the business logic fails, does the delete respond to the API caller correctly?
Yes, that means the software being tested is not very testable. And should probably be refactored.
In my experience when code isn’t OOP, that means all static functions with static (I.e. global) data which isn’t hard to test, it’s actually impossible because you can’t mock out the static data.
OOP functions are usually harder to test because they expect complete objects as arguments, and that tends to require a lot more mocking or fixtures/factories to setup for the test.
FP functions typically operate on less complex and more open data structures. You just construct the minimum thing necessary to satisfy the function, and the test is comparatively simple. None of this has anything to do with global data. Using global data from within any functions is generally a bad idea and has nothing to do with FP or OOP.
Pure functional, yeah, absolutely - that’s not what I usually see though. I see procedural/iterative static functions that connect to static data that connect to live databases and immediately start caching its contents locally.
A macro like
EXPECT_ASSERT(whatever(NULL));
will succeed if whatever(NULL) asserts (e.g. that its argument isn't null). If whatever neglects to assert, then EXPECT_ASSERT will itself assert.Under the hood it works with setjmp and longjmp. The assert handler is temporarily overridden to a function which performs a longjmp which changes some hidden local state to record that the assertion went off.
This will not work with APIs that leave things in a bad state, because there is no unwinding. However, the bulk of the assertion being tested are ones that validate inputs before changing any state.
It's quite convenient to cover half a dozen of these in one function, as a block of six one-liners.
Previously, assertions had to be written as individual tests, because the assert handler was overriden to go to a function which exits the process successfully. The old style tests are then written to set up this handler, and also indicate failure if the bottom of the function is reached.
Integration tests are typically easier to write / maintain and thus are more valuable than small unit tests. Don’t know why the entire premise argues against that.
The first part where he says you can check in the passing test of code that does nothing makes me twitch, but mainly because I see functional testing as the goal for a bit of functionality and unit tests as a way of verifying the parts of achieving that goal. Unit tests should verify the code without requiring integration and functional tests should confirm that the units integrate properly. I wouldn't recommend checking in a test that claims to verify that a deleted item no longer exists when it doesn't actually verify that.
Deciding on the granularity of actual unit tests is probably something that is best decided through trial and error. I think when you break down the "rules", like one assertion per test, you need to understand the goals. In unit testing an API I might have lots of tiny tests that confirm things like input validation, status codes, debugging information, permissions, etc. I don't want a test that's supposed to check the input validation code to fall because it's also checking the logged-in state of the user and their permission to reach that point in the code.
In unit tests, maybe you want to test the validation of the item key. You can have a "testItemKey" test that checks that the validation confirms that the key is not null, not an empty string, not longer than expected, valid base64, etc. Or you could break those into individual tests in a test case. It's all about the balance of ergonomics and the informativeness and robustness of the test suite.
In functional testing, however, you can certainly pepper the tests with lots of assertions along the way to confirm that the test is progressing and you know at what point it broke. In that case, the user being unable to log in would mean that testing deleting an item would not be worthwhile.
I think the better rule is “just test one thing, don’t use a lot of overlapping assertions between tests”.
Number of assertions, as noted by this post and this comment section, is not a smell at all.
The issue with multiple assertions for unit tests is that they hide errors by failing early extending the time to fix (fail at Assert 1, fix, re-run, fail at Assert 2, fix...). If you want multiple assertions you should have a single assertion that can report multiple issues by returning an array of error strings or something. The speed increase of multiple assertions vs multiple tests is usually tiny, but if you do have a long unit test then that's basically the only reason I can think of to use multiple assertions.
Even when I do have multiple failing tests, I focus on and fix them one test at a time.
jUnit provides a helpful assertAll(...) method that allows us to check multiple assertions without stopping and the first failed one.
I my tests I often use "thick" asserts like assertEqualsWithDiff(dtoA, dtoB), that compare 2 objects as a whole and prints property names with values that do not match. Not everyone likes this approach, but for me it is the best balance between time spend on the test and the benefit that I get from it.
Do whatever you need/want to that makes your software successful. The rules are meant to be broken if it means you shipping.
Without a "assert one thing", what tends to happen is there will be cut-and-paste between tests, and the tests overlap in what they assert. This means that completely unrelated tests will have an assertion failure when the code goes wrong.
When you do a refactor, or change some behavior, you have to change _all_ of the tests. Not just the one or two that have the thing you're changing as their focus.
Think of tests that over-assert like screenshot or other diff-based tests, they are brittle.
Test setup is sometimes very complicated AND expensive. Enforcing only 1 assertion per test is moronic beyond description.
def testLotsOfWaysTofail(): d = {"Handle Null": (None, foobar), "Don't allow under 13 to do it": (11, foobar), "or old age pensioner": (77,wobble)} for ... generate a unit test dynamically here
I have built metaclasses, I have tried many different options. I am sure there is a near solution.
But I never seem to have it right
https://github.com/ckp95/pytest-parametrize-cases
@parametrize_cases(
Case("handle null", age=None, x="foobar"),
Case("don't allow under 13s", age=11, x="foobar"),
Case("or old age pension", age=77, x="wobble"),
... # as many as you want
)
def test_lots_of_ways_to_fail(age, x):
with pytest.raises(ValueError):
function_under_test(age, x)If all your assertions are always evaluated, you don’t have this issue?
It must be implemented as a macro so that the line number and assertion expressions are printed, allowing to easily identify the failed assertion.
If a language doesn't support such macros and has no ad-hoc mechanism for this case, it should not be used, or if it must the assert function must take a string parameter identifying the assertion.
If you know the file and line of the assertion, plus the values that are being checked, there's not as much need for a stringified version of the expression.
It does save time. With the actual condition reproduced, half the time I don't even need to check the source of the failed test to know what went bad and where to fix it. Consider the difference between:
FAILED: Expected 1, got 4
In /src/foo/ApiTest.cpp:123
vs. FAILED: response.status evaluated to 4
expected: 1
In /src/foo/ApiTest.cpp:123
vs. FAILED: response.status evaluated to Response::invalidArg (4)
expected: Response::noData (1)
In /src/foo/ApiTest.cpp:123
This is also why I insist on adding custom matchers and printers in Google Test for C++. Without it, 90% of the time a failed assertion/expectation just prints "binary objects differ" and spews a couple lines of hexadecimal digits. Adding a custom printer or matcher takes little work, but makes all such failures print meaningful information instead, allowing one to just eyeball the problem from test output (useful particularly with CI logs). FAILED: Expected RESPONSE_NO_DATA, got RESPONSE_INVALID_ARG
In /src/foo/ApiTest.cpp:123
For non-enumerated integers, the output would be something like: FAILED: Expected ResponseCode(200), got ResponseCode(404)
In /src/foo/ApiTest.cpp:123
If values are formatted with some amount of self-description, then printing the name of the variable being evaluated is not usually informative.Nothing like removing one line, breezing through code review, and bringing down production. "But all tests passed!"
Arrange: whatever you need for the setup. Act: a single line of code that is under test. Assert: whatever you want to assert, multiple statements allowed and usually required.
We put the 3A in code as comments as boundaries and that works more than perfect for the whole team.
Oh and it's readable!
Actual status code: NotFound.
Expected: True
Actual: Falsetbh unit testing is a balancing act between reaping code quality benefits and bogging yourself down with too much testing updating.
I'm constantly thinking where we need to put tests though, and in still not fully convinced I get it right. My rule of thumb is that each test should map to a specification point, and that spec is a necessary documentation line for the test.
Both seem to be driven by misunderstanding through simplification of some otherwise meaningful ideas.
To the point where if they are testing object equality, they check each field in its own test.
I have no idea where they got the notion from.
they were like no, copy paste the entire setup and change the field to false in the new setup.
im like how are you supposed to tell which of the dozen conditions triggered it now? you have to diff the 2 tests in your head and hope they don't get out of sync? ridiculous
On a specific case, what is the gain from not testing the pre and post-conditions of your important test?
1) Test suites are organized as trees, not as lists:
I found one of the most common reasons to have many assertions in a test was that you want to share some complicated setup/teardown logic - or that one testable action depends on another testable action having happened before. (i.e., adding an item - asserting it's there, then removing it, asserting it's gone).
The disadvantage is that you have to lower the granularity of your tests - if you want to debug a specific action, you still have to rerun the whole test.
I think a better way to solve this would be to organize tests as a tree, maybe something like this:
- A single unit test consists of a setup phase, a teardown phase, 0 or more assertions and 0 or more child tests. Each child test is organzed the same, i.e. can have child tests on its own, etc.
- When running a test, first the setup phase and assertions are executed, then each child test recursively, then the teardown phase. Success/failure is tracked for each test separately, but child tests are ran in the same process/context as the parent test.
- Each test can be started individually, including child tests. When a child test (or grandchild test, etc) is ran individually, the test runner will first run the setup phases of all ancestors, then run the test, then run the teardown phases of the ancestors.
- Bonus: In the setup phase, a test can dynamically generate child tests (e.g. as lambdas/closures). Each test must have a unique ID with which it can tracked across different test runs or started individually. This could be useful for parameterized tests or if you want to test a loop invariant across multiple iterations.
This would allow you to write your test script like one big multi-assert test, but still get fine-grained reports and control as if you'd have put each assert in a separate script.
2) Provide "metrics" and "change detection" as an alternative to assertions:
I think one of the most involved parts of writing tests can often be to verify the results - think which particular state you want to assert, how you can access that state in your script, etc.
A way to make this easier would be to provide a second kind of "output" for the test script: The test script simply outputs a list of key/value pairs without any notion whether or not the value is "correct" or "incorrect". The test runner stores the list and compares the values with the list from a previous test run - e.g. the previous commit. Every value that was changed between the runs is shown to the user and can be marked as "correct" or "incorrect".
This way, you could sort of interactively "learn" which values are correct and which aren't instead of having to figure out all of it beforehand.
The runner could also implement more complex conditions instead of "changed"/"did not change", such as "value may only change in one direction" e.g. for quality measures or "value must stay the same within a certain confidence interval" for flaky tests.
This could also let you track more difficult to manage metrics in a test, such as runtime or memory consumption of particular method calls.
- whoever
I definitely write tests with multiple assertions, the rule I try to follow is that the test is testing a single cause/effect. that is, a single set of inputs, run the inputs, then assert as many things as you want to ensure the end state is what's expected. there is no problem working this way.