My Python testing style guide (2017)
blog.thea.codes
blog.thea.codes
It also encourages tests that do too much. If the test is named "test_refresh", well, it says right there in the name that it's a test for any old generic refresh behavior. So why not just keep dumping assertions in there?
I'm much more happy with names like, "test_displays_error_message_when_refresh_times_out". Right there, you know exactly what's being verified, because it's been written in plain English. Which means you can recognize a buggy test implementation when you see it, and you know what behavior you're supposed to be restoring if the test breaks, and you are prepared to recognize an erroneously passing test, and all sorts of useful things like that.
class TestAuthenticateMethod:
class TestWhenUserExists:
class TestWhenPasswordIsValid:
def test_it_returns_200(self, app, user):
assert ...
class TestWhenPasswordIsInvalid:
def test_it_returns_403(self, app, user):
assert ...
class TestWhenUserDoesNotExist:
def test_it_returns_403(self, app, user):
assert ...
It is so worth it. I can look back at my tests 3 months later and know exactly what was intended wthout duplicating explanations with comments. Yes some methods can get a bit long and indented, but it's worth it for me. The other downside is you have to start all classes/methods with "test" but I got used to that very quickly.When you run pytest it's very clear what's being tested and what the expected outcome is since it gives you the full class hierarchy in the output.
E.g. in test_authenticate_method.py:
def test_returns_200_on_valid_password(app, user_exists):
...
def test_returns_403_on_valid_password(app, user_exists):
...
def test_returns_403(app, user_not_exist):
... foo
<snip 1000 lines>
bar
<snip 553 lines>
baz
<snip 243 lines>
quxIn tests, however, I prefer not to have them as my favorite test runners replace the full name of the test method with its docstring, which makes it a lot harder to find the test in the code.
def matriculate_flange(via: Worble):
"Matriculates flange via a Worble."Test function names are less sensitive than regular functions, since they're not explicitly called, but I still don't want to read a_sentence_with_spaces_replaced_by_underscores.
As far as whether or not people do a better job with descriptive test function names, what I've seen is that they do? I of course can't share any data because this is all in private codebases, so I guess this could quickly devolve into a game of dueling anecdotes. But what I've observed is that people tend to take function names - and, by extension, function naming conventions - seriously, and they are more likely to think of comments as expendable clutter. (Probably because they usually are.) Which means that they'll think harder about function names in the first place, and also means that code reviewers are more likely to mention if they think a function name isn't quite up to snuff.
And I just don't like to cut those sorts of human behavior factors out of the picture, even when they're annoying or hard to understand. Because, at the end of the day, it's all about human factors.
I was talking specifically about a `def test_displays_error_message_when_refresh_times_out()` function.
That's too big a name for me to keep in my head, so I'd look for other solutions.
pytest does not though.
How are they, in practical reality, better than comments?
The way we work, it's just a different comment syntax. I do like have a dedicated place for it.
All Python programmers already benefit from the docstrings written in the libraries they use, which generate tooltips and documentation websites.
prior to mypy/type hints it allowed you to document the types of a function
PS: How come your setup and teardown operations are so expensive in the first place? Why do you even need to set up an entire database? Can't you mock out the database, set up only a few tables or use a lighter database layer?
[0]: https://docs.python.org/3/library/functools.html#functools.c...
We've also found over the years that if you want to be sure everything works on version X of database system Y, you run your CI tests on that.
That way, all your unit tests reads: “test that error message is displayed when refresh times out” etc.
I think it’s a really nice way to lay things out and it avoids all the “magic” of some functions being executed by virtue of their name.
In my experience, when a unit test has no clear name, this is usually a sign that the test does too much (meaning that it's actually not a unit test anymore) or even comprises multiple tests, leading to all the known bad consequences this has (leaking state between tests, unclear test assumptions etc.)
Assert foo(), "foo should bar"
So that each failure has a clear meaning.
Indeed, i may have several assert in a test.
Here's an example use case: I have a test suite that tests my application's interactions with the DB. In my experience, the most tedious part of these kinds of tests is setting up the initial DB state. The initial DB state will generally consist of a few populated rows in a few different tables, many linked together through foreign keys. The initial DB state varies in each test.
My approach is to create a pytest fixture for each row of data I want in a test. (I'm using SQLAlchemy, so a row is 1-1 with a populated SQLAlchemy model.) If the row requires another row to exist through a foreign key constraint, the fixture for the child row will depend on the fixture for the parent. This way, if you add the child test fixture to insert the child row, pytest will automatically insert the parent row first. The fixtures ultimately form a dependency tree.
Finally in a test, creating initial DB state is simple: you just add fixtures corresponding to the rows you want to exist in the test. All dependencies will be created automatically behind the scenes by pytest using the fixtures graph. (In the end I have about ~40 fixtures which are used in ~240 tests.)
We started out with like 50 fixtures, but now we have a conftest.py file that has `institution_1`, ..., `institution_10`.
My end conclusion is that fixtures are nice for some things, like managing mocks, and clearing the databases after tests, but for data it's better to write some functions to create stuff.
So instead of `def test_something(institution_with_some_flag_b)` you'd write in your test body:
def test_something() -> None:
institution = create_institution(some_flag="b")
Also another benefit is you can click into the function whereas fixtures you have to grep.I’d argue that too many global fixtures in conftest have a high risk of becoming a “Mystery Guests” or too general fixtures. For a test reader it’s impossible to know the semantics of “institution_10”.
I believe this to be rooted in DRY obsession leading to coupling of tests: “We need a second institution in two modules? Let’s lift it up to global!”
There are so many other (better) ways to achieve the same goal, such as decorators or – as already mentioned by emptysea in their sibling comment – explicitly invoking some function from within the test to do the setup/teardown.
One one hand, I've been impressed with how they compose and have let me do some great things. For example, I had system tests that needed hardware identifiers. I had a `conftest.py` to add CLI args for them. I then made fixtures to wrap the lookup of these. In the fixture, I marked it as Skip if the arg was missing. This was then propagated to all of the tests, only running the ones the end-user had the hardware for.
On the other hand, when I need to vary the data between tests and that data is an input to something that I'd like to abstract the creation of, fixtures break down and I have to instead use a function call.
I'm curious if anyone else who has been drawn in by the allure of the Mock has some strategies to avoid the footguns associated with them? (Python specifically)
I linked this excellent talk in another thread recently. I'll put it here again: https://www.youtube.com/watch?v=EZ05e7EMOLM
"a mock’s job is to say, “You got it, boss” whenever anyone calls it. It will do real work, like raising an exception, when one of its convenience methods is called, like assert_called_once_with. But it won’t do real work when you call a method that only resembles a convenience method, such as assert_called_once (no _with!)."
https://engineeringblog.yelp.com/2015/02/assert_called_once-...
When unsafe=False (the default), accessing an attribute that begins with assert will raise an error.
[1]: https://docs.python.org/3/library/unittest.mock.html#the-moc...
Or any of the names in this surprising hardcoded list of typos https://github.com/python/cpython/blob/fdb9efce6ac211f973088...
https://mock.readthedocs.io/en/latest/changelog.html#and-ear...
Depending on context and implementation details, I'd say DRYing tests can be anywhere from indispensable to toxic.
I'm fine with creating libraries of shared functionality that tests can use, especially when it helps readability. If you've got several tests with the same precondition, having them all call a function named "givenTheUserHasLoggedIn()" in order to do the setup is a nice readability win. And, since it's a function call, it's not too difficult to pick apart if a test's preconditions diverge from the others' at a later date.
What I absolutely cannot stand is using inheritance to make tests DRY. If you've got an inheritance hierarchy for handling test setup, the cost of implementing a change to the test setup requirements is O(N) where N is the hierarchy depth, with constant factors on the order of, "Welp, there goes my afternoon."
I've found that having a class/function as a parameter and explicitly listing the classes/functions that get tested is a small step back and way easier to maintain and read. It sets off some DRY alarms, cause usually that whole list is just "subclasses of X". And it seems like burden to update. "So if I make a new subclass, I have to add it everywhere?". Yes. Yes you do. Familiarity with the test suite is table stakes for development. You'll need to add your class name to like ten lists, and get 90% coverage for your work, then write a few tests about what's special about your class. When something breaks, you'll know exactly what's being tested. And you'll be able to opt out a class from that test with one keystroke.
That being said… I still have a dream of writing a library for generating tests for things that inherit from collections.abc. Something like “oh, you made a MutableSequence? let’s test it works like a list except where you opt-out.”
It does annoy the many programmers who want clear and absolute rules for everything.
Then again they are always annoyed, living in a world where so many things "depend".
I feel like your last point is especially important. Sooooo many times have I seen over-abstracted unit tests that are unreadable and are impossible to reason about, because somebody decided that they needed to be concise (which they don't).
I'd much rather tests be excessively verbose and obvious/straightforward than over abstracted. It also avoids gigantic test helper functions that have a million flags depending on small variations in desired test behaviour...
Personally, I work with some incredibly (100+line) long "unit" tests and they are a nightmare to work with.
Especially when the logic is repeated across multiple tests, and it's incorrect (or needs to be changed).
I really, really like shorter tests with longer names, but I'd imagine there are definitely pathologies at either end.
test_refresh_failure
test_refresh_with_timeout
These get even longer like test_refresh_with_timeout_when_username_is_not_found for example.pytest-describe allows for a much nicer testing syntax. There's a great comparison here: https://github.com/pytest-dev/pytest-describe#why-bother
TL;DR, this is nicer:
def describe_my_function():
def with_default_arguments():
def with_some_other_arguments():
This isn't as nice: def test_my_function_with_default_arguments():
def test_my_function_with_some_other_arguments():Something like this might be a good compromise:
def describe_my_function(register):
@register
def with_this_thing():
...
I think most Python devs understand that "register" can have a side-effect. class Test_my_function:
def with_default_arguments(self):
def with_some_other_arguments(self):
If you can make your eye stop twitching after seeing snake cased class names, this is at least another option of grouping tests for a single function.With rspec, you use the describe and context keywords.
At one level, yes, it's mainly syntactical sugar. As the test-writer, the two approaches may seem interchangeable.
Where I find it really helps is when I'm not the test-writer but rather I'm reviewing another developer's tests, say in PR. I find this syntax and hierarchy produces a much more coherent test suite and makes it easier for me to twig different use cases and test quality generally.
> With pytest, it's possible to organize tests in a similar way with classes. However, I think classes are awkward. I don't think the convention of using camel-case names for classes fit very well when testing functions in different cases. In addition, every test function must take a "self" argument that is never used.
So there's no reason to do this, aside from aesthetics.
I'd recommend against doing un-Pythonic stuff like this, it makes your code harder to pick up for new engineers.
And I've never had a problem with reading class names vs. function names; we do that all the time when reading Python code.
I think this is clearly in the realm of subjective preference for what "looks nicer", which I'd call aesthetics.
On the other hand, this probably breaks your IDE's pytest integration, which would be an objective material downside.
Whatever floats your boat though, definitely not a hill I'd die on.
I've previously used serialized data—JSON, or joblib if there are complex types (eg, numpy)—but these seem pretty brittle...
What do you mean, extra dependencies? The only difference between pytest and unittest in this regard is that tests using unittest declare their dependency explicitly, using an import[0]. Most pytest tests still implicitly require pytest as a dependency, though. (Think of fixtures etc. etc.)
I actually like unittest's approach here – in my book, explicit is better than implicit.