In praise of property-based testing (2019)
increment.com
increment.com
Then recently I had a fairly larger, more error prone piece of work that lend itself very well to property-based testing, and it's been a godsend. It helped me discover a number of bugs, this time with the risk of causing privilege escalation. And since the proptests started succeeding reliably, I've been very confident that a rather complex piece of code now actually does what it's supposed to.
If you're working in JavaScript, I can recommend fast-check [1].
Another interesting approach, that I haven't yet tried, is Quickstrom [2], basically Puppeteer for proptests. It opens a webpage in a browser, performs some random interactions (pressing buttons, entering data, etc.), and then verifies that properties you specified still hold.
Property based testing looks like a really good idea, but I never gained anything from applying it on practice. There are probably some application domains that they are good for, but I still didn't find them.
In the case of the parsing code from the article, the emerging property is that a serialization followed by a deserialization should always yield the original result.
In the case of the 'binary or not' case from the article, the non-trivial, emergent property is that the function never fails with an exception.
Most modern software development is plumbing, stitching together platforms, libraries and frameworks, and there are rarely non-trivial, emergent properties where property-based testing is useful.
- merging/combining of requirements is a monoid operation with an "empty requirement" (e.g. one that accepts every password) as the neutral element. Finding this requirement, I could write properties for the monoid properties (a + b = b+ a, a + 0 = a, and so on) - when generating a password from a requirement, the same requirement must accept the generated password (`requirement.accept(requirement.generate()`)
In general: if you find mathematical rules to your code (commutativity, associativity,. ..), these make a great starting point for property tests.
The first approach is to generate random correct programs and see if they do the same thing with different compilers and/or different optimization flags.
The second approach is to take a correct program, note that some parts of it are not executed on some particular input, then mutate that part and run on the same input. The output should be the same.
Also, any program (not just compilers) should satisfy the properties that no assertions fail and that no sanitizer failures occur.
[0] https://pragprog.com/titles/fhproper/property-based-testing-...
I loved the video but my first experience with this method was with Python's Hypotheis [2], I haven't used it a ton but it's great for finding parsing errors. In the words of Python core developer Raymond Hettinger:
> It's not quite fuzzing, but it hits it with the kind of test cases that a good QA engineer would typically come up with, and it does it in automated fashion.
1. https://www.youtube.com/watch?v=AfaNEebCDos
Then I noticed who wrote it and thought "Ah. Well that makes sense."
Doesn't matter what methodology you use if your test doesn't test what you want it to.
I also fear that in general, people try to make their tests too clever. I don't want the test to generate random data that's different every time, unless it's a fuzz test. I want to see the inputs and the expected outputs as clearly as possible, so that when the test fails I'm not guessing whether or not there is some bug 30 helpers deep, or if I simply added a bug with my change. The real world is complicated and people are going to enter a lot of invalid data into your application, so it is crucial that you test that. But you also need the basics to work before you worry about the advanced cases.
The "shrinking" capability of the test library highlighted is brilliant.
I'm inspired to think of how to start to leverage something like this on some upcoming work.
The first idea you think of for shrinking it to take the randomly generated values and try to make them smaller. But the generator may be imposing constraints on the values, and if you lose those constraints, the input becomes invalid.
An example of this problem is generating valid C programs to test C compilers (by compiling with different compilers or different optimization settings and seeing if the behavior differs). The constraint there is that the C program not show undefined or implementation-specific behavior. Naively shrinking a C program will not in general preserve this property.
Hypothesis takes a different approach to shrinking: it records the sequence of random values used by the generator, and replays the generator on mutations of that sequence that do not increase its length. The only way this can fail is if the generator runs out of values on the mutated sequence. Otherwise, the new output will always satisfy the constraints imposed by the generator. Hypothesis does various clever things to speed this up.
Eg the simplest example might be an application where you input an integer, but a smaller int actually drives up the complexity. This idea gets more complex when we consider a list of ints, where a larger list and larger numbers are simpler. Etcetc.
It would be neat if a proptest framework supported (and maybe they do, again, i'm a novice here) a way to rank complexity. Eg a simple function which gives two input failures and you can choose which is the simpler.
Another really neat way to do that might be to actually compute the path complexity as the program runs based on the count of CPU instructions or something. This wouldn't always be a parallel, but would often be the right default i imagine.
Either way in my limited proptest experience, i've found it a bit difficult to design the tests to test what you actually want, but very helpful once you establish that.
“Property testing is where you vary the parameters, perhaps within certain well defined constraints, and through randomization ensure certain coverage.”
I think there are other ways of handling this, beyond the way described, but I’m not sure if that’s because I just read a field guide to Hypothesis (again, without a firm definition), or if other approaches - such as test parameterization at key boundaries is more effective.
My hunch is that property based testing provides less value on unit tests - where it could be used to reduce coverage - and more value in integrating those units together - where the number of permutations can grow very quickly.
No - this would be akin to saying "unit testing is when you use assert statements".
Varying the parameters -- and using far more than simple randomisation -- is just a technique; the purpose of property testing is that the behaviour of the unit (or system) under test fulfills certain properties, in ma mathematical sense of the word. For a simple example: if you're writing a function that is adding two integer numbers, you want to test that it satisfies all the properties of addition [0]. So your tests need to validate against a number of combinations of input values, many of which are not random (e.g. to test for the "successor" property one of them needs to be 1).
It's also helpful to use it in a piece wise fashion if you're doing TDD. An illustrative example (though perhaps not stellar as it is a synthetic, non-real-world example) uses the diamond kata, TDD, and PBT together [0]. None of the tests on their own fully specify the system, but in total they do.
If you're doing TDD (or attempting to) I think this is an interesting case. Many TDD methods have you start off with an example case like (to stick with this kata, and using Pythonesque pseudocode because I'm still not awake this Saturday morning):
diamond-kata-test-a:
assert(diamond('A') == 'A')
Great, so now someone makes that absolute simplest solution: diamond(c):
return 'A'
Now repeat with a second test case: diamond-kata-test-b:
assert(diamond('B') == ' A \nB B\n A ')
And the function is duly complicated: diamond(c):
switch c:
case 'A': return 'A'
case 'B': return ' A \nB B\n A '
default: return 'blah' // or error, doesn't matter it's not tested
But not actually generalized to reflect the intent of the system. By focusing on properties, I've found, the progression of the UUT is a bit better/more natural.Another interesting thing to do with PBT is model-based testing [1]. The useful thing here is that sometimes the errors are triggered by a peculiar, though plausible, sequence of commands to your system. We've all worked with that one guy who somehow manages to find exactly the right sequence that triggers weird edge cases and errors, but unless we're him having a system which will generate arbitrary execution traces for you. (I actually used FsCheck for this last year in trying to sell PBT to my colleagues and was able to identify where a known issue originated as well as several other problems that hadn't been found by users or testers yet.)
In the end, when these failures are found you can always turn them into distinct unit tests in order to preserve them and prevent regressions. The two modes of testing fit well together.
[0] http://christophethibaut.com/programming/2020/03/18/Diamond-...
For example if I'm making a sqrt function, then I want sqrt(x) * sqrt(x) == x, for any x>=0.
Human beings are good at inferring the general rules from examples but sometimes it's easier to understand if you just say what the general rules is.
Also, unit tests that include example data can sometimes be dominated by that data. Removing the particular examples can sometimes remove a lot of distraction and verbosity.
It's not just that property tests find more bugs.