Is it just bruteforcing tests? That means it's not able to test all possible inputs for any function working on long strings or arrays or dicts...
Is it just bruteforcing tests? That means it's not able to test all possible inputs for any function working on long strings or arrays or dicts...
It also doesn't exhaust the set of all possible code execution branches, for example, like an intelligent fuzzer would.
Saying that it's "just" bruteforcing tests still seems a bit dismissive for what you get, at least to my eyes. Property generators are good at exploring interesting neighbourhoods of the search space. When they find a failure, they're also usually good at reducing the error input to something small.
As an example, it might learn that "@#*CH@R822cr21;'c09J@)RH0 92hr19h" causes an error, but then be able to reduce the input to learn that it's the apostrophe that causes the issue -- a strong hint that you've messed up parameter escaping somewhere in your function.
AFAICT hypothesis is not doing SMT, it's not doing intelligent fuzzing and not doing bruteforcing. What algorithm is it using to explore the search space then?
Having a very effective algorithm is crucially important but the documentation does not compare the tool with other methodologies.
The place where I've applied this (not using Hypothesis, mind you) is testing Common Lisp compilers. There are some general techniques for biasing the inputs to encourage certain kinds of programs, but overall just smothering a compiler in a tidal wave of tests exposes all sorts of weird bugs you'd never think of until you saw them. For those sorts of bugs, trying to partition the input space ahead of time is just useless.
They have a bunch of articles here including a variety that delve into the "how" of Hypothesis. Here is a 2016 article titled "How Hypothesis Works": https://hypothesis.works/articles/how-hypothesis-works/
There is very very clever code that does things like create weird dicts where the keys are weird strings and the values are other weird dicts or arrays or strings..
If it finds a test failure, it then applies a reduction step, where it replaces the extremal test case with something slightly more benign and checks if the test still fails. This allows it to generate a test case which is just hairy enough to trigger the bug, but no more. This makes it easier to understand why that specific test case fails.
Does anyone have any recommendations for papers on this?