With property based testing, it actually CAN be 9 passes and 1 failure, because that one single fail can be hitting an edge case the others just aren't. In fact, only a few failures are more likely than it being all failures
I can tell you from experience that random failures cause loss of trust. People learn to ignore failures and just keep hitting rebuild until the tests pass. People will not investigate test failures in code that they don't understand.
But yeah, probabilistic testing isn't perfect.
Allowing reproducibility for a given change set.
Didn't even enter my mind.
Hypothesis has a nice option, to pick the seed for each property by hashing that property's code. It's a nice idea, but relies on Python's highly dynamic nature; so may not be easy/possible in other languages (especially compiled ones).
The co-worker should add the test to a new branch / test it on main. If it fails that's a new ticket (with the great side effect of having a failing test). If that passes it's a problem in their branch. If not it's the same as having a broken main which happens anyway and you deal with that as you usually do.
You can have a prng seeded from something like the commit hash and have prngs generate the test cases. That still can fail on 1/10 tests for a particular run.
The most memorable discussion I had around PBT was with a colleague (a skip report) who saw "true" randomness as a net benefit and that reproducibility was not a critical characteristic of the test suite (I guess the reasoning was then it could catch things at a later date?). To be honest, it scared the hell out of me and I pushed back pretty hard on them and the broader team.
I have no issue with a psuedo-random set of test cases that are declaratively generated. That makes sense if that is what is meant by PBT. Since it is just a more efficient way of testing (and you would assume this would allow you to cast a wider net).
The idea is you have random testing and the test failures are added as explicit tests that then always get run.
Is that so different from someone else testing?
The main issue is you stumble across a new issue in an unrelated branch, but it's not wildly different from doing that while using your application.
I've been there. It takes many hours to try to guess where the system went wrong to produce the undesirable result, and then you still might not be sure if you are looking at the right place, and then there are always environment issues, you aren't sure of. So, you don't know who to blame.
Only very simple systems will reproduce pathological results 100% of the time given some initial conditions. The bane of complex systems is the timeouts. They are usually very hard to justify and are easy to blame for undesirable behavior.
To test an entire system with reproducable failures, you probably need something more heavyweight like Antithesis. Property tests are more useful for unit tests.
But, in general, yes, unit test failure with many steps would've been just as difficult to interpret.
My experience though was that once such a difficult failure is encountered during property-based testing, one has to write a unit test to reproduce the error anyways. But it's hard to assess the probability of the unexpected behavior of being an actual bug. Sometimes you discover that the system behavior was underspecified, or that you misunderstood how the system is supposed to behave after reading the specification.
In practice, property based testing fails because the organization is not actually interested in delivering correct code. "This bug will never happen in practice so we won't fix it." "If we fix this, we may change some incorrect behavior some customer is depending on." And once that happens, PBT is useless, because it will keep finding that "don't fix" bug over and over.
In that case, we update the properties to reflect the new spec.
What's more PBT doesn't depend on having a spec, just on having some properties that hold. So you very possibly didn't have a spec to start with.
You don't know, but you can ask, you can make educated guesses, etc. As I said, we do this stuff iteratively, on a best-effort basis, using one's own knowledge and experience, with input and feedback from colleagues and stakeholders, etc. That's what most programming is.
> What's more PBT doesn't depend on having a spec, just on having some properties that hold
I'd say that "having some properties that hold" certainly counts as "some sort of spec (whether formal or informal, written or vibes-based, etc.)".
> So you very possibly didn't have a spec to start with.
There's always "some sort of spec"; even if it starts as vague as "let's try to make some money using computers".
It seems to be almost totally forgotten, since the only link I could find is an excerpt from a PalmOS programming book:
https://www.oreilly.com/library/view/palm-programming-the/15...
In bigger programs this is an outright necessity because pure random fuzzing would basically be a lottery.
I've always felt that unit testing frameworks and libraries and even parameterized testing were missing this kind of functionality.
Intellij is able to run my tests and figure out the code coverage, but why isn't it closing the loop and auto-fuzzing/auto-discovering how to mutate tests to cover more?
And don't point me at AI, none of this requires AI and nothing should have to "think" to do this.
It's crazy to me that the vast majority of code running all the time is not exhaustively tested through almost all of it's possible state space with most of it's possible input space. It's not like we are lacking the CPU bandwidth to do it.
Why can't I write a new function and have something tell me within ten minutes "this input param causes an exception" without any effort from me? Instead all those extra cores in my CPU just run javascript trash and crowdstrike scanners
Sure, but that's an optimisation/implementation-detail. Similar to how PBT frameworks tend to use random generation + shrinking: it's not fundamental to the approach, but turns out to be much more effective than e.g. enumerative testing (e.g. Smallcheck), or showing un-shrunk examples.
> I've always felt that unit testing frameworks and libraries and even parameterized testing were missing this kind of functionality.
Coverage-guided PBT seems to have been re-invented several times (e.g. using QuickCheck with HPC in Haskell), though all the examples I've seen are toys or experiments. Hypothesis has experimental support for generating data using an external fuzzer, which presumably uses coverage (though I've not tried that feature yet).
> Why can't I write a new function and have something tell me within ten minutes "this input param causes an exception" without any effort from me?
I agree. One piece of advice is to respect the options provided by PBT frameworks, e.g. for the number of tests to run, the maximum discard:success ratio, the maximum "size" to pass into generators, etc. These can be tweaked per property, e.g. if a particular property is slowing down our CI we might want to only test it 20 times instead of the default of 100. However, I always make sure to transform that default value (in this case dividing it by 5) rather than setting a particular number: that way the test suite can also be run with bigger options to get a more thorough search (e.g. locally in the background, or by another CI job that's run less often, etc.).
Unfortunately I've not come across a framework that will keep on checking properties continuously (say, in a round-robin fashion). Sticking the test command in a loop should probably work though: `while runTests; do sleep 1; done; notify "Tests failed!"`
PS: As for "AI", I think it's better to be more specific. LLMs certainly aren't needed for this; but fuzzers have been using GOFAI techniques like genetic algorithms for a long time!
https://developer.android.com/studio/test/other-testing-tool...