If it's valid for a user to do, you can make a list and have it do those in sequence.
I had this for a UI library. It could call the functions to add and create the library and then afterwards would move through it. It was for the BBC so on TVs and could move u/d/l/r - the logic was regardless of the UI if you moved right and the focus changed then moving left should bring you back to where you were (u/d the same, etc).
That's tricky, yet being able to write
FOR ANY ui a person can construct
FOR ANY path a user takes through it
WHEN a user presses R
AND the focus changes
THEN when the user presses L
THEN the user is on the item they were before
Was actually quite easy to write and yet insanely powerful.
The one that really convinced me on PBT was one in this library where it found the bug, and the bug had an explicit test for it and it was explicitly in the spec but the spec was inconsistent! The spec was broken, and nobody had noticed.
Another that drove out a lot of bugs was similar but was that regardless of how many ui changes we made and how many movements the user made, we always had something in focus.
Anyway, the big thing here I want to stress is a series of API calls and asserting something at the end or all the way through.
Side note - oh my this is so long ago, 15 years ago building a new PBT tool in actionscript
JSON.parse(JSON.stringify(randomObject)) === randomObject
That often works. What also often works is generating the expected output and constructing the input from it. For example, a `stripPrefix` function that removes a known prefix from string, e.g. `stripPrefix("foo", "foobar") === "bar"`. Property test: stripPrefix(randomPrefix, randomPrefix + randomSuffix) === randomSuffix
Note, we "go backwards" and generated the expected output `randomSuffix` directly and then construct the input from it `randomPrefix + randomSuffix`.Reference implementation based properties also work very often. For example, we've been developing a JavaScript rich text editor. That requires a bunch of utility functions on DOM trees that are analogous to standard string functions. For example, on a standard string you can get a char at an index with `"foo bar".charAt(3)` and on a rich text DOM tree we would need something like `treeCharAt(<U><B>foo</B> bar</U>, 3)`. The string functions can serve as a reference implementation for the more complex tree functions:
treeCharAt(randomTree, randomIndex) === extractStringContent(randomTree).charAt(randomIndex)
The same can be done with all string functions like `slice`, `indexOf`, `trim`, ...> Without complex input spaces, there's no explosion of edge cases, which minimizes the actual benefit of PBT. The real benefits come when you have complex input spaces. Unfortunately, you need to be good at PBT to write complex input strategies. I wrote a bit about it here...
Here's the link: https://www.hillelwayne.com/post/property-testing-complex-in...
- "Metamorphic testing" is where analyze how code changes with changing inputs. For example, adding more filters to a query should return a strict subset of the results, or if a computer vision system recognizes a person, it should recognize the same person if you tilt the image.
- Creating a simplified model of the code, and then comparing the code implementation to the model, a la https://matklad.github.io/2024/07/05/properly-testing-concur... or https://johanneslink.net/model-based-testing
There's also this paper, which I haven't read yet but seems intriguing: https://andrewhead.info/assets/pdf/pbt-in-practice.pdf
It found so many bugs: file corruption, crashes, memory leaks, pathological performance issues. The kind of issues that standard unit testing doesn't find.
- Despite the terrible tutorial examples, PBT isn't about running one function on an arbitrary input, then trying to think of assertions about the result. Instead, focus on ways that different parts of your production code fits together, what assumptions are being made at each point, etc.
- You don't need to plug random inputs directly into the code you're testing. There are usually very few things to say regarding truly arbitrary inputs, like `forAll(x) { foo(x) }`; but lots more to say about e.g. "inputs which don't contain Y" (so run the input through a filter first), or "inputs which don't overlap" (so remove any overlapping region first), and so on.
- Don't focus on the random inputs; the whole idea is that they're irrelevant to the statement you're asserting (it's meant to hold regardless of their value). Likewise, if your unit test contains some irrelevant details, use PBT to generate those parts instead.
- It's often useful in business-type software to think of a "sequence of actions" (which could be method calls, REST endpoints, DB queries, or whatever). For example, "any actions taken as User A will not affect the data for User B". Come up with a simple datatype to represent the actions you care about, write a function which "interprets" those actions (i.e. a `switch` to actually call the method, or trigger the endpoint, or submit to query, or whatever). Then we can write properties which take a list of actions as input. Remember, we don't need to run truly arbitrary lists: a property might filter certain things out of the list, prepend/append some particular actions, etc.
- Once we have some assertion, look for ways to generalise it; for example by looking for places to stick extra things which should be irrelevant.
As a simple example, say we have a function like `store(key, value)`; it's hard to say much about the result of that on its own, but we can instead say how it relates to other functions, like `lookup(key)`:
forAll(key, value) {
store(key, value);
assertEqual(lookup(key), Some(value))
}
Yet we don't really care about lookups happening immediately after stores, we want to make a more general statement about values being persisted: forAll(key, value, pre, suf) {
runActions(pre) # Storing shouldn't be affected by anything before it
store(key, value)
runActions(suf.filter(notIsStore(key))) # Do anything except storing the same key
assertEqual(lookup(key), Some(value))
}That last test style you describe can be done with Hypothesis. I've had some good success testing both Python programs and programs written in other languages that could be driven from Python with it. Like a server using gRPC (or CORBA once) as an interface, driven by tests written in Python imitating client behavior.
(For context, some "stateful" things that I've tested using ordinary PBT include browser automation (Hypothesis + Chrome + ChromeDriver), window manager scripts (QuickCheck + polysemy + Yabai), and a dynamic binding implementation for the JVM (ScalaCheck + a bespoke DSL for testing that it works correctly with Futures))
There's a good discussion of stateful property testing at https://stevana.github.io/the_sad_state_of_property-based_te... but personally I'm more interested in the parallel/concurrent/linearisability aspect that's also discussed.