Functional Tests as a Tree of Continuations (2010)
evanmiller.org
evanmiller.org
- You usually have shared initialization nearer the root and the various cases you want to assert at the leaves.
- You want to group related tests logically together, so it's not one huge flat namespace which gets messy
- You want to run groups of tests at the same time, e.g. when testing a related feature
Typically, these different ways of grouping tests all end up with the same grouping, so it makes a lot of sense to have your tests form a tree rather than a flat list of @Test methods or whatever
Naturally you can always emulate this yourself. e.g. Having helper setup methods that call each other and form a hierarchy, or having a tagging discipline that forms a hierarchy to let you call tests that are related, or simply using files as the leaf-level of the larger filesystem tree to organize your tests. All that works, but it is nice to be able to simplify define a tree of tests in a single file and have all that taken care of for you
That's such a great condensation of why automated tests are worthwhile.
"To write your own testing framework based on continuation trees, all you need is a stack of databases (or rather, a database that supports rolling back to an arbitrary revision)."
PostgreSQL and SQLite and MySQL all support SAVEPOINT these days, which is a way to have a transaction nested inside a transaction. I could imagine building a testing system on top of this which could support the tree pattern described by Evan here (as long as your tests don't themselves need to test transaction-related behavior).
Since ChatGPT Code Interpreter works with o3-mini now I had that knock up a very quick proof of concept using Python and SQLite SAVEPOINT, which appears to work: https://chatgpt.com/share/67d36883-4294-8006-b464-4d6f937d99...
And by eventually somebody, I mean two days ago me.
I'd much rather just have a utility that copies the database actually and hands my test the one it's allowed to mess with.
One idea is to separate scraping from verification. The latter would run very fast and be reliable: it only tests against stored state.
Then scraping is just procedural, clicking things, waiting for page loads, and reading page elements into a database.
Some consequences are needing integrity checks to ensure data has been read (first name field selector was updated but not populated), self-healing selectors (AI, et al), and certifying test results against known versions (fixing the scraper amid UI redesign).
A lot of effort is saved by using screenshot diffing of, say, React components, especially edge cases. It also (hopefully) shifts-left test responsibility to the devs.
Ideally, we only have some e2e tests, mostly happy paths, that also act as integration tests.
We could combine these ideas with "stacked databases" from the article and save on duplication.
Finally, the real trick is knowing, in the face of changes, which tests don't have to run, making the whole run take less time.
If they are slow, it means your application is slow. Good thing your tests make you realize it so you can work on it.
If they are flaky either your application is flaky or your UI is hard to use. Anyway that's something your tests tell you you have to fix.
And last: if your tests are all independents, why not run them all in parallel? With IaC you should be able to provision one instance of your architecture per test (or maybe dozen tests) easily.
With IaC, emulating a constellation of all dependent services along with the site is technically feasible. (There are other possible constraints.)
What's your ideal scenario? For example, k8s + cloud, ephemeral db, auto-merged IaC file of a thousand services, push-button perf testing, regression suite with a hundred bots, etc.
On premise k8s cloud where you can deploy many instances of the exact same services you have in prod. Let's say a E2E test takes 5s to run, your deployment takes 2mn and you want to stay under the 5mn line for running your test suite: deploy an instance of all your services and their databases per batch of 30 tests.
I can understand this not really being possible at an Amazon scale. But for most businesses? A good beefy server should be enough.
It turns out that using some evil macro magic, each test re-runs from the start for each inner section [1]. It also makes deduplicating setup code completely painless and natural.
You just have to get over the completely non-standard control flow. It's a good standard bearer for why metaprogramming is great, even if you're forced to do it in C/C++'s awful macro system.
[1] https://github.com/catchorg/Catch2/blob/devel/docs/tutorial....
If you create a query language, then the state can be verified to match expectations at any point.
I'm not sure why we don't program like this.
I don’t really know what you’re talking about, and have a hard time imagining how ideas from relational algebra can be applied to all APIs.
For example, many database-like things already use relational algebra and an actual query language, for sure. But how does this apply to, say, a GUI toolkit or an audio device driver?
it may be that a bug only shows when hundreds of transitions are performed (like an overflow or bug due to large data), but that's more stress testing. many bugs have repros involving a few state transitions.
relational algebra is a useful tool in my opinion because much of programming involves adding/removing things from sets or testing for their membership in a set. also relations are powerful as they can express recursive ideas like which widgets are contained within others (from the GUI example).
relations also allow defining invariants at a high level which must be true at any state. (eg, there should be no state like: audio_buffer_is_empty and audio_playing)
additionally we have languages such as SQL or for example https://alloytools.org/applications.html that can help programmers specify this in a familiar way.
Is it necessary to be able to exhaustively enumerate states? I remember a project I worked on some years ago which had me sketching out a statechart for a system which was not structured as an explicit FSM. The end result was surprisingly complex. I imagine some systems might even have an unbounded number of “states.”
I’ll definitely read more about Alloy too.
test_fork<A, B, C>(a: A, f: (A) -> B, gs: List<(B) -> C>) -> List<C> {
fs.map(g -> g(f(a)))
}
As you can see, the first step of the test is re-run from scratch every other branch of it.Any collaborations having observable side effects such as these are very difficult to prove correct, regardless the test approach employed.
>> Any collaborations having observable side effects such as these are very difficult to prove correct, regardless the test approach employed.
> yes but resetting the environment before each test is one way to deal with them
Not really.
Resetting an environment for each test requires distinct tests exist for all anticipated workflow permutations. While this is onerous when the side effects are limited to in-process mutable state (such as global, session, and thread-local data), it is infeasible when global state is a persistent store[0].
But I can't find a nice way to have pytest make a test per node in the tree. We end up with a single test for the whole tree which is less than ideal for dev experience.
Anyone with pytest hacking skills and an idea?
But I can see how this approach allows for parallelism as well, I especially like the fact that you only get one failure in case one of the steps fail
The lesser evil is to just ”do what you need and test everything once you are arranged”.
You won’t get hundreds of neatly separated well-named test cases which fail for a single reason. But for slow tests that isn’t as important as keeping the redundant setup away.
I like the tree idea but once we have simple pure/immutable we don’t really have the problem of redundant setup being slow, just ugly.
Setup: login. Test 1: delete your account. Test 2: can change username.
The last time I saw this handled was a test copy of db per thread and transactional test rolled back. Not great but it did 10x our pipelines and avoided locks and issues.
BUT obviously you can just test deleting the account at the end of the test that modifies the account.
SetAccountUserName(account, "New name"); GetAccount(account.id).UserName.Should().Be("New name");
SetAccountEmail(account, "new@email"); GetAccount(account.id).Email.Should().Be("new@email");
DeleteAccount(account); GetAccount(account.id).Should().BeNull();
This is a working timeline
But they commit an entirely new test filled with redundant stuff that takes way longer than is necessary and makes it unclear which assertion was the critical one. Because hey, look at all that green text in my PR. I'm so thorough.