A while back I wrote out a list of what we really want out of a test framework (more than any existing framework supports, though many get much closer than TAP):
A while back I wrote out a list of what we really want out of a test framework (more than any existing framework supports, though many get much closer than TAP):
Just start your test's description with your "CUSTOM-KEYWORD", then aggregate on that, instead of "ok/not ok".
You can even easily plug into prove(1) to emit your own custom summary.
You can aggregate on anything, e.g. you could consider all tests taking longer than 500ms failures, regardless of top-level "ok" status.
Your use-case is rather niche, so it's not a failure of the protocol that it provides a boolean "ok/not ok" at the top-level, and not "failed known flaky" or whatever (although that's usually marked as "TODO" test).
It's intentionally minimal exactly to make it easy to support the sort of use cases you're describing here.
If the only controllable level of granularity is "run a whole executable file", you don't have control.
Actually running all tests is about the least interesting thing you can do when testing.
1. Should this result be treated as a pass or a fail? 2. Should a warning of some sort be emitted? (Ex: for FAST and SLOW) 3. Should the result be counted as part of the main test corpus, or as part of some auxiliary grouping?
For example, if you get MISSING outputs from tests under CI, you know something is wrong with your CI configuration. If you get either kind of WIP output, you shouldn't publish (FSVO publish - some projects might require this for every commit, some only for commits on the main branch, some only on tags).
It also makes it easier to diff test results meaningfully.
1..4
ok 1 - Input file opened
not ok 2 - First line of the input valid
ok 3 - Read the rest of the file
not ok 4 - Summarized correctly # TODO Not written yet
and I had two immediate objections.a) Shouldn't tests have names?
b) (more serious) To generate valid TAP output, I need to know up-front how many tests there are. What if I can run the tests in sequence but I don't know how to count them? This can happen if there is a tree of tests, where each node knows what children it has but not what children they have. There's no credible way to generate TAP output streamily, even if there is no parallelization of tests.
No thanks.
That's what the text following the test number is for.
> b) (more serious) To generate valid TAP output, I need to know up-front how many tests there are. What if I can run the tests in sequence but I don't know how to count them? This can happen if there is a tree of tests, where each node knows what children it has but not what children they have.
Typically you'd use subtests for this, e.g.
1..1
# Subtest: root
1..2
Subtest: branch 1
1..2
ok 1 - leaf 1
ok 2 - leaf 2
ok 1 - branch 1
ok 2 - leaf 3
ok 1 - root
But if you really didn't know the number of tests up front for some reason, the test plan is also allowed to appear at the end of the file. TAP version 14
ok tst:zero:1
ok tst:zero:2
..
ok union:1
ok union:2
ok year:len
1..215
Tapview is pretty portable, and works well as a simple output visualizer turning my view into: ./yal examples/lisp-tests.lisp | _misc/tapview
...........................................................
215 tests, 0 failures.b) No, you don't have to know prior how many tests will run. If you read the spec, you'll see there's an option to print the final count at the end.