Test Anything Protocol
testanything.org
testanything.org
Honestly I would prefer if it used a custom line delimited format, why shouldn't the producer be able to produce 'message: assertion failure: foo > 0'? You could have some simple rules to escape new lines and you'd have a format which doesn't require producers to embed a complete YAML library to produce output, and it would be nicer for humans to read.
This is if you need structured metadata at all. It seems alright to me to just have free form metadata text associated with a test.
Otherwise, I kinda like the idea of a standard test output format... but in the end, 'exit code 0 == success, not 0 == failure' is enough for the vast majority of cases such as CI integration. The human can figure out the implications from the test failure based on mostly any form of textual test output format.
Test Anything Protocol - https://news.ycombinator.com/item?id=23473370 - June 2020 (29 comments)
Test Anything Protocol specification (2006) - https://news.ycombinator.com/item?id=11915487 - June 2016 (10 comments)
Test Anything Protocol - https://news.ycombinator.com/item?id=10030889 - Aug 2015 (1 comment)
libtap - Testing library for C, implementing the Test Anything Protocol. - https://news.ycombinator.com/item?id=2936877 - Aug 2011 (6 comments)
A while back I wrote out a list of what we really want out of a test framework (more than any existing framework supports, though many get much closer than TAP):
Just start your test's description with your "CUSTOM-KEYWORD", then aggregate on that, instead of "ok/not ok".
You can even easily plug into prove(1) to emit your own custom summary.
You can aggregate on anything, e.g. you could consider all tests taking longer than 500ms failures, regardless of top-level "ok" status.
Your use-case is rather niche, so it's not a failure of the protocol that it provides a boolean "ok/not ok" at the top-level, and not "failed known flaky" or whatever (although that's usually marked as "TODO" test).
It's intentionally minimal exactly to make it easy to support the sort of use cases you're describing here.
If the only controllable level of granularity is "run a whole executable file", you don't have control.
Actually running all tests is about the least interesting thing you can do when testing.
1. Should this result be treated as a pass or a fail? 2. Should a warning of some sort be emitted? (Ex: for FAST and SLOW) 3. Should the result be counted as part of the main test corpus, or as part of some auxiliary grouping?
For example, if you get MISSING outputs from tests under CI, you know something is wrong with your CI configuration. If you get either kind of WIP output, you shouldn't publish (FSVO publish - some projects might require this for every commit, some only for commits on the main branch, some only on tags).
It also makes it easier to diff test results meaningfully.
1..4
ok 1 - Input file opened
not ok 2 - First line of the input valid
ok 3 - Read the rest of the file
not ok 4 - Summarized correctly # TODO Not written yet
and I had two immediate objections.a) Shouldn't tests have names?
b) (more serious) To generate valid TAP output, I need to know up-front how many tests there are. What if I can run the tests in sequence but I don't know how to count them? This can happen if there is a tree of tests, where each node knows what children it has but not what children they have. There's no credible way to generate TAP output streamily, even if there is no parallelization of tests.
No thanks.
That's what the text following the test number is for.
> b) (more serious) To generate valid TAP output, I need to know up-front how many tests there are. What if I can run the tests in sequence but I don't know how to count them? This can happen if there is a tree of tests, where each node knows what children it has but not what children they have.
Typically you'd use subtests for this, e.g.
1..1
# Subtest: root
1..2
Subtest: branch 1
1..2
ok 1 - leaf 1
ok 2 - leaf 2
ok 1 - branch 1
ok 2 - leaf 3
ok 1 - root
But if you really didn't know the number of tests up front for some reason, the test plan is also allowed to appear at the end of the file. TAP version 14
ok tst:zero:1
ok tst:zero:2
..
ok union:1
ok union:2
ok year:len
1..215
Tapview is pretty portable, and works well as a simple output visualizer turning my view into: ./yal examples/lisp-tests.lisp | _misc/tapview
...........................................................
215 tests, 0 failures.b) No, you don't have to know prior how many tests will run. If you read the spec, you'll see there's an option to print the final count at the end.
[1] https://static.googleusercontent.com/media/research.google.c...
I wish we had a unified format for for reporting benchmark results too. Now I think about it, I guess it should be possible to extend TAP formst to add arbitrary extra measures alongside pass/fail.
I also wish testing tools would do more to force test authors to differentiate between different types of failures (the test thinks it detected a bug / the test setup failed / the test code itself divided by zero). Once your project has a corpus of tests that just "fail" it's quickly impossible to retroactively fix that problem.
Most testing frameworks allow you to skip with a message if setup failed, which can help
> the test code itself divided by zero
I think this is a bug you do want to know and should be rare once code is merged into master (there's a bug in your code, just so happens it is test code!). However if for some reason you rely on a non-deterministic value that can make your test code fail in some way I have used Python decorators in the past to mark a test as such and raise a specific message.
junit's xml format is also similarly valuable as an interchange format and was how I set up unit test reporting on some of my first golang CI pipelines in jenkins/hudson.
Not to criticise those languages, but they are hardly used in test automation anymore (if they ever were). This looks like a really out of date hobby project
I’ve been looking for something to run test assertions that’s available on base distro installs.
Test::More is always there.