I wish we had a unified format for for reporting benchmark results too. Now I think about it, I guess it should be possible to extend TAP formst to add arbitrary extra measures alongside pass/fail.
I also wish testing tools would do more to force test authors to differentiate between different types of failures (the test thinks it detected a bug / the test setup failed / the test code itself divided by zero). Once your project has a corpus of tests that just "fail" it's quickly impossible to retroactively fix that problem.