Code Coverage for Arbitrary Languages Made Easy [pdf]
semdesigns.com
semdesigns.com
Widely accepted, but still surprisingly easy to game if you want to do it (and have lazy or no code review). Code coverage only tells you that some test has called the code, not that it has actually checked the results it gets. So, if you write a unit test that "exercises" all the various branches of a method and simply ignores (or performs only minimal checks on) the return values, you get 100% code coverage and, as a bonus, a test that's unlikely to ever break! This might still be useful to make sure that the code doesn't throw any exceptions etc., but it's contrary to the spirit of how unit tests should be written...
https://en.wikipedia.org/wiki/Mutation_testing
That said, any metric is only useful for people actually using it to improve something, and completely useless when used in an adversarial way against someone willing to game the system. Coverage is pretty useful for a programmer to see where test gaps are, and easily gamed when it's felt to be too onerous.
In a large code base I'm pretty sure you would end up throwing exceptions in lots of places.
At my company we have this language that is not mainstream (Oberon) and therefore there was no code-coverage reporting at all. I found this paper and realized it was super easy to write a minimal code-coverage analysis from our AST parser and code formatter (internal tools). Somehow it had never occured to me that it was this simple.
The transition from "no code coverage at all" to "seing which module is covered and which is not" was really insightful !
Thus every code refactoring becomes a pain, as the tests aren't really testing the expectations of a function call, rather ensuring every line of code is shown as being covered.
bool check_index_in_range(const std::vector<int>&v, size_t index) {
// BUG ON NEXT LINE: SHOULD BE < NOT <=
if(index <= v.size()) {
// BRANCH A
std::cout << "access was safe" << std::endl;
return true;
}
else {
// BRANCH B
std::cout << "access was unsafe" << std::endl;
return false;
}
}
void main()
{
auto v = std::vector<int>();
v.push_back(1);
v.push_back(2);
v.push_back(3);
// Assertion passes, prints that the access is safe.
// Exercises branch A
assert(check_index_in_range(v,1));
// Assertion passes, prints that the access is unsafe
// Exercises branch B
assert(!check_index_in_range(v,5));
// We now have 100% code coverage
// but still failed to catch a trivial bug.
// Even though the buggy line was actually tested twice.
// Code coverage is not a good way to measure testing
}I know that statement is true, but WHY is it true? I have never seen anyone look at a coverage report and take any useful action. Instead I see management looking at those reports for a number to rate people on, but no useful action is ever taken. When I've looked at the reports I keep finding trivial code that would be hard to test for little gain.
I eliminated the report and nobody misses it. Instead we report to management total number of tests run (test cases, and each assert). Which is easy to game, but since the number naturally increases fast isn't worth the bother.
In many, it is more a directional check. 100% is likely not to give you 100% confidence. But, 0% will make everyone less confident. And if you have two teams, one at 50% and one at 80%, it is probably safe to assume that the testing culture is better for the 80% team.
And why branch coverage, as opposed to path coverage? Largely because you hit the easy to measure things before you worry about the hard to measure ones.
We made code coverage a part of CI, and we've been using that to drive up code coverage. It's been working. Our tests have gone from ~25% to ~50% coverage. Breakages making it out to staging have gone down.
Also, some incentives are well aligned with this. If people see that things aren't easy to test, then they get rearranged to make them easy to test. Not all incentives are aligned, but not everything is perfect. This is why reviews are useful.
When I've looked at the reports I keep finding trivial code that would be hard to test for little gain.
If your trivial code is also hard to test, then this indicates architectural code debt. Why isn't the trivial code trivial to test?
I work in embedded where we have to control real hardware.
This method also encourages reusing working specs, code, and tests. We see that in DO-178C with the RTOS’s, etc.
This article suggests a way to have code coverage probes automatically added to various locations in our source. I think I would also need some way to read coverage information from the rust program itself, which seems straightforward. But I wonder if anyone knows if this has already been done? I know about the llvm-based coverage reports, but I don't know if they can easily be adapted to this purpose
You have to buy each piece of functionality they offer separately for each langauge, for example I can buy a Java test coverage tool for $200, and it'll still only run on Windows or in Wine.
https://en.wikipedia.org/wiki/DMS_Software_Reengineering_Too...
https://www.semanticdesigns.com/Products/Parlanse/examples.h...