Even Tetris is hard to test
blog.jwhitham.org
blog.jwhitham.org
"Hard to test" in the submission title didn't mean what I thought: I though it meant it was hard to write tests for Tetris, not that it was hard to recover a complete specification of the game while playing it.
The big problem here is that code coverage tests don't help you cover what you should have explicitly defined or tested but didn't. As a result a lot of things end up still defined by implementation and not specification, as all sorts of important details only got defined during implementation.
any others?
int addOne(int x) {
if(x == 0)
return 1;
else if(x == 1)
return 2;
...
}
and int addOne(int x) {
return x + 1;
}
I always keep in mind the famous Dijkstra quotes about testing and program complexity:"Program testing can be used to show the presence of bugs, but never to show their absence!"
"Simplicity is prerequisite for reliability."
What you want to be really sure is state coverage, or input range coverage assuming your functions are pure. Now testing every function from int.min-int.max might seem unrealistic but what you have to do then is constrain the possible range of input or divide into ranges with special cases that you can somehow group together. Say for example int.min, negative numbers, zero, positive numbers and int.max.
Also, just because you covered a line doesn't mean it's correct, the only thing you've really tested is that the program doesn't crash. For the test to be really useful you also need a correct result of the output, added by a human. You can't just randomize input to increase the coverage.
Imagine now you have a switch with 50 different conditions that have nothing to do with each other and cant be reduced to a simple arithmetic operation.You'd have to test all the paths if you want a high test coverage rate.
All you can do is abstract decision making through FP or OOP(chain of responsability).
For me, it's a code smell to have a function called "updateScore4". It's often possible to come up with an algebraic statement that gives the same result as a bunch of logic (code paths).
I think userbinator's point is that better code is easier to test as a result of having fewer code paths. Of course there are some gotchas to be aware of (e.g., overflow), but overall I agree.
I guess it's easier to use a boolean logic example (in JavaScript):
return myNormalObject || myDefaultObject();
Your code might always execute the myNormalObject half and return it, and even though your code coverage is 100% of the lines, your tests might miss myDefaultObject() code path.
Perhaps some code coverage tools can take this into consideration, though, but then you're back to the original problem..
All of coloring is just arithmetic, after all, but is rather complex
Arithmetic in programming is typically about functions taking in numbers and returning new numbers, ie. returning a new number instead of mutating one of the numbers in-place. That sounds like the spirit of FP, to me. (Though of course arithmetic on fixed-size numbers falls short of this ideal when it comes to overflow and such, and we often don't handle this possibility.)
I don't doubt that there are other abstractions than FP and OOP. But arithmetic looks like FP, to me.
When aikah said "FP", I took that to mean something like using a higher-level "updateScore" function that accepts a "calculateScore" function as an argument. I certainly may have been mistaken, though.
And actually, especially in games, test plans are still poorly communicated. In the old days, it was awful -- you would have the publisher doing all the QA, and barely speaking to the development team apart from bug reports. QA still often doesn't get involved until the last half of the project, before which time nobody has been thinking much about testing.
As studios improve their production, this is getting better. As a programmer, I've had more collaboration with QA as I work, at its best including having our QA liaison talk out a test plan with me while I'm working the feature. With enough communication, hopefully this kind of detective work to figure out what to test can be avoided.
In regards to the article, i would wager the better scores will be found by QA than developers :-P
To have good test coverage you should both test all possible inputs and have proper asserts for those inputs.
Testing is hard, covering lines of code isn't. To put it another way: in a test the hard bit is the assert, not the call into the tested code. And coverage only reports on the former.
At the start of your function, check all the prerequisites, e.g:
if(x<0) throw "x should be non-negative. Got $x"
if(x>=n)throw "x should be smaller than n. Got $x and $n"
Add tests for edge conditions that do not throw, but, say, increase a global counting edge conditions hit: if(x=0) edgeConditionsHit += 1
if(x=n) edgeConditionsHit += 1
Then, write tests so that you hit all paths in the condition tests.If that doesn't hit 100% in the rest of the function, the function has code it doesn't need, or your precondition checks aren't complete.
Think about other edge conditions. For example, does your code special-case x=n/2? Add a check on top. And yes, that is implementation-specific, but there is nothing you can do about that.
Of course, you don't need the edge condition and implementation-specific checks in release builds.
With these in hand, you can also split tests into implementation-specific ones and contract-based ones.
If a e.g game contains a sorting in some place in the renderer, I can replace the quicksort with a mergesort as long as the renderer interface is still testing ok. The new sort algorithm may have new special case paths (even number of items vs odd for example) but it's not a concern of the renderer public interface. I may however have introduced a bug with an odd number of items here and the old code was 100% covered and now it isn't. So there is a potential problem and the 99% has actually helped spot it.
If the sorting is a private implementation detail of the renderer then there is no other place to test it than to add a new test to the renderer component only because the sorting algo requires it for a code path. This is BAD.
The proper action here is NOT to add tests to the renderer component to test the sorting code path, but instead to make the sorting visible and testable in isolation via its own public interface.
So one of the positive things about requiring coverage is that if you do it right, it will lead to smaller and more decoupled modules of code.
The bad thing is that if you do it wrong you will have your God classes and a bunch of tests coupled tightly to them.
Of course, this still isn't a guarantee. But it does make me feel better about my code.
So yeah, when a team claims 100% code coverage, usually that is just a signal that they care about testing and the quality of the code, therefore it tends to be less buggy. Not necessarily because 100% coverage itself made it so.
Really I only use a code coverage tool to check for important places that aren't covered at all, AFTER I have tried to think of the proper behavior/spec of the code from an outside perspective. It's like a secondary check after you think you are already done. That keeps you focused on what correct input and output are, and then patching up the little areas that you missed with a tool.
»Our customers tend to be makers of aircraft or car
parts. Both businesses have strict safety standards
which involve coverage testing, and our tools help you
produce the relevant reports for certification, like
DO-178B for the aviation industry.«
I guess in both industries the value of more testing cannot be understated.And yes, it can. Are you trying to state their value compared to what? A good schema for task division, with encapsulation and whatchdogs? Proofs of correctness? Proofs of halting? Testing is much less valuable than any of those.