A lot of us would consider SQLite to be high quality and relatively bug-free. The extensive test suite they've built up that exercises each release is a huge reason for it.
Of test code quantity, Coplien writes:
>If your coders have more lines of unit tests than of code, it probably means one of several things. They may be paranoid about correctness; paranoia drives out the clear thinking and innovation that bode for high quality.
> - Keep regression tests around for up to a year
> - Throw away tests that haven’t failed in a year.
SQLite appears to keep their tests. Somebody files a bug; SQLite writes a test that reproduces that bug; the test code remains long after the bug is fixed. This prevents the bug from reappearing. (Isn't that the main purpose of "regression" in the phrase "regression test"?)
I also think it's better to keep relevant regression tests for years. E.g. Consider that there's piece of code that's been working correctly for 5 years with 5 years of passing regression tests . Imagine a new programmer wants to rewrite the code to optimize it for speed and reduced memory usage. I think we'd feel much more confidient if the new code passes those same regression tests that were accumulated over 5 years.
As for test code size ratio... A lot of good comprehensive tests will have LOC outnumbering the actual code being tested. This is especially true for library code that's used in many places up the stack. I wrote string parsing routines and a reverse Boyer-Moore search routine where the test code (test edge cases, test nulls, test string sizes at 2^32 boundaries, etc) was 10 times larger than the actual code.
Of testing's utility, Coplien writes:
- Testing can’t replace good development
- [...] Tests don’t improve quality: developers do
... which looks like a strawman and a false dichotomy. Can anyone cite a credible development philosophy that believes testing can replace bad developers or bad process?We could say that about <ANYTECHNOLOGY> such that <ANYTECHNOLOGY> can't replace quality developers. Garbage Collection doesn't improve quality, developers do. Array boundary checking doesn't improve quality, developers do. And so on.
Or maybe there's a difference in terminology? I wonder if Coplien considers SQLite testing "system test" or a "unit test"? Does he consider SQLite "white box testing" or "black box testing"?