Analysis of whether unit testing obviates static type checking (2012)
evanfarrer.blogspot.com
evanfarrer.blogspot.com
(for reference, it's the same year when TypeScript came out. PEP 484 came out in 2014.)
Obviously a compiler can't prove your code is correct in every possible way, but any time the compiler can prove something correct without me needing to write a test for it, that's a win. Stronger type systems and static type checking is a great way to do that.
But one of the aspects of types that I find useful is the documentation they provide. In a dynamic language, you often have to guess what the types of functions are by reading the implementation, and not just of the function in question, but the functions composing that function. As a code base grows in size, this can become annoying. With code base size, modularity becomes increasingly important, and modularity benefits greatly from clearly defined and stable interfaces. Also, types can enable something like LSP to make suggestions about what can appear in a given place. With type inference, this can be quite powerful, allowing you to build code in an interactive fashion (proof assistants make use of this behavior).
Unit tests can also provide value as documentation, in the sense that they can function as tested and working examples of, say, how a library could be used.
So, in other words, both types and unit tests have value, and the value they offer will vary depending on the type system and how types are being leveraged in a given situation. There exists an overlap in the practical value provided by types and by unit tests, and where there is overlap, I would say types have the upper hand in principle, if not always in practice.
For me this is one of the core aspects. Self-documenting code makes me so much more productive, and being able to reason about code thanks to the types without having to consult the documentation is very powerful.
Any time I work with JavaScript, Python or similar I invariably waste time having to do print(dir(foo)) or similar, just to figure out what the hell some function returned and what I can do with it.
My feeling is that this is correct, but with a caveat big enough to drive a truck through: for often very different definitions of "working".
In a language with a strong type system, I can usually get something working (correctly!) fairly quickly with very few (and sometimes no) unit tests at all. In a dynamically-typed language, or one with a weak type system, unit tests quickly become essential to avoid simple errors that end up causing abnormal execution termination.
> without having to wrestle a type system for things you don't care about (e.g. malformed input leading to crashes when parsing machine-generated, standardised inputs, as is the case for three of the test programs
I've found that it's quite often that someone thinks a particular condition can't be triggered, when in fact it can be. Data gets corrupted, through coding errors, bad data entry, and many other things. And when you really do know that something can't ever happen (or have decided that the right thing to do is to terminate in that case anyway), most statically-typed languages have escape hatches, like `Option.unwrap()` in Rust.
> keeping in mind that, in Python, a crash due to a type error produces an orderly, if surprising, program exit rather than, for example, memory corruption
Most statically-typed languages out there will also abort cleanly if you do manage to trigger a type error. Languages like C & C++ are the exception here, not the rule. (Not to mention a type error is much much much less likely in a statically-typed language, so this is kinda a weird comparison.)
In the past year I've seen at least half a dozen posts here on HN making this exact claim. Now I kind of wish I'd saved a few.
and the reference:
[3] J. Spolsky and B. Eckel, “Strong Typing vs. Strong Testing,” in The Best Software Writing I. Apress, pp. 67–77, 2005.
I do agree that development time is an important factor and one that I didn't address simply because that would require a different type experiment. I hope that researchers look into that. I do also think that in addition to development time that overall maintenance time is considered. For example its plausible that dynamic languages are faster to develop, but take longer to make changes due to it being harder to understand, and changes aren't guided by the type system. It is also possible that dynamic languages are both faster and have lower TCO. To be clear I have no idea which is faster, and which has a lower TCO. I'd love to see some scientific evidence on development time and TCO.
In my experience, development time is much less of an important factor than management would like to think it is. I've seen companies burn customer goodwill by releasing a quickly-built, bug-ridden product too early, because they wanted to be first to market or whatever.
And I've also found that the companies that push for faster releases (without reducing scope, of course) end up shipping late anyway. If they'd accepted a later deadline in the first place, fewer coding mistakes would end up getting made along the way. (Stressed-out developers rushing toward an unachievable deadline make more mistakes than those whose time estimates are listened to.)
I know I don't represent all developers, but I do much better when I write in a "slower" language when that language's compiler verifies more things before we get to the point of running the code. The end product is more stable and more correct than what it'd be otherwise, and I don't think I deliver slower in a way that's significant to the business.
"If your unit tests provide good code coverage, don't feel too paranoid about giving up compile-time type checking." (pp 69)
"The only guarantee of correctness [...] is whether it passes all tests which define the correctness of your program." (pp 75)
"To claim that static type checking constraints in C++, Java, or C# will prevent you from writing broken programs is clearly an illusion" (pp 76)
The closest I could find to your quote was at the end: "a dynamically typed language could be much more productive but create programs that are just as robust as those written in statically typed languages." (pp 77) In the context of the article (see previous quote), this is presumably referring to the goal of "defining the correctness of your program", and the article in fact makes the point which I made reference to, which is that incorrect inputs and API usage could well be out of scope for this goal (and perhaps were in the programs you tested). Or, to put it another way, it doesn't make sense to talk about robustness unless you also give a context.
The worst I could say about the article is that it is kind of tautological, because it's easy to no-true-Scotsman the concept of "sufficiently good" unit tests. But to be honest this also happens with type systems.
I realise there's a lot of ... let's say "enthusiasm" in this particular area, but I found the article quite a bit more measured than what was implied by the blog post.
I have seen literally that exact claim, word-for-word, from people who it seems reasonable to characterize as (rather zealous) "proponents of dynamically typed programming languages". It might be a weak man[0][1], but it's definitely not a straw man.
0: https://slatestarcodex.com/2014/11/03/all-in-all-another-bri...
1: https://slatestarcodex.com/2014/05/12/weak-men-are-superweap...
That's an interesting question for any program or app in the move-fast-break-things domain. But perhaps a suboptimal, and potentially inaccurate, one for high-assurance systems that can't break without losing substantial financial value or lives, and/or for which system-wide refactoring may be necessary as the system evolves. Know your domain, choose your tools and methodologies appropriately.
But that said, dynamically typed code is OK for small code bases you can fit in your head.
The real question is does "tie yourself in knows" type systems like Haskell have a benefit over "80-20 does it" type systems like Python with type hints.
Paraphrasing Ed Catmull in his book Creativity Inc, the cost of preventing errors is usually much higher than the cost of fixing them.
I actually think there is literally no downside to making typing opt out rather than opt in - beginners benefit from the compiler saying they accidentally swapped two arguments around - how could they not?
I have helped beginners with untyped Python codebases and it's a lot harder to tell what function parameters are supposed to be than in Java or C++, leaving aside any other considerations.
Mentally parsing the type errors is an art though - generally you want to ignore pretty much all but the bottom 1 or 2 lines - the rest is a bunch of very verbose information at a much higher level of abstraction than you care about.
I suggest that testing that's enough to have confidence that a non-trivial program is correct is going to cover just about everything static type checking could find as well (except in dead code, and how does that matter?) Or, to put it another way, if you didn't care enough about the correctness of your code to test it adequately, why would you care that static type checking could find some bugs? You've already determined quality is not too important.
I think you got this backwards. If you didn't care enough about the correctness of your code to even do something as simple and straightforward as using a proper type system, why would you spend enormous amount of time to write and then support unit tests just so that they could find some bugs?
The static type checker is only useful (in the sense of making the software more reliable) if the testing is inadequate. In which case, the software is going to be crap anyway.
You might argue that enforcing static type checking cause the software to be structured in some way that's helpful, but that's a different argument.
Automated testing is only useful (in the sense of making the software more reliable) if manual testing is inadequate. In which case, software is going to be crap anyway.
You might argue that enforcing unit tests cause the software to be structured in some way that's helpful, but that's a different argument.
---
If you have a counter-argument against this in manual vs automated testing debate, then you also have a counter-argument against your own position above.
Also, from my personal experience: I've started to write more and more functional code lately. Not only static type checking, but "almost-Haskell" in Typescript with fp-ts. And I find that when I write code in this way, the only place where I actually find and fix bugs are... unit tests.
Unit testing is not the only form of automated testing. You can have automated regression tests, integration tests, and even automated exploratory tests (similar to what you get with the "million QAs typing on keyboards" approach of manual exploratory testing).
Automated tests are reproducible and reliable (particularly for regression testing or with tight timing constraints) in a way that manual testing will never be. You sound a bit like my old boss who had to be convinced to let us automate tests by sitting him at the relay box and telling him, "This test procedure says 'flip the switch 10 times per second for 10 seconds and a fault indicator will appear', can you do that?" The answer, of course, was no. No one can, automated tests are useful for making reliable software (this was a safety critical system, we cared very much about reliability and if you've ever flown you ought to be happy that we cared).
To give some idea of what automated testing can do these days, I point to the testing I regularly do on Common Lisp implementations, and particularly on the compilers of these implementations. I can quickly set it up to run some 200 tests per second in a single thread on SBCL, with randomly generated code of moderate size (a few hundred cons cells). Over the years I've run billions of distinct tests with this approach.
This sort of testing found large numbers of bugs in every implementation I've ever tried it on. Similar testing on compilers for other languages (like Csmith for C) found bugs in every compiler it was ever tried on, even those implemented in statically typed languages.
Depends on the type system. For example, if we have parametricity then we can constrain possible values quite a lot; e.g. a value with type 'forall t. t -> t' is the identity function (or diverges, if our language is non-total). If we don't have parametricity, such values could do all sorts of shenanigans, e.g.
def f[T](x: T): T = x match {
case n: BigInt if isOdd(n) && isPerfect(n) => n+1
case _ => x
}
This acts as the identity function for all values of all types; except if we give it a BigInt which is an odd perfect number. This is probably the identity function, but ensuring there are no edge cases would require our test suite to solve a famously hard problem https://en.wikipedia.org/wiki/Perfect_number#Odd_perfect_num...(This sort of guarantee is known as a "free theorem", e.g. see https://bartoszmilewski.com/2014/09/22/parametricity-money-f... )
Have any of the python projects acquired type annotations?
The false premises is that anyone ships bug free code. It can be done but it's stupidly expensive. Formal methods Z proofs, etc.
Unit testing and static typing are 2 different techniques which eliminate bugs.
The question that needs to be answered is which of the 2 techniques is the most effective per developer hour spent.
Unit tests do not caught all the bugs that static typing would catch nor does static typing catch all the bugs unit tests would catch. That's not an useful observation.
Nor is there a moral here that you should be doing both, doing both is expensive and costs $$$.
These are different things. I do think a type system dramatically reduces the size and complexity of unit tests, though. I don’t miss writing tests to check types.
I think it's absolutely certain that, to achieve the same level of confidence in your code, you need more unit tests for, say, a Python code base, than for a Java code base.
"If it compiles, it’s correct" is obviously false, and no one should be saying it about Haskell code. "If it compiles, it works" is true in certain contexts and for suitably weak definition of "works" (such as, behaves in a sensible way and provides a sensible answer, just not necessarily a correct one) and is an experience attested to by many, many Haskell programmers.
---
I don't get how this post says anything about unit testing or static typing at all. Doesn't it just imply that bugs are easier to detect with one method than another method? Most statically typed languages out there have a type system too weak to make the expression 1/0 a compile-time error. But you work around this with linting, unit tests, using a language with a stronger type system, etc.
Yes, but then how can you say
> I don't get how this post says anything about unit testing or static typing at all.
It does say something about unit testing and static typing!
And that's all that matters. Unit tests are the more effective technique.
Yes obviously static typing doesn't catch all the bugs caught by unit tests or vice versa. But it's a meaningless thing to say.
I find conversations with him enlightening because there does not exist a mainstream type system which can clearly an concisely express what he's doing there - even though these are fairly mundane things in dynamically typed languages, like arguments with correlated types.
I'm a firm believer in exercise as a means of maintaining skill, so I've been doing a hobby project in vanilla JavaScript. Of course I make type errors, but they're usually the sort that makes me go "aw shit, this field is wrong" and move on. Meanwhile I relearned some characteristics of this environment, which are:
-Tests are mandatory.
-You have to be radically explicit, otherwise you lose track of what you've been doing.
-Documentation is the cornerstone of remaining on track in the long run.
-The biggest mistakes are not in the types, but, unsurprisingly, the logic behind the code.
-General advice, like short-circuiting edge cases, using pure functions and immutable data structures still applies.
-As per Grug (https://grugbrain.dev/) - "grug very like type systems make programming easier. for grug, type systems most value when grug hit dot on keyboard and list of things grug can do pop up magic. this 90% of value of type system or more to grug" - wise tell this, me concur
Some interesting features of the codebase so far:
-There's considerably less code than in my imagined type-annotated version.
-I still use JSDoc to label inputs and outputs.
-I'm sort of cheating, because I rely heavily on classes, which have the `myObject.constructor.name` property, which in turn serves as runtime type information.
-Data structures are flatter - I have e.g. an array (named "path") where even indexes are strings, odd - numbers. There might be a way in TypeScript to express this(and there's definitely one in some über "strongly" typed language that compiles to JS), but my experience is that finding this out would be the perfect nerd snipe for me. I could have an array of objects with proper fields, but these are not initialized in the same place and a "path" ending in a string is also valid, so I would have to deal with the bureaucracy of optional fields.
---
Overall you probably don't want to do the same in a team - unless you're doing a group exercise and this is not code meant for production. Static types are to me a whip for those who take the path of least resistance - every developer has such moments, some just more frequently than others. That being said you might want to occasionally try working without the whip cracking above you and see what you create then - my guess is that if you don't, you'll miss out.