On Rigorous Error Handling
250bpm.com
250bpm.com
This is the single best piece of advice in this article. The second thing you have to document is the postcondition in case of an error - what state is the program left in?
With both a normal and an error postcondition you can fully specify your program. Like the author, I'm convinced that most of the pain with error handling stems from programmers ignoring the consequences of an error. That's the reason why approaches that force you to deal with errors explicitly (Maybe a, Result e a, etc.) end up being more robust. Otherwise, a part of the program just ends up missing.
However, from a theoretical perspective, exceptions are superior. The reason is that the error postcondition really does represent a non-local exit. Just like an ordinary return statement, it should be implemented as one instead of forcing programmers to walk the stack by hand. The latter is both less efficient and more error prone. Additionally, resource management must be integrated with error handling anyway and exceptions provide a clean opportunity to connect the two. This is one of the things that C++ gets right.
If the stack is walked up automatically programmers aren't going to deal with errors explicitly. No?
* No throw: the function will not fail. Full stop.
* Strong exception safety: if the function fails the state of the object(s) is acting on is unchanged. This is similar to transactional atomicity guarantee.
* Basic guarantee. If the function fails the state of the objects is unspecified but valid (i.e. no invariant is violated), but data might be lost.
From the point of view of the caller of course no throw is the most desirable property, the strong and finally basic. Anything less than that (i.e. corruption, leaks, dangling pointers) is considered unacceptable.
Another important insight is realising that exception guarantees have little to do with exceptions and everything to do with postconditions in the return path: for example the same techniques used to guarantee strong safety on the face of exceptions also work to guarantee postconditions on the faceof multiple explicit retun paths.
This might be a great approach for some (plausibly very large) subset of cases, but it can't handle everything.
True, but you can always offer at least the basic guarantee, and you can always document what you are guaranteeing to the caller.
Also abort sequences are a thing so you can kinda-sorta unfire them (talk about compensating sequence!).
Not in a way that truly solves the problem. Any time you are coordinating multiple actions that are irreversible and may fail, you'll need some contract other than "either your transaction exceeds or everything is rolled back."
This is one of my favorite empirical studies of software: Simple Testing Can Prevent Most Critical Failures: An Analysis of Production Failures in Distributed Data-Intensive Systems (2014)[1] It says that, indeed, many catastrophic errors happen because of ignoring the consequences of errors when the handling code was either empty (explicit ignore) or just logged the condition, even when the language enforced error handling. A simple tool they wrote to recognize it would have prevented 33% of the catastrophic failures they'd studied in Cassandra, HBase, HDFS, and MapReduce. So even when programmers are forced to explicitly respond to an error, they handle it with what amounts to a ¯\_(ツ)_/¯. I speculate that it's because psychologically we don't want to think hard enough about what to do when things that seem exceptional happen.
[1] https://www.usenix.org/system/files/conference/osdi14/osdi14...
Go's error handling isn't perfect, but at least they defined a single type (error) and made it idiomatic to use it everywhere. The same approach could have worked with checked exceptions, resulting in a language that has two kinds of methods: those that always succeed and those that can fail. This would result in a "what color is your function" problem [1], but with only two, obvious choices, it's liveable.
But there is no consensus in Java for how to say "this method can fail for a variety of reasons". (Many Java programmers believe that declaring a method to throw Exception is bad.) So you have a tower of Babel situation where methods can throw dozens or hundreds of different checked exceptions, many of which are incompatible, and lots of exception adapting at the boundaries, and long chains of wrapped exceptions.
[1] http://journal.stuffwithstuff.com/2015/02/01/what-color-is-y...
BTW, exceptions don't exactly introduce the colored-function problem, because it's easy to catch an exception, handle the error, and stop the "color chain." In fact, that's the whole idea. With async/sync this either cannot be done, or, if it can, it comes at a significant cost.
BTW, that it's "almost universally hated" is more myth than reality. When there are polls at conferences, most developers actually say they like it a lot. The complaints are mostly not about the feature, but the choice of which exceptions thrown by methods in the standard library are marked as checked and which are not.
A good example is the Monad instance of Either in Haskell.
Once you're going to go to the effort to walk the stack, why not give the caller the opportunity to continue? It's a huge increment in power for what seems like a minimal addition.
My understanding is that, at the cost of a significant penalty for the (hopefully rare) exceptional case, error handling with exceptions can be faster than returning values like in C in the (hopefully common) successful case.
In TXR Lisp, I unified conditions and restarts into a single mechanism, which is called exceptions.
There are two kinds of handling frames: ones for which an unwinding takes place first, and ones which just intercept the search. Both are identified by an exception symbol which exists in an inheritance hierarchy.
It's all documented in detail here: http://nongnu.org/txr/txr-manpage.html#N-0146B946
There are dialect notes comparing with ANSI CL, and an example program shown in both TXR Lisp and a CL translation for comparison.
Generally the advice of “pick a few failure types and stick to them” is exactly right. You not only encourage error handling to take place but that code is likely to remain correct/complete over time.
Easily avoided by using the latest Python feature, added in 3.6: Formatted string literals, a.k.a. f-strings¹. Instead of using
"foo {} baz".format(bar)
you use f"foo {bar} baz"
1. https://www.python.org/dev/peps/pep-0498A coverage tool can be used to find any error handling code that wasn't tested. (But it won't help you find error handling that's missing altogether.)
For me, personally, this is backwards. As a programmer, I want to write error handling (especially for infrastructure), because it means I'll be able to work more quickly later. I won't have to debug through all these abstraction layers. It's the manager who always says "It (the demo = happy path) looks good, so it's time to move on to the next feature".
The nice thing about Erlang is that practically every line of code is an assertion and they’re all live in production. Such a huge advantage over development assertions that get thrown away for prod.
Another nice feature of the Erlang model is that, often, you can code the happy path and forget the error checking. Makes for much tighter/cleaner/easier-to-read code.
Is that the best one can hope for - to leave traces for my successor to pick up the pieces?
fun doSomething(arg: X): Try<Y>
Since the return signature needs adjusting, this leads to developers very consciously making the choice to either handle the error in the function, therefore avoiding adjusting the return type, or let the caller deal with it if it isn't logical to handle the error there.1. Things go wrong that are out of your control - network down, database down, etc.
2. Coding mistakes. Either in your code or input arguments.
In either case, why not let the end user decide how to handle the error? Sometimes it’s some type of retry pattern, others it just to have a big try catch block that logs the fatal error and alerts someone.
These should be handled gracefully at each layer, and an appropriate error thrown to the layer above.
If the system is database or network dependent, there is no graceful way to handle it automatically most of the time.
In particular, Python solves this by having a “raise … from” language construct:
try:
dangerous_operation()
except DangerousException as e:
raise ThisModulesException("Dangerous operation failed") from e
Then, the original exception (including its stack trace, etc.) is available as an attribute on the exception you caught.By wrapping the exception, the user of “thismodule” can simply call it by writing
try:
thismodule.do_thing()
except thismodule.ThisModulesException:
logging.exception("Failed to do thing")
othermodule.do_other_thing_instead()
And this user of “thismodule” is free from having to know that thismodule calls dangerous_operation() and/or raises DangerousException (which are probably both from a different module). This information will be shown automatically in the exception’s backtrace, including all line numbers of all wrapped exceptions, so it is not lost. But the code which uses the module is both shorter, simpler, and has less knowledge about internals of the module it is using.Only in the case of code where I really need to do something if a specific operation fails, for whatever reason, do I use “except Exception:” or its even more catch-all variant, the bare “except:” clause. And even then, I very often just use it to log a message or send an e-mail, and re-raise the exception again afterwards.