Java Exception Handling
neverworkintheory.org
neverworkintheory.org
The worse production bugs I have seen involve complicated exception handlers that generated a second exception while trying to write to the filesystem or something. In that case some other exception handler now starts to execute further up the stack. It really is as if a new program is now compiled and installed on your production servers. A new program that has read/write access to the production db. A program that most likely was never executed in QA and has never been seen by any person.
It is a bit hyperbolic but I have never seen an exception handler over four lines long that did not contain a bug.
My attitude towards exceptions is like a fire alarm. If I hear a fire alarm it means one thing and I have one action to take. Any notion of providing different pitched fire alarms to encode information would be absurd and dangerous, that is how I view exception hierarchies.
It is my understands that the Go language from google decided not to include this language feature.
I have heard that Java's checked exceptions are now seen as an important, failed experiment.
Sadly, this is what people often say these days, I still think they provide far more value than unchecked exceptions. I really liked it when I didn't have to provide any exception signatures and didn't have to think about where to catch them. I didn't like it so much anymore when unchecked exceptions started crashing my code, because there were exceptions thrown at points I never expected them to be thrown.
Either no exceptions or checked exceptions, but unchecked? My head :(
You must always assume that code in the exception handler can fail unless you have written every single line and it allocates no memory at all.
Let's say "more or less all": Yes, if there's an OOM because the system has not enough memory at all you maybe have to tell the user what to do instead of just crashing because the environment failed you, but in almost every other case the moment you accept the environment as parameter it becomes part of your code: If someone gives you a null object you can throw an IllegalArgumentException and show the calling program that it misused your API, but if you accept null and then call some method without checking for null the NPE you get is your bug, not that of your caller.
Edit: Rewritten to make my point clearer.
I've concluded that unchecked exceptions (above the runtime) and exception chaining are painful artifacts necessitated by using overwrought frameworks like Spring, J2EE, JPA, etc.
Most useful anti-checked exception TL;DR I've gotten is "What do you do when your code can't act (recover) on a checked exception, like some library fails on I/O?"
My answer remains "I handle it." After demo'ing my code, buddies conceded that my strategy works because I'm coding "close to the metal". In that case, it was an ETL/workflow engine (framework) that used hand-coded state machines (eg get next, data error, retry, restart).
Looking forward, I'm keen to learn Erlang and Elixir. I have the impression they would provide for free all the error handling and recovery I painstakingly hand-coded.
Basically, at those boundaries, you either have to say that the called code can throw anything - but then what about your own contract? you will either have to swallow everything; or rethrow everything, making it a part of your contract; or wrap everything in an exception type that you declare vas part of your contract, which is really a disguised version of rethrowing that doesn't add any meaningful value. Or else you say that the callee cannot throw anything, and then it is forced to swallow exceptions at the boundary.
Here's a very simple example of how it affects design: try implementing Java's List<T>, or even Iterable<T>, on top of a file. Something really simple, like one string element per line.
You'll find that it all works, except for error handling - because every read can fail, and neither List.get() nor Iterator.next() let you throw an appropriate exception type.
In contrast, in .NET, File.ReadLines will happily return an IEnumerable<string>, which will throw IOException when a given line couldn't be read - because IEnumerator is not limited with respect to what exceptions it cannot throw.
To properly deal with this problem, you need value-dependent types. Then you can do things like write, say, a generic implementation of map() that takes any random sequence, and throws the same exceptions that said sequence can throw when it's iterated.
The point of interfaces is to abstract from the concrete implementation. But that also means that the API designer cannot possibly know what types of exceptions concrete interface implementations can throw.
You're left with the choice of either declaring "throws Exception" on all/most interface methods or using unchecked exceptions.
That's why I tend to catch all exceptions and then either produce a fallback value, or signal back to the caller by returning an Optional (if the only useful thing to tell it is that it wasn't possible to produce a value) or by throwing a new specifically tailored exception and making that checked. In either case, the caller must be explicit about what they want to do when things go wrong.
I've seen good example of it just few hours ago, when my IntelliJ WebStorm (Java application) signalled, that it doesn't have enough memory (OoME) and suggested to update Xmx config setting and restart. Apparently, such exceptions should be handled at the root of execution stack, but sometimes, like in this case, there are ways to handle them with minimal damage.
Everything is unwound, everything is safe, special logic and wrapping not required (though wrapping can be valuable if there is information to add along the way).
So I wouldn't use that as proof that the world has settled on global "stop the world" exceptions vs checked exceptions.
Rust is another one that requires specific Result<DataType, ErrorType> to be passed back to a calling function, that either contains the return value or an error. But unlike Go, it's very difficult to ignore the error entirely, but it does require code to handle and propagate errors up the stack.
Although both Go and Rust also have the concept of a panic, that does stop the world (or at least the current thread, unless the panic is captured but that's another story...).
The problem with Go's error handling is that it doesn't play nice with using functions as part of larger expressions.
[0] Followed by multiple return values, duck-typed interfaces, automatic pointer dereferencing, channels, and a rich standard library. Sorry, this didn't begin as a rah-rah Go post.
result := f() * g()
becomes x, err := f()
if err != nil {
//handle f() err
}
y, err := g()
if err != nil {
//handle g() err
}
result := x * y
Even without exceptions I can imagine to have something like result, err := f() * g()
if err != nil {
//handle err
}
Of course that raises a lot of other questions if functions can have side-effects (but possibly no more so than shortcut logical operators)Collecting errors is just part of it. You also have to make sure nothing that has side effects gets called after the first one.
Without language support this is tortuous and error prone.
You mean delegating exception processing to a separate code path, a la the Strategy pattern? Or just straight up catch blocks?
Also, being from the Java era, an Exception was mostly a boring way to ensure error checking, while lispers did use exception as cut / tree jmp to provide solution values to an algorithm.
ps: funny today youtube suggested this https://www.youtube.com/watch?v=zp0OEDcAro0
"when your code discovers that something that was supposed to be impossible just happened, your program is no longer viable. Anything it does from this point forward becomes suspect, so terminate it as soon as possible. A dead program normally does a lot less damage than a crippled one."
I'd say in many cases it is perfectly desired to either swallow or re-throw exceptions. As just one example, many business rules can be correctly implemented as best-effort and if some resource (network or service) is unavailable simply let the failure propogate, or log, or backoff and retry some other time.
And most certainly but in fewer cases more robust handling is called for. Even if only 10% of exception handling is in this category, then try/catch is still a valuable tool.
I find it really faulty to reason that "this feature is only used 'correctly' 10% of the time, therefore it's a failed feature."
Let's say you want to read a file. You may get an exception for malformed paths, permission errors, filesystem errors and so on. It's unlikely that I will recover from any of these errors during runtime. The only thing I really can do is log the error, inform the user and send a report.
The checked exceptions problem on Java is one of overall language and libraries design, not inherent on the checked exceptions idea.
I prefer to catch Exceptions that I know how I want to handle and handle them accordingly. For other exceptions, I prefer to let them fly. If and when they happen, I can then understand the scenario and implement the desired logic.
e.g. rather than being able to write:
foodOnPlate.forEach(this::eat);
You end up having to write: foodOnPlate.forEach(food -> {
try {
eat(food)
} catch (Barf e) {
// handle? rethrow as unchecked?
}
});
This library does a pretty god job handling that case:https://github.com/diffplug/durian/blob/master/test/com/diff...
try {
Thread.sleep(1000);
} catch (InterruptedException e) {
throw new IllegalStateException(e);
}
try {
Thread.sleep(1000);
} catch (InterruptedException e) {
// never happens
}
In other words, the correct handling is to swallow or throw. The former is likely preferred by those who do assertion-based coding. The latter for people that prefer to discover their bugs in other ways.When I see code like the latter in PRs it drives me nuts, to the point that I start building up a case to have the knucklehead that does it moved off onto another project.
try {
Thread.sleep(1000);
} catch (final InterruptedException e) {
//never happens
throw new IllegalStateException(e);
}
or try {
Thread.sleep(1000);
} catch (final InterruptedException e) {
LOG.info("Woken up early", e);
}The problem with exceptions is that they allow you to put arbitrary distance between the point of detection and the point of handling.
Some applications need to fail fast, fail hard, and restart automatically.
Some applications need to present any failure or error condition to the user and give him power to decide.
Some applications need to be restored to some sensible state at all cost.
Needless to say, all possible error handling strategies, from C return codes to Java checked exceptions, have their fair use cases; discussing which one is better and which one is worse without specific project in mind, at least as an example, is completely pointless.
Another anti-pattern is subclassing Exceptions to have things like "UserNameInUse" exceptions during a sign up path to indicate that a selected username is not unique in the db. This becomes a poor man's message passing implementation. Using an exception to signal that is very inefficient as the exception has the entire stack trace as part of its data.
Technically, all the exception throw needs do is pop stack frames until it finds a handler. Getting a stack trace is optional, but keeping track of the return addresses (pushed by the CPU's call instruction) is sufficient to materialize that on demand. An array of integers isn't expensive either.
I don't believe transactional programs using exceptions for rollback is an anti-pattern. It's not the best pattern - a transactional language, with transactional data structures, is much harder to get wrong - but it's not worst, not by a long shot.
The article objects to the fact that applications handle many exceptions by logging them and giving up. What's wrong with that? That's perfect. For the class of exceptions that I can just log and give up, that's often what I want to do. If I had an open TCP socket to a remote host, then I can't "recover" from an exception indicating that the connection was broken. Instead, I just want to tear down everything related to that connection and carry on.
Exceptions can rarely be recovered from at the level of abstraction where the exception occurred. For example, I can't repair a socket that's been broken. I cannot retry the request at that level - it might not be idempotent; the socket processing code can't know. The connection pool doesn't know. The HTTP client might or might not know. However, I can recover from the exception at a higher level -- at the level of the code that understands what the task is, it might retry the top level processing job, which might invoke a call on a service client, which will invoke a call on an HTTP client, which will invoke a call on an HTTP connection pool to open a new TCP socket as needed. This is exactly the pattern that we see commonly in Java applications: a higher level module calls a lower level one, and then variously logs exceptions and tries again (possibly with exponential backoff over multiple attempts), or logs them and give up, or does something else.
You don't really "repair" or "recover" from exceptions. You mitigate them by handling them and carrying on with the job. To achieve this, you unwind up to the level where it makes sense to retry or abort a task. When you retry, you retry from the right level of abstraction needed to start over.
The extremely valuable property that exceptions have, that they give you, is confidence about what specifically went wrong. They constrain it. If I try to interact with a socket and I get an IO exception, then I know the error is constrained specifically to that socket. The rest of the program is working fine, and I can either try to recover from that error (sometimes possible), or close the connection and open a new one, or give up processing the task.
Exceptions are also valuable because you don't have to add error-handling logic at each layer in an application. Many layers can be ignorant or agnostic of the exceptions that will propagate through them, and that's a good thing because it reduces coupling. These layers might be inserted between the high-level logic that understands task processing, and that will make the judgment call about retry/backoff/give up, and the specific code that can fail and throw an exception. I don't want all layers of code to have to care about exceptions that might pass through them; they can't and shouldn't do anything about it.
It is valuable to be able to build abstractions and build code that exceptions pass through, without handling them. You might add logic to handle them later, or you might do so at a higher level, or a lower level.
If any of my low-level exceptions bubble up into the top-level catch block, when we're expecting to handle all exceptions at lower levels, then I have the ability to do something reasonable in my catch block: abort the task, emit metrics, and log the full stack trace as part of my "unexpected exception" routine. That's really useful: it's one way to discover exceptions that the system ought to handle but isn't handling yet. You might write one of these catch blocks with the expectation that it "should never happen", but if it does we'll detect it and supply a reasonable default behavior.
For example, if my application is a microservice, then I will wrap each microservice operation in a try/catch block, and in the catch block I will include logic to convert the exception into an appropriate failure return code from the operation. Or even better, my service framework will handle this for me -- that's an example of the decoupling that exceptions provide as an abstraction. If any microservice operation throws an undeclared exception, then the framework can surface it as the equivalent of HTTP 503. We might set the expectation that all exceptions should be declared exceptions, but if we ever miss one and have an undeclared exception, then the framework can return an appropriate error response.
The right thing happens by default: we don't need detect-and-propagate logic in each call frame. Some of those layers could also catch and retry the exception, or catch it and refine the exception type, or just catch and log it and emit task-specific log entries or metrics. For example, perhaps I have a try/catch block around some logic that calls S3. A catch block around the S3 calls can emit metrics describing the failure rate of calls to S3, which are useful to alarm on, and to examine distinct from the overall failure rate of my microservice operation which might call other services too. This is an example where a simple catch-observe-propagate-or-rethrow adds meaningful value.
At the same time, those failures could happen for many different reasons. Perhaps my credentials have expired, and so I'll see a 100% failure rate with that step: emergency. Perhaps I'm using a presigned URL, and it's past the expiration time. Perhaps there is a random transient failure. Perhaps the file I'm uploading is too big, or perhaps its checksum doesn't match, and so on. We can start with a basic naive catch block around S3 and refine it over time to handle these cases if we observe that it's relevant to do so.
Exceptions are an excellent foundation for confidently evolving system behavior in these sorts of ways. Plus, I can implement this evolution in the most convenient place for me. I don't have to go to a lot of trouble propagate all down the call stack to where it's thrown; that code can pass it on without knowing what it is. Yes, any of these exception handling blocks is coupled to the exceptions it catches, but higher level blocks aren't, and code in between can pass them on without understanding them -- that's valuable.
(If you feel that "the network should be reliable" is a valid answer, see: https://en.wikipedia.org/wiki/Fallacies_of_distributed_compu...)
I grant you that there is a third class of exceptions, "remote unavailable", for which the error handling is to retry X times, then promote to "system broken" and let the user retry later.
For all intents and purposes, "system broken" is a better name for "code error", covering also cases where the DevOps team should fix the hardware or dependencies on third party systems. [Configs are funnily written code].
Often times the "retry X times" logic is wrapped nicely in one's RPC infrastructure of choice, so what the app developer has to deal with is either "user error" or "system broken".
Exceptions happen all the time because of operational errors.
They allow you to confine errors on any single request to that request. Many languages/platforms suggest you to use multiprocessing for that.
Also helps that servers usually have well-understood failure modes (DB error, data validation error, network failure) compared to other kinds of software (e.g. a missing file may lead to a whole investigation there)
For e.g. desktop programming, exceptions are kinda less useful.
I don't understand this, returning error codes has the same effect. In fact you can do that with most forms of error handling, short of segfaulting?
At the moment, I'm mildly intrigued by Rusts result types which force you from the start to handle "all cases". They strongly remind me of checked exceptions.
Many languages support multiple return values, algebraic data types, or box types, allowing you to return an error value out of band.
And if the study is to be believed, checked exception are almost always dropped on the floor, just with more boilerplate and ceremony.
One benefit of not throwing an exception is that you can maintain referential transparency, which exceptions generally break. You can't substitute a function call with its value because throwing an exception will have different effects depending on the context, which consequently hinders the composability of such functions.
Exceptions are also generally not type-safe, e.g. the return type of a function tells you nothing about what exceptions it may throw. If you go the checked exception route like Java, you do get some degree of type safety, but you do so at the expense of higher-order functions, which can't reasonably be expected to know about the specific exception types its arguments may throw.
On the other hand, ADTs are composable. They work for functions that aren't defined for some inputs (maybe an input type that can't be constrained by the type system). They avoid bugs because they force the caller to deal with the exceptional case (unlike returning null, a sentinel value, or throwing a RuntimeException), but without the boilerplate of exceptions, particularly in languages with pattern matching. The caller gets to decide when, if, and how to handle the exceptional case.
Exception, on other hand, just happen and get caught in a few places.
Errors can be grouped into 3 categories: programming mistake, semantic error and non-deterministic error.
The first category are bugs that are discovered through sanity checking and defensive programming - conditional code that is only invoked in the case of a mistake and is always avoidable. Depending on the environment, a whole program abort or restart might be a valid response; a localized response (like handling an exception) isn't going to be useful. Fall back to the event loop / request-response loop, log and continue usually suffices for non-critical apps. Checked exceptions are not a good idea, but that's OK - Java uses unchecked exceptions for these, even if not all libraries do the same.
The second category are semantic errors, typically where a client or end user has tried to do something invalid, and something like exceptions come in handy to unwind a partial operation, a bit like rolling back a transaction. Starting out with optimism but rolling back once a problem has been discovered isn't an unusual or technically flawed approach; exception handling is perhaps an error-prone way to do it, but it is a way. And it's a case where localised exception handling again isn't a good idea. This is a problem area in Java, as far too many applications use checked exceptions here.
The third category is non-deterministic errors: something outside the world of the program didn't conform to expectations. The other side of a network socket disappeared; a file was deleted; a device was disconnected; etc. You can't detect the error before attempting the operation, and the resulting error is, in a way, a type of return value - unexpected information - from the attempt. Fixing the problem is almost certainly only likely to be successful if it's local (hacks like lazy creation of files, automatic reconnects of sockets, etc. aside). Propagating the error is fine if it can't be fixed locally - most code isn't written to have alternative strategies for these scenarios - but forcing callers to handle this specific type of error - that's not terribly useful. Again, checked exceptions - specifically, typed checked exceptions - not giving up the benefit they promised. The fact that something might fail in a non-deterministic way is useful information at the API level, but it also violates modularity. The "fix" in Java is for every module to throw its own hierarchy of checked exceptions, which, of course, is no fix at all: now you have even less prima facie information, and useless exception handlers and exception specifications proliferate until the meaning behind them is lost.
Other languages that use option types to propagate errors work fine for the third category; they usually have a different scheme for the first category; but on the second category, they are usually silent, and that's a bad thing in my opinion. Something exception-like is very valuable in the absence of strong tools for transactional modifications of program state. You shouldn't take away exceptions without providing a solid way of throwing away partial work. Erlang - tear down the process - that's OK; functional languages - immutable state necessitating immutable, persistent data structures - that's OK; something like Rust - well, it's going to need panic for more than just the first category.
No surprise. I've seen books that fail miserably explaining exceptions, some of then using wrong code for the examples.
Now there is also people that opposes the concept of exceptions. Not sure what the article conclusion is:
The only conclusion I can draw from this is that exception-based error handling has failed to achieve what its creators intended.
Is he talking about mandatory checked exceptions or any exception in general?
First, don't catch exceptions unless you really can handle them. Swallowing and emitting a log line is not handling an exception.
Second, and this is a corollary, only catch the exception classes you can handle.
The classic antipatterns of swallow-and-log or swallow-and-convert are very frustrating.
Swallow-and-log means that you don't see the true state of the world upon exceptional conditions arising. It is particularly harmful for gutting the usefulness of acceptance and unit testing.
Swallow-and-convert means that you deny the rest of the stack a chance to properly handle the exception. You also suppress the information which checked exceptions provide to consumers of your method, converting something that can be solved at coding time into runtime error lotteries.
Edit: per the comments below, my meaning for "swallow-and-convert" is the practice of turning meaningful exceptions into `RuntimeException` because it's "too hard" to add exceptions to consuming method signatures.
There is a book solely dedicated to the topic: http://www.amazon.com/dp/0131008528/?tag=stackoverfl08-20
Had it mentioned (together with another only-one, but fpor ASP) in an http://programmers.stackexchange.com/questions/14831/how-to-... answer, which the proud professionals then dewnvoted and eventually deleted :)
A.) String to Integer with null handling
for HTTP GET/POST transformation isn't
worth the effort, because you're users
are drunk idiots, and they won't
pay attention to your carefully stratified
error messages anyway. So just tell them
to sober up, straighten out, and fly
right, and barf a hideous, uninformative
stack trace at them to scare them off
until they go away, and come back in
the morning.
B.) Developers are so rushed to just get shit
done that NullPointerExceptions,
NumberFormatExceptions,
ArrayIndexOutOfBoundsExceptions and so
forth, are useful enough, and probably
don't need much more differentiation
anyway. As middleware developers, we're
really not dealing with anything too
exotic, and so we mostly don't need to
re-interpret things like sensor or transducer
state, or spin dial knob rotation angle,
or leaf spring load, or magnetic flux.
So, just because 90% of server side web developers don't use it, doesn't mean it's a failure. It just means that interpreting discete character data from noodniks on the other side of a keyboard and pointing device, behind a TCP/IP connection can only get so complicated.