Error stack traces in Go with x/xerror
brandur.org
brandur.org
I used go mod replace to strip the stack trace collection out of pkg/errors, with a fork that does a no-op for that function call, and it was a significant improvement for our use case.
Is there a write-up or references on how one can achieve this? Sounds like a good practical use case for the replace directive.
and (2) a way to cut down on if err != nil { ... }
boilerplate that pervades all Go code.
Am I weird in liking the explicit error handling? :/There is no standard solution though, so you have to rely on 3rd party libraries or use platform specific solutions.
I have marginally more experience with Rust (not that much of either) which instead gives Result types, tuples of 'ok' and error values. And, crucially for this thread, they must be explicitly consumed.
That enforced error handling is novel to Rust afaik, and after a bit of getting used to, an excellent feature I think.
But I'm not sure that the idiomatic (because that's all it is really, rather than language feature?) Go is materially different from deciding all your Python functions will return Tuple[TOk, TErr] and sticking to it, or even that different from returning ok types only and raising exceptions, really.
Hardly, though Rust of course has merit in making it popular, to the point that people think it was introduced by the language.
Whereas if you compare Rust to Python, C(++), or Go as I was - having to consume a returned 'result' is more notable.
Well that's not true, that's completely orthogonal?
There's a go vet pass as well. https://pkg.go.dev/github.com/golangci/govet#hdr-Unused_resu...
Scala's Either[A, B] is another example.
Is Scala's Try[A] "specific"?
1) The aforementioned boilerplate code that makes up a significant chunk of many Go projects and adds no value.
2) With explicit error handling, however deep your call stack is, you are relying on everything in that stack to have done the right thing with regards to error handling. With exceptions, you have to go out of your way to screw them up.
I can't tell you how many times I've gotten dumb error messages like "Invalid string" from a 3rd party library that take forever to debug, and accidentally swallowing an error is even worse. A simple (even crappy) exception message + a stack trace is much easier for me to use for debugging than even the best handcrafted error message in 99% of cases.
I write somewhat equivalent amounts of Python, Java, and Go lately, and each one has their good and bad parts, but there are few things I dislike as much as Go's error handling patterns.
I remember well java codebases littered with:
try { .. } catch() { // Todo }
Usually cheaply inserted by the IDE.I'm pretty sure nowadays there are linters that will ensure you have to do some extra work to actually check in such code, but still...
But either way, you have to go out of your way to do that, which is sort of my point.
The one place where you do see dumb boilerplate and chances to screw up is in dealing with checked exceptions, which have been controversial since the very beginning. I think they are one of those things that seem like a great idea in theory, but end up not working out so well in practice, but (like people who appreciate Go's error handling model) there are people who disagree.
However it is very easy to do the wrong thing using Golang. Just one random library doing `fmt.Errorf` without the `%w` verb is sufficient to lost all information from there on. I would much prefer some kinda of annotation that does the correct thing by default (wrapping the error) instead of the "every error should be explicitly" approach of Golang.
yeah this indeed happens with checked exception and careless developers just wanting to get stuff to compile (we'll handle that property later) and the IDE helpfully providing the boilerplate that resolves the compilation error (and assumes you'll fill the body)
I don't think I agree with this argument. In many languages, there's no obvious indicator that any given function may or may not raise / throw an exception. And even if it doesn't throw an exception today, it might tomorrow, and it's easy to forget to update all callers.
Since Go makes errors a return value you have to actively discard the result (by replacing it with a underscore, e.g. `res, _ := getResults()`) or take the time to handle the error.
And just like an intermediary library in Go can swallow the error, so can an intermediary library in an exception language catch and discard an exception.
It seems to me that the result is more errors are properly handled in Go - because they are explicit - while uncaught exceptions often cause bugs that make it to production.
For context, I recently switched jobs from one where I wrote Python for 4.5 years to a job where I've been writing Go for about 5 months.
Right, but because `err` ends up getting re-used in so many cases, you can re-assign it and forget to do anything about it (which is what I've found in most cases where an error was inappropriately suppressed in Go).
In most cases, simply bubbling up the error to something that will generically handle all errors is the right thing to do, and in the case of exceptions, even if you don't know about one, that is what will happen. I just sampled a handful of the top Golang repos on Github, and found very few cases of anything more than the standard `if err != nil; return nil, err' pattern that just bubbles the error up.
This is a bad thing because it’s not narrow. It’s like a language’s “catch” statement not allowing you to catch specific exceptions. Once someone has decided to throw away an error with the underscore assignment, all future errors the underlying might return are swallowed as well.
In other words, it takes even more boilerplate to have narrower error exemptions than to have the generic exemption.
> It seems to me that the result is more errors are properly handled in Go - because they are explicit - while uncaught exceptions often cause bugs that make it to production.
An error in go that is blindly returned all of the way up the stack has no difference with an uncaught exception.
The alternative is Zig or Rust's thing, where you still have explicit error handling but you don't repeat the same four lines all the time. This is maybe helpful because in Go one has to read those lines carefully to see if anything unexpected is happening.
LoadData(&data)
Does this function return nothing or an error ?
https://play.golang.org/p/k-eTs3dC_wK
Note that Go playground also runs go vet, a static analysis tool.
- the code has failed and the developer needs to be informed what to fix
- the environment has failed and the user needs to be informed
Exceptions are for the former, which Go still supports via Panic(). Errors are for the latter. Which is why all the boilerplate Err code creeps in.
I don’t like getting stack traces from applications when I’m a user. It always feels to me like the application is only half finished. When the error is environmental (eg file system permissions) a stack trace just muddies the water with unnecessary output.
This is why I like Go’s distinction between errors and exceptions.
All this was figured out decades ago, that's why exception systems were invented, but language designers keep making the same mistake over and over again in an effort to simplify the unsimplifiable.
The main benefit of Go is that the people that defined the library ecosystem actually decided to handle errors. It's the bare minimum, but it's better than just forgetting that the error exists and not documenting it either.
The language is ... not the best, but the libraries tend to be more robust.
stream.map(f)
doesn’t allow f to declare any checked exceptions. We could catch and wrap everything but that obscures the useful code without any improvement in safety.What language are you thinking of when saying Go libraries do better at error handling?
Exceptions do exist in Go, it’s called “Panic”. This whole “Go doesn’t have exceptions” meme is completely untrue. It’s just not the preferred way of error handling because, unsurprisingly, most application errors are going to be environmental. But if you select an out of bounds index in a slice or don’t handle a nil pointer correctly you’d get a panic — because those are clearly developer issues that warrant a stack trace rather than environmental problems that need a friendly user facing error message.
Whereas exceptions fail safe, ignoring the error means not catching it, which means if the error does occur, it will load up the debugger.
This is the whole Java "checked exceptions" misfeature again.
The panics are there to be used only in cases where the further code execution is impossible. Not to inform the developer about an error.
Errors are a general solution to report an unexpected event during execution. The destination of such report depends on the part of code the program failed.
Personally my one gripe with Gos errors is that EOF is handled as an error.
Not sure I got this part. If you want to stop the program and don't to print the usual panic messages into the console - you can log\print whatever you want and call os.Exit().
At the same time you can log stack trace without panicing.
>Or at least I think it does.
Maybe, it's just that "the code has failed and the developer needs to be informed what to fix" applies to errors just a much as to panics. The only difference is that an error is the situation during the execution where you can continue the job and a panic requires execution to stop (in most cases).
At the same time panic are 'developers only'.
In SaaS, mobile, and appliances, the environment is entirely or almost entirely controlled by the programmer - a missing file is the developer's problem, and the user can't do anything about it.
In enterprise software, there is a third entity, the Administrator, that has (almost) full control over the environment, and usually there are layers of 3rd party software that control it directly. A missing file or permission error is useless to the user, most likely it needs to be logged somewhere that Admins check, and the user simply told to contact their Admin.
Finally, even in personal desktop software, the environment is jointly owned by the user and the developer - the developer is normally responsible for setting up the initial environment through some kind of installer (msi, deb, make configure etc.), and many environment issues are bugs in the installer, not user errors.
Of course, the same library may very well be used in all of these vastly disparate deployment scenarios - the library can't decide what is the source of an error, so making it arbitrarily decide between two possible kinds of errors is wrong.
Note that even the reverse is not clear; an index out of bounds error could be a user problem - if they are trying to access the 7th element of a 5 element list. The fact that you'd normally do
if userSelectedElement > len(arr) {
return nil, fmt.Errorf("Array index out of bounds")
}
return arr[userSelectedElement], nil
Is only because of the convention you were proposing - `return arr[userSelectedElement]` does the same thing alone, if we disregard the convention (in a memory safe language, of course!).Trust me, searching through a stack trace on a centralised logging system isn’t fun.
> Finally, even in personal desktop software, the environment is jointly owned by the user and the developer - the developer is normally responsible for setting up the initial environment through some kind of installer (msi, deb, make configure etc.), and many environment issues are bugs in the installer, not user errors.
Desktop Linux, yeah. It’s seldom that simple in servers though. SELinux, custom config, custom iptables rules, network wide UIDs (eg shared storage volumes), there’s so much that can go wrong the moment you do enterprise.
> Note that even the reverse is not clear; an index out of bounds error could be a user problem - if they are trying to access the 7th element of a 5 element list. The fact that you'd normally do
If you’re writing software that doesn’t do input validation and bounds check then you’re a failure of a developer. Sorry but this is the bare minimum I’d expect a developer to do.
A stack trace is still better than a one line error with no context. It's a kind of 80% solution - it's not ideal (a perfect error includes only the relevant context), but getting it's much better for 0 effort than error codes/values give you for free.
> If you’re writing software that doesn’t do input validation and bounds check then you’re a failure of a developer. Sorry but this is the bare minimum I’d expect a developer to do.
So what is the profound difference between doing input validation in your own code vs letting the array accessor do it?
Yeah, a stack trace is better then an “undefined error” type message. But the point of forcing error messages over exceptions is you’re enabling you’re developers to write meaningful error messages. So your point is moot.
> So what is the profound difference between doing input validation in your own code vs letting the array accessor do it?
- Meaningful error messages (eg has the code failed because of a bug or because of invalid user input?),
- security hardening,
- reducing potential undefined behaviours,
- thorough unit testing,
- self documenting code (the code clearly defines what the happy path is and when it’s possible for a user to break out from that)
Etc
I really liked their `check` proposal. I hope they will bring it back to life.
result, err := foo()
if err != nil {
return nil, fmt.Errorf("Error message: %w", err)
}
This pattern is just so commonplace, but it's four times as long.Contrast with Rust, which handles errors as `foo()?` or `foo().context("Error message")?`*. It's still explicit — I like it better than silently-propagating exceptions — but my functions don't wind up being 75% error handling.
* Where the `context` function comes from an error-handling library.
What saddens me is that these kinds of matters will never be fixed in Go. Go has a stubborn anti-feature mentality and while it does help preventing feature creep, it overall harms the language in the long run. For a language created and maintained by Google this is a huge missed opportunity.
I was curious about this and it’s true, I didn’t know these functions could err, TIL.
The Godoc states:
> It is conventional not to worry about any error returned by Printf
The stability of Go is something I value a lot, so I don’t mind stop-gaps like this too much on balance.
val, err := int_or_err(foo)
if err != nil { return nil, err}
return val.add(bar), nil
Vs result := int_or_err(foo)
return val.map(add,bar)
I'd even take some sort of result := int_or_err(foo)
return val.map(add,bar).tuple()
to return a val,err pair in order to match conventional go.Consistency is more important than conciseness. Clear is better than clever. Plus, a function invocation / .map() is heavier than an err != nil check / condition.
Don't get me wrong, I appreciate the Either pattern as an alternative to if err != nil, but I also appreciate the really dumb and straightforward approach of non-clever Go.
The vast majority of errors only stop or rollback the current action. An incredibly small amount of code uses errors to detect stop conditions or to retry.
However, Go's `if err != nil` is not more explicit than, for example, Rust's question mark operator. The former is more verbose than the latter, yes. But both are explicit in that we can know which line can return an error.
The `if err != nil` is probably okay if there are only a few lines that can return an error, but if most lines in a function can return an error, it will result in too much noise. A real world example where panic is abused as an exception-like approach because the "proper" error handling using `if err != nil` is way too verbose: https://pkg.go.dev/github.com/apple/foundationdb/bindings/go... (Actually Go's stdlib too sometimes abuses panic in a similar way.)
You just need to have some thought when writing your code:
let a;
try {
a = someFailingFunc();
} catch (e) {
// handle
}
// use a
// do work func doRequests(url1, url2 string) resp throws HttpException {
resp1 := http.DoRequest("POST", url1)
resp2 := http.DoRequest("GET", url2 + resp1.ID)
return resp2
}
vs func doRequests(url1, url2 string) (resp, error) {
resp1, err := http.DoRequest("POST", url1)
if err != nil {
return nil, fmt.Errorf("Error in req1: %w", err)
}
resp2, err := http.DoRequest("GET", url2 + resp1.ID)
if err != nil {
return nil, fmt.Errorf("Error in req2: %w", err)
}
return resp2, nil
}
Of course, we can write the second example with try/catch as well, but the whole point of exceptions is to be able to write the first one when appropriate.Of course. I'm talking about when you actually handle your exceptions you should be specific about the lines you're handling.
The issue is not explicit error handling it’s specifically Go’s, which is verbose and half-assed. Not entirely unlike java’s checked exceptions though unlike checked exceptions we have plenty of other (and I’d argue better) implementations of “explicit error handling”.
Go’s error handling is a relatively minor improvement on C’s, but we’ve gotten quite a ways beyond that since.
I think people's issue is that there's no "I don't care about errors, just show me a stack trace" option like you get with exceptions, or more or less with Rust's `?` if you don't use `.context()`.
I think stack traces should be mostly reserved for interactive debugging and should not be included in user-facing errors. Why should the users of your program care that foo() called bar() called baz() which then produced an error? They want to know what went wrong and how to fix it (if it's "their" fault), and that is much easier if they get proper, context-specific errors (e.g. "CLI argument 'limit' must be between 1-10"). And if you need a stack trace for debugging a problem you should simply use panic().
A common problem is that when multiple producers produce failing work items and send them to a consumer -- a stack trace will just show "panic -> consumer.doWork() -> created by consumer.startWork()". Gee, thanks. You need to track the source, so that you have an actionable error message. If the consumer is broken, fine, you maybe have enough information. If a producer is producing invalid work items, you won't have enough information to find which one. You'll want that.
The idea of an error object is for the code to make a decision about how to handle that error, and if it fails, escalate it to a human for analysis. The application should be able to distinguish between classes of failures, and the human should be able to understand the state of the program that caused the failure, so they can immediately begin fixing the failure. It's up to you to capture that state, and make sure that you consistently capture the state.
Rather than leaving it to chance, I have an opinionated procedure:
1) Every error should be wrapped. This is where all the context for the operator of your software comes from, and you have to do it every time to capture the state of the application at the time of the error.
2) The error need not say "error" or "failure" or "problem". It's an error, you know it failed. As an example, prefer "upgrade foos: %w" over "problem upgrading foos: %w". (The reason is that in a long chain, if everyone does this, it's just redundant: "problem frobbing baz: problem fooing bars: problem quuxing glork: i/o timeout". Compare that to "frob baz: foo bars: quux glork: i/o timeout".)
But if you're logging an error, I pretty much always put some sort of error-sounding words in there. Makes it clear to operators that may not be as zen about failures as you that this is the line that identifies something not working. "2021-08-23T20:45:00.123 PANIC problem connecting to database postgres://1.2.3.4/: no route to host". I'm open to an argument that if you're logging at level >= WARNING that the reader knows it's a problem, though. (I also tend to phrase them as "problem x-ing y" instead of "error x-ing y" or "x-ing y failed". Not going to prescribe that to others though, use the wording that you like, or that you think causes the right level of panic.)
3) Error wrapping shouldn't duplicate any information that the caller already has. The caller knows the arguments passed to the function, and the name of the function. If its error wrapping needs those things to produce an actionable error message, it will add them. It doesn't know what sub-function failed, and it doesn't know what internally-generated state there is, and those are going to be the interesting parts for the person debugging the problem. So if you're incrementing a counter, you might do a transaction, inside of which is a read and a write -- return "commit txn: %w", "rollback txn: %w", "read: %w", "write: %w", etc. The caller can't know which part failed, or that you decided to commit vs. roll back, but it does know the record ID, that the function is "update view count", etc.
The standard library violates this rule, probably because people "return err" instead of wrapping the error, and this gives them a shred of hope. And, it's my made-up rule, not the Go team's actual rule! Investigate those cases and don't add redundant information. (For example, os.ReadFile will have the filename in the error message, because it returns an fs.PathError, which contains that. net.Dial is another culprit. Make a list of these and break the rule in these cases.)
4) Any error that the program is going to handle programmatically should have a sentinel (`var ErrFoo = errors.New("foo")`), so that you can unambiguously handle the error correctly. (People seem to handle io.EOF quite well; emulate that.)
You can describe special cases of your error by wrapping it before returning it, `fmt.Errorf("bar the quux: %w", ErrFoo)`.
Finally, since I talked about logging, please talk about things that DID work in your logs. Your logs are the primary user interface for operators of your software, but often the most neglected interface point. When you see something like "problem connecting to foo\nproblem connecting to foo", you're going to think there's a problem connecting to foo. But if you write "problem connecting to foo (attempt 1/3)\nproblem connecting to foo (attempt 2/3)\nconnected to foo", then the operator knows not to investigate that. It worked, and the program expected it to take 3 attempts. Perfect. (Generally, for any long-running operation a log "starting XXX" and "finished XXX" are great. That way, you can start looking for missing "finished" messages, rather than relying on self-reported errors.)
(And, outside of HN comments, I would make that a structured log, so that someone can easily select(.attempt == .max_attempts) and things like that. It's ugly if you just read the output, but great if you have a tool to pretty-print the logs: https://github.com/jrockway/json-logs/releases/tag/v0.0.3)
Anyway, I guess where this rant goes is -- errors are not an afterthought. They're as much a part of your UI as all the buttons and widgets. They will happen to all software. They will happen to yours. Some poor sap who is not you will be responsible for fixing it. Give them everything you need, and you'll be rewarded with a pull request that makes your program slightly more reliable. Give them "unexpected error: success", and you'll have an interesting bug report to context-switch to over the course of the next month, killing anything cool you wanted to make while you track that down.
$time, INFO, reading from /tmp/some-file
$time, ERROR, FileNotFoundException at $whole-stack-trace
).Conversely, the Go ecosystem is actually extremely bad with errors, using error strings almost everywhere, and adding very little context at that.
If your team is spending time on good error messages, then it doesn't matter as much what language-level error support there is. If this time isn't being spent, I think you'll have a much nicer debugging experience in a system which at least collects stack traces.
Even for your example, let's say the stack trace is:
NoSuchHostException in stdlib.LookupHost() line 341
called from stdlib.HttpGet() line 6131
called from applicationLogic.DoSomething() line 123
called from main.main() line 3
In Go, with similar effort spent on error reporting along the way (if err != nil { return nil, err } vs automatic exception propagation) you'd find Failed to do something: "no such host"
I think there's a much better chance to figure out what happened in the first case.Totally agree.
For a package like SciPipe, which we have managed to develop with zero dependencies, to maximize future reproducibility of scientific pipelines, it would very much hurt to bring in the first external (Go-) dependency, while at the same time, we really really would be much helped by stack traces.
If all errors had stacktraces, everything would be much slower.
I've been doing something similar for a while, using `errors.WithStack` from https://github.com/pkg/errors
The error can then be logged with https://github.com/rs/zerolog like this `log.Error().Stack().Err(err).Msg("")`
For human readable output (instead of the standard JSON) use a console writer, see https://github.com/mozey/logutil