Read Every Single Error
pulumi.com
pulumi.com
An error is either:
1. A real bug someone ran in to.
2. Not a bug and just a spurious "error", which should be a log message, some statistic somewhere, or just disabled outright. Errors should only be errors – an uncaught exception, explicit error log, stuff like that. Everything else should be something else.
What I've found – and this certainly won't hold true for all cases, but it does seem to be the case often – is that half or more all in the second category: it's just stuff that shouldn't be logged as an error. This is usually trivial to fix.
This still leaves you with 5,000 errors/day, but typically ~10 bugs will cause ~4,950 of them. This takes more time to fix, depending on the bug, but usually also takes less time than "5,000 errors every day!" makes it sound.
The last ~50 errors are usually the hardest, especially if it's frontend JavaScript stuff. I "read every error", but I have a few that pop up regularly that I just can't reproduce and don't know how to fix. I think it may be extensions people use or something, but idk. Since this post talks about APIs I guess that's not really applicable to them.
It can be even worse than that. At one point I worked on an app that generated north of 500 lines in the PHP error logs per minute per server. In the course of implementing a feature I cleaned up a few of the low hanging warnings near the code I was modifying. This resulted in about a 3 percent reduction in the log file velocity.
When we deployed to prod, everyone panicked, because log file velocity staying about the same was one of the primary things they alerted on, as a proxy signal for server health. That was the one and only time in my career I was reprimanded for cleaning up warnings.
I don't want to bash on PHP like it's 2002, but ... that's probably one of the most PHP things I ever heard.
I worked somewhere where people were against raising error_reporting because someone was in the habit of writing "flase" instead of "false" – which worked because PHP – and it wasn't deemed worth fixing :-/
Old PHP days were wild.
But what I’ve found is a whole lot of errors are things timing out which are kind of an issue, but not something you can really investigate or solve.
> a whole lot of errors are things timing out which are kind of an issue, but not something you can really investigate or solve.
Depends what is timing out; if it's SQL queries then you should probably do something somewhere, if it's HTTP connections then you probably shouldn't log them as an error but as info/debug, keep some aggregate statistic somewhere, or just not do anything with it at all (depending on your scale and what you need).
And then you have nodejs warnings
If I may broaden this, I think the same principle applies for:
- poor performance
- tech debt/otherwise hard to manage extant functionality
- gaps in documentation
And probably a lot more. If it’s one or a few narrowly identifiable problems, it can be easy to escalate and rally and solve. If it feels pervasive, even trying to address one thing feels insurmountable. And speaking to this from a perspective of being highly motivated to tackle problems like those I mention, and having some fairly significant successes in that regard over recent months, it still feels daunting. But it feels a lot more reasonably possible and more importantly demonstrably valuable the more I chip away at it. But it has taken actual years to get even that far.
At that scale, you get some errors even purely from bits flipping randomly in memory due to cosmic rays. (The latter is speculation. I don't remember how many of those we got, or whether we used enough error correcting memory. It's been a while.)
Pointing out that errors should be counted is tangential to the argument at hand.
We did, however, aggregate and group similar 500s and those did get looked at, but no way could we have looked at all errors.
The other thing, is that with resilient infrastructure, who cares about an occasional 500. Back off and retry. No harm done.
OTOH, the correct number of connection reset by peer errors between your web servers and your database servers is likely zero, and there may be something to do.
I will however guarantee that the author does not mean that they actually look at every peer reset individually.
Broken database connections - yes, a code red and they must be examined urgently.
The whole point is to classify your errors, and especially to know when a new class of error had been seen.
Even at a fairly small scale you will see "errors" of this kind regardless – you don't need to be a huge CDN for that.
But what the author is talking about – in my reading anyway – is programming errors: exceptions, out of bounds access, explicit report_error() calls, things like that which will result in 5xx errors for the client. You've got a lot more control over that. Sometimes random failures beyond your control will cause a 5xx error, and IMHO that's a bug and this usually should be a 4xx or something else (e.g. if the client sends wrong or junk data you shouldn't respond with a 5xx).
I suppose that's the corollary to this article: not all "errors" are the same, and you should distinguish between different types of errors.
The advantage of software engineering over medical treatment is that for the SRE, solving one case also solves all cases of that type, whereas the doctor has to administer the same treatment to each individual case. But don’t let that scaling trick you into tolerating a triage system that results in tasks like “re-architect this service”.
The error budget does not mean that you don't read the errors. It is just a way to adjust the amount of time you spend fixing bugs versus pushing new features.
Also, as mentioned...
> Over the next six months, the team iterates, listens to customers, and delivers a ton of value. As a result, traffic levels grow to a rate of 10,000,000 requests per year.
The people who advocate looking at errors in aggregate may have much, much busier services.
10MM requests per year is 20 requests/minute. Most people will probably be working on services at this scale, or less. So the advice is good, but not global, and the qualification is important.
We have an error budget at work and we still read essentially every error. There is a long tail of errors that might show up once or twice a week or month and I'll admit that I might have missed those, but within any practical sense, my team is still aware of any error that is contributing against our error budget. We resolve errors too, even while under budgetm because if we didn't then we would go over our error budget since new errors pop up all the time, so you need to resolve existing errors because new ones are introduced constantly.
Pulumi seems to think that companies with Error Budgets just wait until they are out of compliance to fix errors and then must only resolve enough to get back under budget.
If you've got 100,000 concurrent users then this isn't going to work.
1,000,000 actions per minute means that all of the actions need a 99.9999% success rate to generate 1 error per minute.
Lots of 9's of reliability is hard.
Also. At that kind of scale you start observing lots of non-deterministic errors such as rare race conditions. As you drill into these rare race conditions you realise that they're affecting 0.001% of customers and you'd have to re-architect your system to eliminate them. You hit a point where it's not worth fixing those errors, so you build an error aggregating system to help ignore them.
You don't need to read every single connection reset by peer or similar network errors..
This is counter-intuitive, but seems to be what the article is suggesting.
The solution was to carefully audit all log messages, reducing the number of messages, making sure they were at the right level, and replacing log messages with trace attributes. We also implemented trace sampling to avoid paying for redundant success messages. That reduced the bill by 80%, saving tens of thousands of dollars a month.
Runtime errors are even worse, for this we do have a CI with a 90% test coverage, or better a model checker like cbmc with 100% coverage. Every single error counts. Esp. on hard Realtime, where we reboot on a single error
But there’s no need to send every single error to Slack individually, just keep an eye on counts by error message.
This is what error aggregators (e.g. Rollbar) do. Now you've got 10-1000x leverage on your time because you're reading unique errors instead of every occurrence, and you've got more data to debug with because you've got data about the aggregates ("when did it start? what request params are common? what servers is it on?") in addition to data about the individual occurrences. You can still commit to solving them all ("Rollbar Zero"). It's just a whole lot easier.
Their primary point of "read every error" is one that almost no one would truly debate (within reason). Most companies do try to read every error as is reasonable. Being aware of an error is different than immediately fixing it (since that can usually be impractical). Where companies tend to differ is in their ability to triage and correct errors. That is the secret sauce that people want to hear.
One of their secondaries points is mocking companies that use error system aggregators. Why is this so bad? They claim it prevents you from reading every error, but I'd beg to differ.
Let's say your system gets 5,000 errors today. It is highly likely that many of those errors are replicate instances of the same error. Out of 5,000 errors you probably have 5 errors that are 800-1,000 instances each, which consumes 4,800 or so of the error instances. Then you probably have a long tail of errors that only show up once or twice a day. So why is an error aggregator bad? Is there any point in reading the same error 1,000 times? Wouldn't it make more sense to read the error once, have the system show you that it shows up 1,000 which is the highest in your system. Then you can spend less time reading more duplicate errors and spend that time fixing them. You can look at a graph and see that the error increases over time, which tells you to prioritize resolution of this error.
If you have 2 hours a day to fix bugs, would you rather fix the bug that causes 2 errors per day or the one that fixes 1,000?
They seem proud of their error burndown chart that shows errors decreasing over time, but this is likely not the result of "reading every error" as much as actually putting time and attention into error resolution. My company for example put a team together at the end of last year who is a cross-functional team who's sole job is fixing prominent errors. We use an error aggregator and our error burndown chart looks the same as Pulumi's. Why? Because we are dedicating resources and attention to fixing it.
Where most companies go wrong is not that they don't read errors. It is that they are so busy pumping out features that they don't dedicate time and resources to fixing things or they don't provide enough time to build less error-prone code before it hits production in the first place.
Lastly, the "haute culture" of error budgets are funny to me. Yes I get that this is the SRE mantra. But what Pulumi describes here is basically an error budget by a different name. You have just set a lower threshold (ie 'budget') for acceptable errors. Just because a team has an error budget doesn't mean that they don't read the errors until they go over. They are reading and fixing errors to stay within the budget.