What the Fastly outage can teach us about writing error messages
onlineornot.com
onlineornot.com
The first problem is that different viewers need different context, which is especially true for service providers like CDNs. Telling an end user "the page is taking too long" makes sense. But what if its the CDN customer (developer), theyre going to need request IDs and other diagnostics just like Fastly did display. Playing the "we need the RID from the headers you didnt know about to capture" game is a losing proposition. The post and it'd "good" examples suggest a way to contact support, but there is ~0% chance support can help without additional information that isnt included.
Expanding on that the post is making really dangerous assumptions and carrying those through to the semantics of the text. Exposing authnz failures as Access Denied vs Not Found is problematic in its own right. Going further to tell the user "You need to login first" is asking for pain in the other 10% of cases where it's a credential or authz problem. Again, different users need different context. And there's no way to know what context to provide apriori.
Lastly it's somewhere between incredibly hard to impossible to distinguish the cause ("why") on a per request basis. Especially for service providers like CDNs. The CDN has a cache miss, but can't retrieve the content. Is it a 504 Gateway Timeout because the origin is not responding, 502 Bad Gateway because the origin TLS is broken, 502 because the local clock is off, or a 503 Service Unavailable because of an internal service timeout? Even if you can distinguish what does the CDN represent to the end user; a general fault in the CDN? A specific failure of the origin? Or a simplistic "unavailable." Again, different viewers need different context and semantics. Which you're not going to do at thousands or millions or tps.
there's no reason that a good error message AND pertinent support information are mutually exclusive. show them both.
show the errors that make sense for that action and that error. no one is saying that there should be a small list of "approved" errors that everyone would see. show what makes sense in that situation, and don't show something that doesn't help.
you don't need to have prophetic insight into the actual causes. just show the flipping errors. don't hide things behind friendly messages; show the friendly message and the error. it's insulting when an error message is hidden from me in favor of something like "oopsies! we broke something! so sorry!!" insulting and disrespectful.
there is no need for a single solution. do what's right for you, and most of all, start thinking about ways that things can work, instead of going straight to the reasons that they might not. don't talk yourself out of a good decision because it might not work for every last situation.
> there's no reason that a good error message AND pertinent support information are mutually exclusive. show them both
As the OP stated, there is quite literally nothing support can do for system-wide outages like this, and when you show support information on these pages, the regular users that see them end up asking for support. There are a non-zero amount of posts monthly on the Cloudflare community forum from people asking about error messages they see for sites they don’t own.
I also agree with GP's point that extremely vague error messages are aggravating. Maybe I did something stupid / unexpected, and with a usable error message I'll be able to fix it on my own. Like sending the wrong file format, or whatever.
I absolutely hate when a system will just say "there's an error". And having cute animals or phrases doesn't change the fact that the error is useless. This is one of the reasons why I absolutely hate working with Windows. "An error occurred. Call your sysadmin". Right. I'm the sysadmin. What now?
This also trains users to never say what's wrong when they call for support. "Yeah, there's an issue with [thing]. What issue? It doesn't work. Oh, of course. There's only exactly one failure mode, so I'll get right on it."
I work in operations and get those fairly often even from people who should know better (software engineers and such) and I find it frustrating to no end that I always have to pull information from them even when there are clear error codes. So no, hiding this kind of info serves absolutely no purpose.
Event Viewer. (You already know, but for the benefit of whomever else.) Hiding things here stinks if you don't know there's something to see there.
I rarely use windows, so I don't know the conventions (if there are any), but I remember one case of someone attempting a remote desktop connection that had an error along the lines of what I said: "Cannot connect". I've never found where the logs for the RDP client were. For servers, we dump the logs in Elasticsearch, so a naïve query will at least set me on the right track.
I'm also not a particularly patient person, so the unbearable slowness of the event viewer's search function and the fact that I never know if I'm looking for Thing, Microsoft Thing, Windows Thing, or ThngSvc, I usually give up in frustration.
And, again, had I had an event id or something, the Find in the event viewer might have been more helpful.
No offence but as an end user I really despise this attitude. Sure there might be automatic monitoring at a place like fastly, but even they can use help in tracking down the problem. Also, if the error is distinct people can put it into Google and call on the vast power on the Internet to help figure out how to fix it or work around it, especially on smaller services where a fix may or may not be forthcoming anytime soon. A good error message leads to a Stackexchange page leads to a solution. A vague error leads to a support call with a bewildered frontline tech and a lot of work for some sorry engineer who has to dig through log files.
Also note that SE went down yesterday
Obviously this does require your developers not to be complete idiots by putting their passwords in the error messages or something, but this is an extremely easy bar to hurdle.
At the very least you can put an error code up. Just make sure it is a reasonably long string so you don't have collisions. If you are big enough people will work out what the codes mean and what they can do. For example, I know 0x80D02017 on Windows means Microsoft has broken their IPv6 service endpoint for the Windows Store again, and you can temporary disable IPv6 support to work around it. Even though the error message is the monumentally unhelpful "Unknown", the Internet can come to the rescue. Of course if the error message had been something like "Windows Store update connection to address 2603:1061::9f4d failed: Service responded with protocol version 4, this client only supports version 5" it would have been even better.
Would it disclose data about how Microsoft internals? Maybe a little, but most of that was observable anyway and maybe if you had error messages like this maybe it wouldn't take them 8 goddamn months to find and fix the problem?
Fair point, I mainly picked the first contrived example that came into mind, rather than thinking long and hard about what the "correct" error message should be.
Will update the article to clarify that.
You can choose plenty of time periods to satisfy a 0% error rate for most services - somewhere from milliseconds to days or even years.
But the faulty node should not be called in a well designed system.
Error rates (as alert thresholds and end user reporting) are a service thing, not a node thing.
Max has many good points about error messages in general, but all of them require access to out of band information.
We in the Varnish Cache Project do not have access to that information, we dont know who runs the varnish instance or what kind of information they serve to what kind of clients.
This is why the default '503 message only exposes the "XID" nonce: That allows the administrator of this cache instance to find all the details in the log files.
Varnish Cache users who want to present something else can do that from VCL, and I'm pretty sure Fastly normally does.
But when all else fails, and here it must have, Varnish Cache errs on the side of caution.
http://varnish-cache.org/docs/trunk/phk/503aroundtheworld.ht...
https://twitter.com/cherrikissu/status/972524442600558594?s=...
Give me “Unspecified error” any day over that.
"Gearing up the dildonator"
"Implicating the fairies"
"Hogtying George Bush"
Dude - just give me a spinner or a progress bar, and if something errors during the load out give me some sort of stack trace or error ID I can use to help
1. "Nice" is an adjective with nearly zero meaning. 2. Either you have access to my camera outside of calls and are analysing my appearance (WTF) or you're making stuff up. 3. You're making chat software. Stop trying to butter me up.
No, really, you should find a fire extinguisher.
> Without good products, life would be a mistake.
> Make products that matter.
> Product excellence isn't a point in time. It is a state of mind, it is a way of life.
> So many feature requests, so little time.
Maybe they're being silly? I hope so. But I've also met product managers and designers who seem to think like that. It makes me internally sigh every time I have to use their product.
(Actually I only now learn that they're related in series - always thought they were competitors.)
Are people really that dumb? I'm afraid to use apps made by those people.
Maybe if we want to be a bit funnier we will write "changing the break pads" in the loading screen, but in case of error we have to write something clearer, like "error accessing to the game server" or "error reading cars data files", to being able to solve the problem.
It is no use to the user since they can’t do anything and is actually dangerous to give out.
It's just something about being thanked in advance for an action I did not intend to do that irks me. It implies to me that the reward for doing it has been given to me without my consent, and now I'm obligated to follow up on my part to prove I deserve it. It ultimately makes me less likely to file a bug report, stemming from this unease.
It's a really minor thing, and definitely not worth getting angry over, but for some reason I always remember it whenever a discussion about out-of-touch error messages happens.
I'll bet whoever put that in there feels extra silly about misspelling "meditation", too.
Software is a highly personal and creative thing, it should have a personality. The everything gray enterprise spaces have gone too far already. Also, I really miss Linux yelling on panics.
Now, the same joke on a highly visible position or on the beginning of the message is harmful and will impede people from solving the problem.
Not it isn't and it shouldn't. The process of creating the software is; huge distinction.
The product of your efforts should not have a personality or feel personal, it should just work as intended. Software is hard as it is and we don't need to make it more whimsical. A little experience can teach us, that whether we like or not, it will ultimately exhibit its own whims anyway.
Genuine question: why it can't have both? I know it's hard to convey tone on the web but I'm asking the question because I'm genuinely interested in knowing what you think.
I personally think you can create software the right way, ship something that works and still incorporate some personality and make it less boring. I don't see why the two can't live happily together.
Imagine if construction/aviation/etc engineers wanted their building/bridge/airplane to be whimsical and have its own personality. Are you scared yet?
Imagine if your out-of-band(not in the initial requirements/spec) and whimsical software contribution was responsible for a bug that brought down an airplane, or killed a patient. How whimsical would you be then? Well at least you wouldn't feel bored at your day job right? Anyway, I think you get my point.
Are you aware that those things go through a design phase with the explicit objective of giving them a personality, right?
Mechanical engineers have an habit of breaking that personality due to their profession constraints, so most airplanes lose the original ones, but bridges usually are built just as intended.
Anyway, it's not like you can avoid giving your software a personality. You can't. What you can decide is if it will behave like a dull humorless thing, a holier than you all knowing braggart, or something people like having around. And yes, some software should have those two first options too, it depends on their application.
Not really. I see quite the opposite; during the design phase the team sets the rules in order to guarantee consistency and cohesiveness and avoid any deviation from what has been agreed upon. This rules out personal or whimsical contributions, because by definition it would ruin the process.
> Anyway, it's not like you can avoid giving your software a personality. You can't.
I alluded to that, if you re-read my comment, but a software having its unintended whims, compared to intentionally trying to give it some "personality" is not the same thing at all.
The bridge architects you are talking about behave very differently from the ones I've met.
But do they build bridges by themselves? Are bridges built by a one man show?
And they don't have rules? And how do they get anything done?
About rules, no, nobody pass rules down the stream. People communicate full designs of some issue (that is not the full design of the thing, designs are "sectorial" where people add their concerns into the overall thing). When it's done right, the design goes to and from those sectors changing the entire time. When it's done badly, somebody finishes a "general" design and sends it downstream for people to fill the other parts. A team does not work on the same issue, that would be chaos.
I think the majority know what 404 is, and possibly 403, but I agree about the more obscure ones.
That said, I don't think it's a bad idea to rely on the "default exception handling behaviour" that the majority of users, even non-computer-literate ones, will have: they'll retry a few times, see that it doesn't work, and go elsewhere for a while.
No more www, no more protocol in the address bar and apple is selling iMac colors in it's commercials...
I spent a good 1 hour to explain the difference to a tech illiterate on why typing "mywebsite.com" was different from typing "mywebsite com" and picking the first result on Google.
I'm not sure he understood, but I really have to admit that this trend of dumbing down things is only make them worse, in a way.
One of my friends has recently learnt that lesson when she went to buy flights on Ryanair; just typing that into Google and clicking the first link. £30 service charge.
Do you think we should use hex or binary instead of their ASCII or UTF-8 equivalents? After-all, ASCII must be dumbing down as it makes it easier for non specialists to read.
Not having people learning new things and learning to approach things with some logic results in them never having to do this in their lives. In the end you get a dumbed down public voting on their emotions and against their interests.
When fastly was broken and telling me that london was broken (lon3356 or something), that told me I could Reroute via Cleveland and have a chance of it working. It also made me comfortable it was a CDN error rather than a site error.
That’s far better than “oops something went wrong, we’re trying to fix it”
Sadly they've been joined by Google. 'Something went wrong' could have come from the Sirius Cybernetics Corporation.
Thanks god for coffee, and nice views out of the window
Isn't that too early for pay mortem? I didn't experience it myself, but I think it happened at most few hours ago.
Posted about 17 hours after the incident.
In short, a valid customer configuration change triggered a bug. One thing I don't see in this writeup is a commitment to ensure that customer configurations cannot break the whole system. Cloudflare does seem to make this promise with their zero trust architecture,
https://www.cloudflare.com/learning/security/glossary/what-i...
Fastly's downtime seems to be caused by an automatically generated config that got deployed in production as a result of change requested by a legitimate customer.
This happened to CF in its early days and I really doubt that ZT had anything to do with the fact that they do not have this kind of problem anymore. It's probably some sanity checks before they deploy updated lua scripts to their fleet of nginx's if anything.
The generic topic your looking for is probably something like "customer isolation" ("service isolation" might also be relevant, but is used also in the context of "tenant isolation" which isn't really what you want). See this thread: https://news.ycombinator.com/item?id=25237836 for some talk about how AWS does "cellularization" which is a form of workload/service isolation/partitioning.
In general I don't think there's much discussion of this issue on the wider web.
With the main exception of user submitted data validations where it is up to the user to submit correct data. For the average web service, there the error is a stack trace and there is nothing the user can do so a blank screen with a "something went wrong, we have logged this erorr" is the only real option.
A number of DNS failures I’ve worked around with the hosts file.
Or if the host name has .eu.domain or something else indicating a geographic location changing that can get around a localized failure sometimes.
It’s more than possible to embed a small vector image to add some humanity to an error page without breaking the bank, bandwidth-wise
-__LINE__
which, if you know C++, means your error code is just the negation of the current line number. That's super convenient when writing. It's really annoying when someone actually sees such an error code and emails you about it -- because they obviously can't do anything with it in software. The obvious problem is that the semantics of an error code depend on the software revision they built with, but you also need to figure out which source file has an error return at that line that the user could have reached at their context.The new error codes are almost as ergonomic to write (involving a little splash of code generation to make it so, not ideal but worth it). However, they can actually be handled in software and there's a functional perror-equivalent.
Edit: trying to reverse engineer the “magic”, architecture, and the failure from a varnish error message is folly, and misleads others. How do I know the comment is patently false? I ran those teams at Fastly for 3 years.
Also helps others pick up if you do go out of business or end up acquired, and before you do it even leads to a more level playing field.
However I guess a level playing field wouldn't have given you a NYSE listing... so I guess you're just being selfish?
Looks like it, except that Varnish has it spelled “Guru Meditation”, not “Guru Mediation”. Anyone know why that would be?
https://github.com/varnishcache/varnish-cache/search?q=medit...
Edit: As to why Fastly wants to expose the relatively stock error message instead of custom text, I dont know. My guess would be Faslty (or their customers) have tooling built around parsing the http response and they're preserving compatibility.
It was to identify fastly's Varnish vs customer's Varnish.
my fault. I would sometimes monitor for "guru mediations" popping up to tell if we were throwing errors without it being caught by other systems. Among other reasons.
Error messages can't possibly explain the problem, because the product can't know what the dumb user did. So don't bother really trying.
This from the perspective of someone doing customer support.
Because when you develop things you are the end consumer.
Technical errors are way better than just "I'm sorry we couldn't process that right now."
But anyway I was really using "backtrace" as a synonym for "technical details that users don't understand". Not very clear, sorry!
[0] https://owasp.org/www-project-top-ten/2017/A6_2017-Security_...
Security by obfuscation is generally not a good option if you can avoid it.
I strongly urge everyone to hide their stack traces in production. This will reduce your application's attack surface.
Of course, but obscurity increases security
You are the guru. And you are meditating on the problem.
Better to give some plain debug info than tell the user "have you tried turning it off and on again?" in my opinion.
For example: Cannot serve website. (What?) Reason: could not connect to database. (Why?)
Most of the time, it is very easy to programmatically assemble such messages. It is much harder to automatically figure out who caused it and when it will be fixed.
I quite liked the windows 0x800*** hex error numbers, though some were more useful than others.
New error codes are only useful if they are generally understood. Maybe instead of using more codes for sub-use-cases, use the permissible error text to express these (HTTP-Status: 503 0x63F0 Data corrupt ?)
I suppose no one does it.