Scary Outage Stories from CTOs
thenewstack.io
thenewstack.io
https://github.com/stickfigure/blog/wiki/Beware-cutesy-two-l...
I will never, ever again use a non-.com domain for a business.
One day, out of the blue, the authority started refusing any requests towards the domain. We tried to contacting them, but nothing. It took a week for the domain to come back. If this was our main domain, I think we would be not in business anymore.
The affected users don't care that it works for others, and you can only afford to ignore it if you are OK with not only losing those users forever, but also the negative reviews and word of mouth.
The probably biggest (in terms of company size/affected users) example of this would be Twitter's "something went wrong" on mobile. I'm sure that they have some sort of metrics showing them that everything is fine, but plenty of users on HN complain that they are pretty consistently getting a generic error message when they open a Twitter link until they reload the whole page. (I assume it's separate from the issue mentioned in the article, since it was still ongoing recently.)
Also, because errors are much more noticeable than everything working, a 10% error rate is pretty much perceived as "never works".
A subtopic of that: The affected users don't care why it is broken, and it is broken until it is fixed for them. If you fix the issue, but the fixed version hasn't yet made it to your users' phones (e.g. because it's only fixed in the development version, or your version hasn't made it through some review process), then your app is still broken. (Looking at you, Newpipe and Firefox: Newpipe was broken for weeks, apparently because F-Droid, their official distribution channel, is shipping an outdated version, and when trying to fix that by downloading the newest APK, I ran into a bug in Firefox Mobile that prevents downloads from working - which is fixed in some developer version but not the public one, and has been broken for at least a week).
A simple "rm -rf" turfed a directory full of configuration files. This was immediately (well, 2004 immediately) replicated to several countries. Thus several telephony systems went poof. Obviously the same rm -rf took out the CVS repository, and on a Saturday night. Want to make it more fun - the ENTIRE (worldwide) development team was offsite for a team-building week, in Bonn.
Two of my friends got on a plane, flew to Ottawa, and started having a go at computers to find a recent repository. Hard drives were ripped out of machines, jacked into others, mounted, and prayers made. This wouldn't have had to happen had the tape robots (yup, plural), with backups not both their arms stuck, and the vendor simply not providing 24x7x365 support because the company was too small. On the bright side, everything was up relatively quickly all things considered.
Way better than the time Mike BBQed 20000 chickens accidentally.
Needless to say that was a pretty rough week.