Google OAuth Login Issues
status.cloud.google.com
status.cloud.google.com
- generates a per-incident Google Doc
- opens a per-incident Google Hangout
- pages appropriate people via PagerDuty
During this incident (which affected our customers who use Google to log in), people had difficulty accessing all three (our PagerDuty uses Google Oauth).
This was a good reminder to have alternatives for all crucial incident management services. (We have backup plans in case Slack or our automation scripts are down, but didn’t have backups for Google...)
Lots of people having the same issue: https://support.google.com/chrome/thread/4399961?hl=en
EDIT: minutes instead of hours
A small org can do about the same, and they're not paying what Google pays for infra and staff. Does Google have it harder at their scale? Of course, but they also argue they're better than most at it, at scale.
[1] https://downdetector.com/status/gmail/archive
[2] https://www.theguardian.com/technology/2019/mar/13/googles-g... (Google's Gmail and Drive suffer global outages; Users in Australia, the US, Europe and Asia report problems with various applications for several hours)
EDIT: Disclaimer: FastMail customer. Whatever downtime they've had (and I'm sure they've had some), I don't notice.
Also worth highlighting that for the "email user experience" pretty much any system is better then Gmail these days.
It doesn't even really make sense to compare. And since we're dealing with cloud services where the large players have billions of users increasingly providing these services to everybody from small shops (I work for a startup and we use gSuite) to enterprises (many large companies host all their email on one of 2-3 major email providers).
Nothing I said above has anything to do with user experience as you describe it- that's a non sequitur.
Simple systems are the most reliable. Luck has nothing to do with it. If you create complex systems, you have only yourself to blame when they require more upkeep to meet the same requirements simple systems can meet.
Do I care Gmail can serve a billion people? Not in the slightest. I care about my org’s availability and downtime alone, which is comparatively easy to deliver on when scoped to a local domain.
Off the shelf computers have so many layers of complexity in them that are just not visible to you, sometimes even if you sign NDA blood pacts to ask.
But since they work 99.9% of the time for many new components after the first month or so, and you can just treat them as opaque building blocks to swap on failure, you don't have to care about the complexity, because it behaves as expected and can be replaced if it ceases.
Then the more complicated systems you build on top of this, the less reliable they can be, because they're all lower-bounded by the 99.9 of each computer in the system's part, and each thing you build out of all of them can be no more than, say, 95%.
So then you build a bunch of the 95 things and cleverly arrange it so you can distribute among them to not care if some go away, and get a reliability that's a function of the thing you're using to do the distributing and those 95s, let's say 99.9.
Your org is probably on a line between these three - just be sure to watch and be wary that at some point, you will likely see things going more toward 95% or less from human error or other factors, and be aware that you'll be ticking toward the cost of outsourcing your stack to a provider being less than the cost of running it yourself.
(Work for elgooG, not on anything related to any of this, opinions my own.)
We used to have these arguments with people every day with respect to email. "Simple" is more reliable until it isn't. The company running email out of a closet or small server room in an office space has a great, reliable experience until the toilet clogs upstairs, or you need to add 50 people and scaling up is too expensive.
For cloud services, my experience has been mixed. I’ve been alarmed at the frequency of issues, but they do fix them quickly and the same issues don’t happen again. As an engineer, I can empathize with this and understand. But explaining it to my PM/execs is definitely a challenge.
It only seems like your organization can do better because you're nowhere near the same requirements. Whether this is a valid trade-off by taking on more risk and operational overhead is a judgement call but it rarely works out unless it's a business-critical service that you must control and can afford the resources to do so.