Google Compute Engine Incident #16015
status.cloud.google.com
status.cloud.google.com
The postmortems published by Google, Amazon, and Azure (as well as postmortems internal to the company I work for) are nearly always due to some type of change (code or configuration) being rolled out. It seems to me that we need some help from computers to make these systems reliable. - something like static type checking in a programming language, but applied to a distributed system. Perhaps your architecture wouldn't "compile" if the network traffic will go the wrong place, or if a rate limit is above the capacity something is expected to handle, or if the change would impact too many servers at once.
This is an interesting idea. I'm always terrified when I have to deploy a minor configuration change into a production system that gets distributed out; if there was an easy way to apply some sort of check against all of it (you know without having to build something explicitly to do this) that would be awesome.
I wonder if the future of something like this might be using containers. Each piece that gets deployed is in its own container networked together and you stand it up in a staging environment, ensure it's talking together correctly through some sort of set of integration tests then push it into the production environment.
Hell even without automated the justice system just doing as you suggested alone would be amazing. I've often wondered about putting laws in GIT or similar.
Would you be willing to trust a program with your sentence? I wouldn't in the slightest.
So that's a lot of ifs and provisions...so I wouldn't expect this to happen ever.
And be self-consistent I mean that it doesn't contradict itself.
Tiny nitpick, it's git, not GIT :)
This seems crazy to me when intelli-sense for many languages can seemingly read my mind but configuration is in many ways still a trial and error process..
Does anyone know of ways to parse many of the common configuration file formats? From linting to code completion to best practice (or worst practice) checks. IF not and somebody wants my money yesterday, theirs your startup.
Abstracting the problem to another layer isn't always the answer.
I'm sure there is lots of research happening in this area within Microsoft and Google. Can't wait to have this sooner.
When you have a really large distributed system that's primarily running on metal, smaller-scale copies of the whole system per developer or even a single companywide staging environment that mirrors production are really hard, and they don't always exist.
Developers work on their components in isolation by mocking out the rest of the system, hopefully there's a thorough code review, and then "integration testing" happens by flipping a feature flag and watching the logs/metrics in production. You might design the feature flagging so it initially only hits test accounts, but some things (like service communication layers, Puppet configs, router configs) don't work that way.
The cost of outages resulting from the lack of a staging environment may well be less than re-architecting production (100% automation, no snowflakes, and probably some kind of IaaS) to allow for disposable dev/test environments which would exhibit the same bugs.
If you can't really test your change, then the other thing you can do is try to rationally analyse its consequences. After all thats why "hopefully there's a throrough code review".
Wouldn't it be nice to automate some of that analysis? I have no idea how to go about such a thing, or whether it is even possible -- but the harder testing gets, the more attractive this option is.
External network is BGP though and sounds like they didn't detect it until user reports hit. They can't predict the problem and there detection isn't working well either.
> Perhaps your architecture wouldn't "compile" if the network traffic will go the wrong place, or if a rate limit is above the capacity something is expected to handle, or if the change would impact too many servers at once.
So in the first instance, I tend to like this sort of idea. However: we are already substantially ahead of the sort of things that you're thinking of.
Full static simulation of a system as complicated as all the components involved here is... well, I can sort of see how it could be done, but it would be a herculean effort; I don't think it would ever be good enough to catch cases like this the first time they happen. There are systems where this sort of thing can be done, but all the ones I can think of are much smaller in scope.
I can't talk about the details, but you can assume that "static analysis" of the form being talked about here is something we've already done, and it's not enough to handle cases this complicated.
https://cloudharmony.com/status-for-google
Prior global outage - April 11: https://status.cloud.google.com/incident/compute/16007
Disclaimer: I am the founder of CloudHarmony Edit: Outages triggered due to ICMP timeouts
It states "1.67 hours" of downtime, which is not the same as increased latency due to sub-optimal routing.
That means that Google is not dogfooding GCE to the degree that I would hope and expect for me to risk my business on using the service. Disappointing, to say the least.
Say what you will about AWS (and I've said many critical things, and will continue to do so, publicly and privately), but when AWS has a major outage it also affects Amazon digital products. They have major skin in the game, while it appears that GCE is a special snowflake service completely separate from the important, money-making services at Google.
GCP is pretty new and there are plenty of big customers running on GCP that you can reference like Spotify and Snapchat. Also it is a rather important and money-making service, on track to potentially eclipse their entire ad business.
"[O]n track to potentially eclipse..." sounds like the worst form of quarterly-report spin.
This is just the size of the market, cloud computing is already a major industry and has just started. The potential upside for a major cloud player dwarfs the entirety of digital advertising.
It might be good if Youtube or something else big-but-tangential was on the line when GCE changes went out. Search, though? It dropped globally for 2 minutes in 2013, and took out 40% of internet traffic in the process. There's a real argument that it's critical infrastructure like little else, since Google and Bing are the way most users have of reaching sites other than the core Facebook-Amazon-Yahoo locations.
More dogfooding wouldn't be a bad change overall, but people counting on Search and Gmail would probably prefer a minimal-downtime solution no matter how it's implemented.
But like Steve Yegge said in his classic rant[1], when you're behind in the market you have to launch your cloud product (which is really just a subset of the "platform" problem Yegge was actually writing about) "...all at once, for real, no cheating".
If you don't have the organizational courage to take big risks, well I guess you deserve your 2.5% market share[2].
0. http://marketingland.com/google-revenues-beat-expectations-w...
1. https://gist.github.com/chitchcock/1281611
2. http://www.datacenterknowledge.com/wp-content/uploads/2016/0...
> resulted in announcing some Google Cloud Platform IP addresses from a single point of presence in the southwestern US.
> not yet configured to handle Cloud Platform traffic and applied an overly-restrictive packet filter.
that caused the real pain. Other traffic that ended up in the Southwest would have routed crazily, but wouldn't get stopped (filter). If it had just been a route thing for GCE (and related services) we would have just seen slowdown, etc. like any other routing mistake.
Disclosure (and Disclaimer): I work on Compute Engine, but I'm not on our networking teams.
Even more simple, is DNS round robin (just putting more than one entry in the reply for a dns request). So in terms of failover, if one of these servers in the reply stops working, generally another one is, and browsers and other devices generally understand that they can try all addresses listed for a given host.
FWIW I'm making it a thing today to read up on DNS as this post made me woefully aware of how little I understand it. So feel free to skip this question and I'll probably answer it myself tonight.
In a perfect world, you use short TTLs (30-60 seconds), and withdraw a record from being advertised as soon as your health checks are failing for it.
Also for it to be that jarring we are talking about a failure at the exact moment you are using the data-heavy app, and not 1 minute before. Every time you go to use the app its likely the dns is re-queried.
So, at the end of the day if you have a mission critical thing running where you need to try very hard and can't tolerate 60 seconds of failure, you would bake multiple paths to a backend in the client configuration at some point. So the client would know if backend A is down it can do client side failover to backend B.
Sometimes this technique is used in conjunction with a highly fault tolerant server stuff, and not as a replacement for it.
I've wondered if it's simply because I don't know it well enough or if we need an easier, more simple solution to replace DNS in internal networks. There are a ton of ways to do service discovery.
IMO, and emphatically not ripping on you: it's this one. DNS is a pretty simple protocol and at "you don't have ten thousand developers" scale managing it is pretty easy. It might take you a while the first time, for sure--but it won't the second, third, or twelfth time. =)
In particular, even setting aside explicit failover cases (which do require some outside monitoring and staging to make useful, i.e. when do you failover?), round-robin DNS actually does a lot to provide redundancy, internal to a single data center or deployment, when coupled with SRV records.
And there is some beauty in it. It's fairly simple and with more servers in more places you get fewer affecting clients when one of the places fails and fewer disrespectful clients as well. It scales very well and cheaply.
Kudos.
In internal post-mortems names are named when necessary for disambiguation, without consequences to the persons named. This results in all participants being candid, and the real source of the problem more easily uncovered. (I wrote two post-mortems for small incidents I triggered.)
I would hope so (everyone makes mistakes some of which may lead to HUGE issues) and not giving them a chance to learn from it is a bit unrealistic. At the same time, however, I have been part of organizations where something has gone wrong and us engineers basically all had to "take the blame" otherwise the single person likely would have been fired. I'm confident Google is not like that but it's always something I think about.
It sucks when somebody starts to point you at their SLA when you ask them about outages. I'm looking at you, AWS.
This level of detail and honesty, this level of specific steps that are being taken to prevent recurrence.
75% of startups try to sweep downtime under the rug and don't even acknowledge problems happened at all -- or only do so privately in a personal support email after hours of interrogation. Another 20% just write "we had some minor issues but everything is fine now."
Be like Google.
http://www.cnet.com/news/how-pakistan-knocked-youtube-offlin...
Let's go with "yes", as the most accurate answer. As soon as I or whoever is oncall has figured out what change was responsible, we can usually revert it quickly and easily. Usually, if I'm oncall and I have reason to even suspect a recent change might be the cause, I'll revert it and see if the problem goes away.
The difficulty becomes more apparent when you realise the sheer number of infrastructure changes being made every hour, some of which will be fixes to other outages, and some of which will be things you can't revert because they are of the form "that location has fallen offline; probably lost networking" or "we are now at peak time and there are more users online". So if your question is "can we just roll the whole world back one day" - no, too much has changed in that time.