Google Outage in Europe
google.com
google.com
My post is just a speculation, let's wait for the actual technical details doc from Google.
Regarding the technical details doc, Google will never state that outright in individual postmortems. And they will definitely not draw this to the logical conclusion regarding the spiky yearly activity.
Why is it logical? There’s tons of changes being deployed at all times, at all large companies.
Across products, verticals, everything - hundreds of changes at any given point. Some of these changes can introduce hard-to-predict bugs in globally distributed systems. Most of the time, external users don’t even notice before they’re fixed.
Like another commenter said, performance reviews do not coincide with with the year-end at Google and other companies.
When are they? My bet is that they're roughly around the same time period, even if they don't start or close on December 31st.
Of course, I wouldn't go as far as to call this a "logical conclusion" as the evidence I've seen is very slim.
We wrote evals in August. Anybody racing to get things launched before perf did so months ago. The timing described in the top post is just false.
We write the next one in February. We are almost exactly half way between perf cycles. Your hunch that this lines up with the end of the year is false.
100% of all devs make huge mistakes, at least once.
I'm not entirely sure that's always true. For example, i've seen people introduce N+1 issues into a codebase, spend evenings fixing them and refactoring code to fix production issues... just to later introduce those very same types of issues.
Sure, you can learn from mistakes, have post-mortems and so on (provided that your org even does those and that anyone listens and cares about the conclusions from those), but to me it feels like the most foolproof way is to ensure that no-one can make these mistakes again, be it with a checklist (which tend to be ignored, honestly), or better yet, an automated CI step or a new test suite.
In my eyes, it's basically the same as with unit tests - everyone agrees that you need them, but people rarely write enough of them. So if you introduce something to prevent them from not doing what they should, e.g. a quality gate within a CI step which will disallow a merge once the coverage falls below a set margin, suddenly things are a lot better in the long run.
>> a quality gate
Yes, this, also.
Depends on the project, i guess: if you're unlucky enough to be working on a monolith and suddenly a page takes 5'000 SQL queries to load as opposed to 100, because someone thought that initializing data through service/DB calls in a loop is "easier" than writing views in the DB, it might still kill the entire system anyways, depending on the count of users.
And once this data initialization is sufficiently complicated and convoluted for you not to be able to rewrite it and them not wanting to rewrite it, all while "the business" is breathing down on your neck, you might either want to introduce caching (and possibly run into cache invalidation problems down the road), or just freshen up your CV.
I guess i'd also like to expand on the previous suggestion and advise others to consider performance/load testing as well, especially when coupled with APM solutions like Skywalking or even Matomo analytics, both of which can allow you to aggregate the historical page load times, CPM and overall performance of your applications, to figure out what went wrong when.
The only time someone should be fired for causing an outage is if they're negligent or sloppy or mess things up all the time. This is rare. Almost always outages in large systems are the combination of many factors — latent bugs, design flaws, abnormal load, etc, any one or two of which wouldn't take the site down. But when the combine in a perfect storm that nobody foresaw things fall over.
$140 billion dollars. On training.
On the one hand... you know what, I'd love to work in an environment like that. Seriously.
On the other hand... what's the argument you make to the CFO in support of this? Honest question, interested to hear answers.
I don't know how much that costs all-in, including the salaries, instructors, facilities, but might be starting to approach a million.
That's valuing training!
But I'm glad that learning to kill people (military) is not taught that way.
For instance, if Google fails and cant profit, it cant just shoot at their client until they pay. Your organisation can.
Well we had to build an Army to win against fascism in the Second World War or we all would have perished.
And 'perished' means literally dead or subject to fascism, not just going out of business.
You want the Army to be... less efficient? Spend more for less capability?
If your company fires people in situations like this, run away and never look back.
People do absolutely NOT get fired over incidents. Making mistakes is human. An incident will prompt a review of the systems and safeguards in place to prevent such an incident, much like an airline incident investigation -
basically "somebody fat-fingered it" is never the answer, postmortems are always blameless
EDIT: now that I think of it, the opposite thing happens after a major incident - a systemic failure should be identified, people are being hired to fix it :)
Why do employees at big tech names (FAANG et al.) are so often so cautious as to include this as a foreword everywhere? Twitter bios are full of that, for instance.
It is crazy to me; who would expect anything else that our opinions being your own and nothing more? Who would expect that your word (with all due respect) is worth anything with regards to the company's PR?
Is there an actual risk in the US? Have there been trials or anything that push people to add such statements?
I didn't want to emphasize on that on the first comment yet to be honest, I find it pedantic because it's pointless, legally speaking.
Basically the company has specially trained people that speak on behalf of the company, and that message should not be confounded by personal opinions of other employees. For example, on the recent FB outage, there was an employee posting inside information on reddit - media companies just took it at face value and ran around with it reporting as it was what FB itself was saying about the outage.
I'm not aware of any actual risks in the US, but then again I'm not in the US. For me this seems a minor point, and I actually enjoy separating my public persona from the company for which I work, being it Google or a small startup.
To be fair, I think the media would have done that even with a "speaking only for myself" disclaimer.
If no, why are you doing this for the company you're trading skills and time against money?
Though I almost got to be the official spokesperson for British Telecom responding on the alt.2600 news group about the the Met police VMB hack - press office was cool but the internal security was not.
Dont name your company if you intend to speak for yourself.
This isn’t just ‘big tech’ - I work at a relatively small tech company, but I’d never want anything I say about the company to be mistaken as some sort of ‘official statement’ especially if it related to an incident that possibly had a financial impact on external parties, and could conceivably be misused in that context in the future.
I go as far as never writing private emails from my work mail for the same sort of reasons - although that is from a possibly over-abundance of caution.
That's why we constantly get "it's just my opinion" used in reference to type-ii opinions (personal understanding of descriptive fact), when it's only really appropriate to type-i opinions (normative value).
Many conversations would be far clearer if it were abandoned in favour of more precise language, IMPUOTDF.
Rumor mill journalists will mine social media and forum comments and write entire articles about "so and so FAANG employee gives hint at future merger" when some dev comments how much they've enjoyed using some library recently.
I think it's a bit over-applied in some cases. Does it not commit you to the theorem that every process can be made so perfect as to be completely invulnerable to one human being making a mistake? (At least, in the form exemplified by the common tweets to the effect that "your processes are to blame for $incident, not your interns/engineers/etc".)
Even if you required two-person auth for every single thing, two people will make a mistake now and then, and in reality - due to our being social animals - the two probabilities are not truly independent.
I just don't see how this is feasible in reality. A more realistic principle feels like: "people will infrequently make mistakes, and that's of course natural and human and forgiveable, but far fewer incidents should be vulnerable to human error than currently are".
Also, such cases are rarely the "fault" of a single person. Or, the direct/immediate cause is often not the main one.
The people who did the stuff have either gotten their promotions and moved on or didn't get a promotion and have given up.
:-)
Excellent time to slack off a bit, make a few mistakes, then come next eval you can point to a marked improvement over the intervening six-to-nine months!
I am not sure where to look for it
We barely understand the systems we’re writing code for - I don’t see how you could objectively judge the code to manage those systems without fully understanding the systems first.
You could have wave some metrics like number of bugs or test coverage, but I can’t think of any (aforementioned included) that aren’t subject to tons of confounding variables.
(Googler, opinions are my own)
Google's more interested in placating their primadonna engineers than solving customer problems.
I don't know how or if this would impact Google, but I'm sure someone there has at least thought about those dates.
Googler here, that is not when those promotion/perf evaluation windows happen.
Assuming the chance of an outage is actually static. If there is one (highly reliable and trusted) provider with five products one year, and three providers with ten products the next year, the chance of you seeing an outage has gone way up because of surface area.
for those with budgets to hire server administrators (or pay for third-party managed services)? yes.
second category includes almost anyone with a large following like institutional users.
hell the incumbent social media service operators can white-label their existing software and sell this as a service.
Self-hosting is a good option, especially when mixed with multi-cloud offerings.
The big rub is that it's really hard to approach the same level of availability the cloud offerings already give you. Depending on work-load, self-hosting is typically more expensive.
This is why the SRE book talks about availability budget. You can't have 100% uptime, so how much do you want to pay to get close to it?
Quite boring commentary tbh.
(Yes, I know they discontinued it, but I haven't seen the service "over capacity" in some time. I say this as a casual web user of it, not an active account holder or app user).
Facebook had an October outage.
This happens quite regularly with distributed and complex systems.
(glares at AWS)
Everyone is bad at them. Trading desks, international trade offices, space related offices all have a set of clocks on the wall because even really smart people just suck at figuring out time zones. I even use https://everytimezone.com/ in lieu of a set of clocks.
What a needless complication. I wish we could all just switch to UTC and stop daylight savings time.
So from the perspective of a regular person it is very reasonable to conclude that it is their internet provider that is down.
9 times out of 10 if you can’t connect to hn the problem is you, not them. Probably even 99/100
A single machine won't give you the same level of scaling and management features, but also won't have hundreds of distributed moving parts that could break and take your service down as a side-effect.
However, when your small the economics flip. You want to run as little as possible as cheap as possible. And a load balancer pointing at a two asgs in different availability zones is by far the cheapest way to get something production ready on the internet. You can go from your garage to a 100M business and probably never need to graduate from this architecture.
It was very frustrating to see DNS working perfectly fine (since DNS is always the problem) and the connections just timing out.
Here's a €2.4 billion fine for you.
AI to EU:
OK, human.
Next stop AWS, firewall maintenance division?!
what makes you think so, actually?
In my experience it's also very easy to think you've got all your bases covered when actually you haven't. I'm protected against mains power failures by a UPS and a generator - but a UPS switchgear fault can cut my power even without a mains power outage. My server has dual power supplies and network cards - but that won't help me if a clumsy worker sent to replace the server above mine unplugs mine by mistake. And so on.
If I think I'm doing a better job than a billion-dollar corporation with hundreds of thousands of servers, does that mean I am? Or is it more likely I'm fooling myself?
For you average person working in an average office if the power goes out you're not going to be working anyway, so it doesn't matter if your server is offline too.
More importantly, I do know I had an error loading twitter at Thu 11 Nov 04:17:01 GMT 2021, however my websites (and google and hackernews) were working. At 17:53:01 GMT Twitter took over 2 seconds to load it's first http page, far beyond the normal 400ms. Google took 23 seconds to load this morning, BBC News just 0.071.
On the other hand I also need to provide services which can't cope with outages measured in milliseconds. Good luck with complaining to a cloud provider that your traffic vanished for 4 seconds. Those services thus have multiple connections on independent hardware and circuits with no single point of failure
And an excellent corollary to this is that when you have lucky 100% uptime there is no incentive to optimize mean time to recovery.
Sure the raspberry pi in your closet has been running fine for years, 100% uptime, but then a component fails. Do you have a replacement on hand? Are you continuously monitoring it to know it went down? The component failed at 3am, did it page you? Did you hop right out of bed to rush to fix it?
Single systems can have really nice uptime until they don’t. Then you are hoping that the people on hand can repair what’s going on after months or years of never having to do that. Mean time to recovery might be a week while you wait for new hardware or a few hours while you google some error message you’ve never encountered.
People can run their own systems if they want to, but they shouldn’t confuse good luck with rigorous engineering.
The server not failing is not the only outage mode.
Of course I'm not trying to run a massively scalable service coping with millions of customers, because I don't need that.
"It’s still less downtime than your own server"
Claim disproven
I say that in many cases your own server is better. My company runs its own on-prem confluence, it's been taken down for updates on a regular basis at a known maintenence time. That's far better than losing it because a cloud based one was hosted on GCP this morning when we actually need the data
Obviously in many other cases the cloud is better. You wouldn't want to server a million cusotmers across the world from a single server in your own basement. That's not the only model.
My only gripe was with your gross over-simplification that read like "hurr durr, my server hasn't gone down this year, so self-hosting has better uptime than Google". It's such an unnecessary and baseless argument.
Every few months another major outage of another cloud provider hits the headlines, meanwhile millions of small companies have no problem with the uptime of their 'legacy' services.
I was at a farm a few weeks ago, the farmer had a server in a closet. It did break on occasion, when there was a power outage. His desktop and internet broke too, so what would the point in his server working.
If it was hosted on google it wouldn't have been working this morning, despite his computer being fine.
If you build your business processes around accepting failure, it's not a problem. It's far easier to keep at least one out of 3 machines online for 99.999% of the time than to keep a single service running for the same time.
When you have a "simple" (already quite complex) BGP(TCP(HTTP)) tunnel, chances are things just work and it's easy enough to diagnose issues. Between anycast, auto-scaling, and WAFs (among others) you have added so many layers of complexity that the chance of a "random" error somewhere in the stack has dramatically increased and diagnosis is close to impossible.
It used to be simple that web services are either up or down. Now it's not so obvious anymore, and that's definitely not being reflected in the 5 9's SLAs falsely promised by cloud providers. I can say with confidence most selfhosted systems i've been close to have much better uptime than modern cloud services and are much cheaper to service and maintain.
Firstly, I think that the GP comment ever so slightly misapplies the localized awareness available in smaller environments to the cloud. Yes, there are tons more engineers to look at stuff, but those engineers are a) already bogged down keeping up with internal infra, politics and bureaucratic machinery, and b) quite some distance further away from actual errors occurring on the ground because everything's aggregated to the hilt so the stats remain comprehensible at the higher scale.
I also agree that the newer stuff that is less mature has an order of magnitude more intrinsic leaky abstractions than old designs, which I do think were more hygienic and espoused more effective separation of concerns than what is used today.
Also, not to nitpick, but the "classical" canonical definition would probably look closer to BGP(TCP(TLSv1.3(HTTP/1.1))), with the modern equivalent being BGP(UDP(HTTP2)) and the future being BGP(UDP(QUIC)). I do agree that the rapid consolidation of HTTP and TLS, without wide-scale awareness and slow, methodological development of general introspection tooling, does make things net worse in general. I suspect infra will still be using HTTP/1.1 for a long time into the future until this materially changes.
My private website on my RPI is running now 2 Years without a problem and only minimal downtime due to rebooting for the new kernels.
It is amazing how much uptime you can achieve with a 5$ Computer in comparison to a 1730000000000$ (1,73 tera $) Company. Even if you compensate for dynamic content.
In contrast, in the cloud you typically take on the risk of the underlying distributed platform (optimized for managing thousands of VMs, etc) even though you only need a single machine for yourself.
I think it purported Google to be hovering around 15Gbps. That sounded humongously wrong, like I'd expect global traffic <-> Google to at least be a couple terabits, right?
I'd be very curious if there's a way to actually ballpark the number. Or maybe it is in fact possible to implement services that reliably track this sort of thing, and my futile searches earlier just weren't finding them...
Was your home internet available all of the time? How many times did you reboot your modem?
It’s a pity as some of its services are frankly among the best, and I depend on them.
Do you mean about the "Don't be evil" slogan it started with? Even back then, it was a pretty obvious move for any movie villain.
Or do you mean you actually bought into the reliability promises of cloud providers thinking downtime would not exist anymore?
Unrelated: not sure what is going on on HN since a few months but almost all comments I get to write get automatically downvoted, even little things that no one should really care much about.