As someone else said, alert on symptoms (user-facing errors), not causes, because you're never going to be able to enumerate all the potential causes ahead of time and dealing with the alert onslaught from trying is unworkable.
Sorry, I'm merely a lowly sysadmin and not a godly Google SRE, but environment monitoring is a solved problem, humans need not apply.
When you have enough machines, 10 of them are always running hot.
What might not make sense, resource-wise, to smaller companies does start to make sense at their scale.
And worse, you have no training data, so ML model you train simply learns whatever badness is allowed to persist in production as normal.
Few of them are worth paging someone.
For me, when a machine reports as abnormal it will decommission itself and file a ticket to be looked at, we have enough extra capacity that it’s ok to do this for 5% of compute before we have issues, and we have an alert if it gets close to that.
If you’re doing anything other than that then either your machines are abnormal and functioning, which is scary because they’re now in an unknown state- or you’re able to just throw out anything that seems weird, which might also hide issues with the environment.
Sure, but isn't there rack-level toplogy information that should warn when all of the elements in a rack are complaining? That should statistically be a pretty rare occurrence, except over short periods, say, < 5-10 hours.
Don't really like how this is framed - plenty of Google SREs are born and bred sysadmins.
That said; the google approach _does_ appear to quite literally throw away institutional knowledge from high throughput/high uptime and automation focused systems engineering.
I don’t discount that google SREs are highly intelligent and often know systems internals to the same or a greater extent than highly performing sysadmins. But there does appear to be desire to discredit anything seen as “traditional” and I think that can be harmful, like in this case.
Sometimes they know better, but there is a failing to ask the question: “what does it cost to only act on symptoms not causes”
I imagine hardware breaking is the cost -- but as long as an engineer's time is worth more than the hardware, that tradeoff will be made.
Hardware metrics on the other hand should IMO be monitored and alerts should be in place for when things go wrong, as the consequences of it could potentially be way more serious.
(Disclosure: I work for Google)
As they said, it was within the error budget, which means that they could instead focus on delivering new features. When the errors go out of budget, they stop delivering features and start delivering fixes to bring it back in budget.
It's the kind of attitude you can have at that scale.
Spotting it by digging down from service-level monitoring instead isn't terrible, but it isn't something to be proud of either.
(GFEs are a reverse proxy that terminates TCP for Google. https://landing.google.com/sre/sre-book/chapters/production-... )
I agree, this does seem like an incredibly circuitous and inefficient route to diagnosing the issue, as opposed to the hardware team simply receiving an alert that a server was continuously in an overheated, and then fixing it.
It's entirely possible a server overheating ticket was in the queue so to speak, and the investigation ends up re-prioritizing an alarm that recovered before anyone was on site to fix it.
I wonder what the initial errors being logged were. Based on the post >These errors indicated CPU throttling
It sounds like the first tool is "show error counts".
It might have been "production is slower than normal, please investigate"
Your suggestion sounds like a good way to flag future issues "1 in 50 racks for X is running a temp", however I'm curious what the false positive rate would be. What if it isn't causing problems? Maybe you've spotted smoke.
A remote SRE shouldn't need to monitor the hardware health, an on-site person should have caught this sooner.
I wonder how the person responsible for fixing/replacing the wheels felt about the follow-up
Disclaimer: Google employee. No knowledge of this event beyond the blog post.
But like, it's not impossible to measure temp on devices. It's entirely possible to design a temp monitor that tracks the physical grouping of equipment: datacenter -> aisle -> rack -> server
The real question is one of prioritization: if a rack's temp has raised but there's no customer error, is it an immediate problem?
Here, the problem was on a machine running a google application, so they noticed. But this is a post on the google cloud blog. This just makes me think that Google isn’t monitoring the health of the hardware they provide to customers in the cloud. It is a change you have to make when you change the layer at which you are providing services to customers. If I’m using the google maps web site, I don’t care if they are monitoring cpu temperature if layers above insulate me from impact. If I’m spinning up a virtual machine, I will be directly impacted.
That post is more a statement on how errors which can be handled at the app layer can have catastrophic effects on lower level components. You cannot assume end customers are running thing at the scale or with the fault tolerance of google, or Netflix.
Its completely possible to have different alerting setups for different binaries, and consider CPU throttling to not be a page-worthy issue for GFEs, but to be page-worthy for GCP jobs.
> You cannot assume end customers are running thing at the scale or with the fault tolerance of google, or Netflix.
Correct, but there's a gradient between 'we have 10 copies of the service in 10 different countries and use Akamai GTM in case of outage' and Dave's one-off-VM. One-off VMs are fine if you know what you're getting into, and I use that setup for my personal, lowstakes & zero revenue website. But if you are a paying cloud customer, it makes sense to pay attention to availability zones regardless of scale.
And sure, there might be a market somewhere for a more durable VM setup. At a past non-profit job we provided customers the illusion of a single HA VM using Ganeti (http://www.ganeti.org/). But it's not clear to me that the segment is viable -- customers at the low and top end don't need the HA.
A server CPU that's thermal throttling is about one step less serious than shutting off. While it's not urgent to deal with dead machines, a dying machine should be in the top priority bracket you see on a day-to-day basis.
Still, stuff like this is really hard. It's easy to think of what should have been done in hindsight.