Finding a problem at the bottom of the Google stack
cloud.google.com
cloud.google.com
I suspect this isn't a water cooling problem, but instead a heat pipe system, with a phase change material in (often butane). Heat pipes are used to conduct heat from CPU's to heatsinks in most laptops. They have a liquid in which boils, and then condenses at the other end of the pipe, and then flows back again as liquid. Heat pipes usually look like a thick bar of copper, but are in fact way more thermally conductive than copper.
The inside of the pipe usually has a felt-like material to 'wick' water from the wet end to the dry end, but wicking is quite slow compared to guaranteeing the pipe is perfectly level and using gravity to just let the liquid flow downhill.
I'm 99% sure that's the reason this system doesn't work with a slight slope.
Something as benign as that can take out an entire infra if unchecked.
I've seen racking fail in the past (usually someone loading a full 42U with high density disk systems putting the rack over weight capacity by a notable factor) and it is definitely a disaster situation. One datacenter manager recounted a story of a rack falling through the raised flooring back in the days when that was the standard (surprise - it kept running until noticed by a tech on a walkthrough).
Good story but comes across as back-patting a bit.
It's my understanding that switchable fans are standard on cisco, arista, and mellanox switches (have no idea about juniper) and besides, front-to-back or back-to-front depends on whether you mount in the back or in the front (both are valid) AND even if you go the "wrong way" patch panels are pretty easy to find and are cheap (at the expense of 1U)
It was not uncommon to see half-depth switches in the era prior to 'hot aisles' and I definitely recall them pulling air in directions that you might not like. They were never deep enough to siphon air from the front of the rack.
If you have a switch fully populated rail-to-rail with ethernet jacks, it's not entirely clear to me where you should route air in a hot/cold aisle setup.
Hardware metrics on the other hand should IMO be monitored and alerts should be in place for when things go wrong, as the consequences of it could potentially be way more serious.
(Disclosure: I work for Google)
Still, stuff like this is really hard. It's easy to think of what should have been done in hindsight.
I wonder what the initial errors being logged were. Based on the post >These errors indicated CPU throttling
It sounds like the first tool is "show error counts".
It might have been "production is slower than normal, please investigate"
Your suggestion sounds like a good way to flag future issues "1 in 50 racks for X is running a temp", however I'm curious what the false positive rate would be. What if it isn't causing problems? Maybe you've spotted smoke.
A remote SRE shouldn't need to monitor the hardware health, an on-site person should have caught this sooner.
I wonder how the person responsible for fixing/replacing the wheels felt about the follow-up
Disclaimer: Google employee. No knowledge of this event beyond the blog post.
But like, it's not impossible to measure temp on devices. It's entirely possible to design a temp monitor that tracks the physical grouping of equipment: datacenter -> aisle -> rack -> server
The real question is one of prioritization: if a rack's temp has raised but there's no customer error, is it an immediate problem?
Here, the problem was on a machine running a google application, so they noticed. But this is a post on the google cloud blog. This just makes me think that Google isn’t monitoring the health of the hardware they provide to customers in the cloud. It is a change you have to make when you change the layer at which you are providing services to customers. If I’m using the google maps web site, I don’t care if they are monitoring cpu temperature if layers above insulate me from impact. If I’m spinning up a virtual machine, I will be directly impacted.
That post is more a statement on how errors which can be handled at the app layer can have catastrophic effects on lower level components. You cannot assume end customers are running thing at the scale or with the fault tolerance of google, or Netflix.
Its completely possible to have different alerting setups for different binaries, and consider CPU throttling to not be a page-worthy issue for GFEs, but to be page-worthy for GCP jobs.
> You cannot assume end customers are running thing at the scale or with the fault tolerance of google, or Netflix.
Correct, but there's a gradient between 'we have 10 copies of the service in 10 different countries and use Akamai GTM in case of outage' and Dave's one-off-VM. One-off VMs are fine if you know what you're getting into, and I use that setup for my personal, lowstakes & zero revenue website. But if you are a paying cloud customer, it makes sense to pay attention to availability zones regardless of scale.
And sure, there might be a market somewhere for a more durable VM setup. At a past non-profit job we provided customers the illusion of a single HA VM using Ganeti (http://www.ganeti.org/). But it's not clear to me that the segment is viable -- customers at the low and top end don't need the HA.
A server CPU that's thermal throttling is about one step less serious than shutting off. While it's not urgent to deal with dead machines, a dying machine should be in the top priority bracket you see on a day-to-day basis.
Spotting it by digging down from service-level monitoring instead isn't terrible, but it isn't something to be proud of either.
(GFEs are a reverse proxy that terminates TCP for Google. https://landing.google.com/sre/sre-book/chapters/production-... )
As they said, it was within the error budget, which means that they could instead focus on delivering new features. When the errors go out of budget, they stop delivering features and start delivering fixes to bring it back in budget.
It's the kind of attitude you can have at that scale.
I agree, this does seem like an incredibly circuitous and inefficient route to diagnosing the issue, as opposed to the hardware team simply receiving an alert that a server was continuously in an overheated, and then fixing it.
It's entirely possible a server overheating ticket was in the queue so to speak, and the investigation ends up re-prioritizing an alarm that recovered before anyone was on site to fix it.
As someone else said, alert on symptoms (user-facing errors), not causes, because you're never going to be able to enumerate all the potential causes ahead of time and dealing with the alert onslaught from trying is unworkable.
Sorry, I'm merely a lowly sysadmin and not a godly Google SRE, but environment monitoring is a solved problem, humans need not apply.
When you have enough machines, 10 of them are always running hot.
What might not make sense, resource-wise, to smaller companies does start to make sense at their scale.
And worse, you have no training data, so ML model you train simply learns whatever badness is allowed to persist in production as normal.
Few of them are worth paging someone.
For me, when a machine reports as abnormal it will decommission itself and file a ticket to be looked at, we have enough extra capacity that it’s ok to do this for 5% of compute before we have issues, and we have an alert if it gets close to that.
If you’re doing anything other than that then either your machines are abnormal and functioning, which is scary because they’re now in an unknown state- or you’re able to just throw out anything that seems weird, which might also hide issues with the environment.
Sure, but isn't there rack-level toplogy information that should warn when all of the elements in a rack are complaining? That should statistically be a pretty rare occurrence, except over short periods, say, < 5-10 hours.
Don't really like how this is framed - plenty of Google SREs are born and bred sysadmins.
That said; the google approach _does_ appear to quite literally throw away institutional knowledge from high throughput/high uptime and automation focused systems engineering.
I don’t discount that google SREs are highly intelligent and often know systems internals to the same or a greater extent than highly performing sysadmins. But there does appear to be desire to discredit anything seen as “traditional” and I think that can be harmful, like in this case.
Sometimes they know better, but there is a failing to ask the question: “what does it cost to only act on symptoms not causes”
I imagine hardware breaking is the cost -- but as long as an engineer's time is worth more than the hardware, that tradeoff will be made.
I can't imagine a better way the phrase "bottom of the Google stack" could have been used. That phrase can now be retired.
What does "statelessly cache" mean? Stateless means it has no state, and cache means it saves frequently requested operations. How can it save anything without state?
> Presumably it caches copies of frequently-requested content on nodes closer to the end-user, rather than always fetching that content from an origin node
> It's "stateless" because stateful requests are always sent directly to the origin node
What is a stateful request versus a stateless one? Every request has some sort of state.
If the request is "give me the profile photo of the logged-in user" - the response depends on the stateful question of who is the logged-in user.
If the request is "give me the gmail logo" - the response is identical no matter who's making the request. These can be cached on an edge node without regard to application state.
I get what you mean, but calling a cache "stateless" seems like a contradiction in terms to me.
The blog post mentioned it this way to convey that the cache operation should be relatively simple and generally not return many errors, and that disabling the servers would be low risk (no data loss or severe performance impact).
Whether or not data is in the cache can change, but if the data is there, the result is always the same.
An example of stateless data would be a video on YouTube or a particular version of a file in git.
If the data cannot change but is controlled by AuthZ/AuthN then it is pretty clearly not stateless.
I get what you mean, but calling a cache "stateless" seems like a contradiction in terms to me.
The temperature spike should have been the first alert. The throttling should have been the second alert. The high error count should have been third.
If you're thermal throttling, you have many problems that could give give all sorts of puzzling indications.
For rationale, see:
- https://docs.google.com/document/d/199PqyG3UsyXlwieHaqbGiWVa...
- https://landing.google.com/sre/sre-book/chapters/monitoring-...
It seems a shame that the blog post wasn't able to follow
« They immediately removed ("drained") the machines from serving, thus eliminating the errors that might result in a degraded state for customers »
with something like « and they were shown a notification that there was a low-priority automatic ticket showing a possible hardware problem on those machines, so they subscribed to that ticket and didn't waste any more time ».
I highly doubt that the edge oncaller who was paged debugged this issue down to the hardware. The post even said that they worked with other teams to figure this out. Once the SRE figured out it wasn't their problem, the likely moved onto something else
This is a completely backwards. What you are describing is caused based alerting, which is strongly discouraged by SRE. SREs prefer symptom based alerting (e.i are users seeing errors) because you only get alerted when you know that there is a problem that effects your business. If you only have caused based alerts, you will get false alarms and un-actionable pages all of the time.
Imagine being an SRE who manages the GFEs in this situation. Now imagine getting a page telling you that the tempurature in a rack where your job is running is hot. What are you supposed to do? Is the issue effecting users? If so, how many? Is it anomalous? Now go poke around the system and try to figure out what is wrong. What would be the first thing you check? Probably error rate/ratio, the thing that you actually care about. If that is your workflow, why not just alert on your error rate in the first place and figure out the rest as you go along?
Edit: typos
I follow Google's symptom-based alerting philosophy in general, but will make an exception when there's a chance to catch something getting dangerously close to an unambiguous hard-failure limit (e.g. 90% quota utilization).
Do you really think that the folks in the Google datacenters are not monitoring rack temp? Implemnting Symptom based alerting does not mean you should not be monitoring other system metrics. Their monitoring system probably filed a P1 or P2 ticket for someone to go take a look at it at some point. But should a person be paged at 2 A.M to repair this? Absolutely not.
For who?
For an application, the job will just get rescheduled onto a different machine, and you're N+2 or whatever so no one will notice the restart.
For the datacenter, since (for most jobs) you can assume that no one will care, you don't necessarily monitor for single machines overheating, but rack or row level issues, if a machine keeps overheating, eventually "nothing runs on this machine for more than an hour without it restarting" -symptom based alerting kicks in and tells someone to look at the machine.
Go and take a look at the rack? You'd have found and fixed the problem immediately in this case.
Where warnings are useful are when "something is actually now going wrong" - they provide a valuable context that can help an engineer figure out what the issue might be, which is exactly how it worked in this case.
This is exactly what the blog post described
The truth, of course, is somewhere is the middle. Symptom-based alerting is extremely tricky to get right in the first place. What is a “within budget” for a service, can be total outage for a customer. Real example from the past: we paged GCS “error rate within this bucket is high, snapshot restore timeouts”, reply is “it is within our SLO”.
Another situation is when problems in background jobs will cause massive outages later. Again, real examples from the past: zone out of quota and, on another instance, multi-million bill due to GC not being monitored correctly.
I agree, as I have stated in other comments, cause based signals should not alert, they should file a ticket, or be displayed along side a symptom based alert if there is a correlation. You should be monitoring as much as possible so you can observe your service in realtime.
> Symptom-based alerting is extremely tricky to get right in the first place.
Depends on the service. If you operating a HTTP API, a simple SLI would just be the error ratio of 500s to total requests. Its not perfect but it is much better than alerting based on CPU percentage or ram usage for a particular machine. Things can get more complex if your API is k8s style or you are operating a data plane that is serving customer traffic.
> What is a “within budget” for a service, can be total outage for a customer.
Now you are getting into SLO territory, which is similar to, but separate from symptom based alerting. Yes, defining good SLOs is very hard and varies greatly by the type of service you are running
> another situation is when problems in background jobs will cause massive outages later.
How do you know what signals to alert with? If you knew what causes/signals to alert, then wouldn't you design your backend jobs to not do the things that would cause an outage? Hindsight is always 20/20. If you are worried that a caused based signal from your backend _could_ be telling you that there _might_ be a problem in the future but your customers are not seeing errors, it is not business critical and you can just file a ticket and someone can look into it later.
Imagine if you didn't replace drives in your RAID until it hurt the user!
The tool also had the idea of noise built in so it would suggest that in general three monitors could be out of a channel and might not alert unless six were. Still we very rarely turned on alerting out of the tool because so little of the data was actionable.
It was still a huge improvement over the Tivoli system it replaced that had one size fits all alerts which generated tens of thousands of unactionable tickets monthly.
The most interesting thing it caught for me was when a system deviated on the low side of a channel and I discovered one of three DNS servers was no longer answering queries.
So Google seems to have the right of it (at least for their business) by measuring user experience and reacting to those metrics and leaving infrastructure metrics for post mortem.
Sounds more embarrassing to me.
That's fairly basic mistakes.
In a huge system with lots of failover you won't notice all the contingencies they identified planed for and successfully mitigated. Then something breaks which; you thought would be covered under another redundancy, was covered until a recent change (firmware change to the HVAC system), you didn't plan for etc...
And everyone points at the failure and says "that seems obvious", meanwhile the mountain of tests and monitoring and redundancy goes unnoticed.
[Not to say that every company out there would do this, nor that this wasn't necessarily considered, but definitely questioning the merit/objectivity of a brag-piece that is trying to rebuild the waning google hype]
I (personally) would not page in the middle of the night for a temperature issue. The systems should be resilient enough to handle some cpu throttling. A non-transient issue like this would probably end up in a ticket queue.
* why do the racks have wheels at all? Doesn't seem like a standard build, and turns out the be risky
* there should be at least daily checks on the data center, including a visual inspection of cooling systems and the likes. I don't know if daily visual inspection of racks is also a thing, but should find that pretty quickly.
* monitoring temperatures in a data center is pretty essential, though I must admit I don't know enough if rack-level temperature monitoring would have caused overheating of CPUs in the rack.
They probably aren't "standard" in any kind of common sense. Google and the other big tech firms design their own hardware and likely have it assembled remotely and delivered as a complete rack. Wheels make positioning it easier. It's also not uncommon to need to move racks of machines around (say, during a cluster refit or to change capacity in different DC halls). Not having wheels is likely more of a pain than the potential gains of omitting them.
> * there should be at least daily checks on the data center, including a visual inspection of cooling systems and the likes. I don't know if daily visual inspection of racks is also a thing, but should find that pretty quickly.
Google has millions of machines and dozens of datacenters. Its techs are likely already 100% busy replacing HDDs and other routine stuff. Noticing a single rack failing (within error budget) isn't going to be a priority so it's unlikely someone is employed to do that. Cooling/power systems are definitely monitored on a more macro level.
> * monitoring temperatures in a data center is pretty essential, though I must admit I don't know enough if rack-level temperature monitoring would have caused overheating of CPUs in the rack.
There's likely DC-level temperature monitoring. Again, systems are designed to cope with some amount of failure so it's not important to monitor everything and alert someone every time anything changes. When dealing with millions of things you need to look at aggregations.
If you thought "couldn't Bob of shimmed it with a block of plywood?", you might want to read up on continuous-improvement. Have Bob put the shim in to fix the problem right quick, then start up the continuous improvement engine...
Ah, how nice would it be if every company had infinite engineering capacity.
"Someone at Google bought cheap casters designed to hold up an office table, and put them in the bottom of a server rack that weighed half a tonne. They failed. Tens of thousands of dollars were spent replacing them all"
A more neutral way to tell it might be: "Some of our servers were overheating. We took them out of service, then discovered they were overheating because of an issue where the rack hardware had failed and tiled the rack. We eliminated this problem from this and all other racks."
The way it's currently written seems to try to set Google apart, but I'm not sure why it wouldn't work this way at any organization. Maybe I have missed the point.