Google Cloud Global Loadbalancer Outage
status.cloud.google.com
status.cloud.google.com
I don't want to remember if "US/Pacific" currently has daylight savings or not.
Its a very strange decision especially considering that GCP has numerous regions outside of "US/Pacific".
Coordinating a tz change over a network of that size is probably infeasible, so may as well push the pain on to customers
The leap second troubles in 2005 (GFS?) that led to NTP smearing were more memorable.
Wow. Just... wow. Thanks for that. That phrase brought back a flood of old memories I'd forgotten.
AWS has a similar issue internally, IIRC (oncall around DST switching times was fun with teams in three continents), but they are making baby steps to fix it. It's harder to fix than it sounds once in place though...
Yes, and it's still a <bad adjective here> decision. Actually, there should be no decisions on those things. It doesn't matter if it is your personal blog. Store events in UTC, no ifs or buts. Then convert for display. You better be able to fully articulate why you are not using UTC (scheduling real world events?), otherwise use UTC. Make it a tattoo if that helps. Again, UTC. Yes I know you are a single person and you don't have more than one server today, still use UTC. Thank your former self later on when you have to match events across timezones...
Text representation is another example. Use Unicode for strings unless you can clearly explain why that's the wrong representation. (Which Unicode? Pick UTF-8, unless you have a reason not to, in which case you'll probably know what the reason is).
Even the UTF-8/Unicode example doesn't hold 18-20 years ago. UTF-8 was completely niche in 2003 and there were a number of flavors of Unicode, not all of which nicely fellback to ASCII like UTF-8 does. If I'm distributing software in 2003, I would have stuck to ASCII unless I was sure I needed Unicode, and if then I would have likely opted for UTF-16 because that was what was supported then.
I wouldn't question anyones decision for using UTF-16 in 2018 for software that had been supported since 2003.
Hang on.. what are we talking about?
According to Wikipedia, AWS was launched in 2006, which is only 12 years ago, and GCP was launched in 2008, only 10 years ago.
You could argue that, even in 2006, UTF-8 wasn't firmly enough established as the "winner" to be a clear best choice, but defending an for avoiding Unicode entirely at that point, for a company with global ambitions, just seems disingenuous.
For the question of UTC, claiming naivete also lacks credibility, if only because Y2K brought to and kept at the forefront of everyone's minds the topic of date/time representation in computers, starting around 20 years ago.
This misinformation is perpetuated by AWS evangelists who claim SQS was the first service.
https://aws.amazon.com/about-aws/whats-new/2004/10/04/introd...
More importantly, for these services to run roughshod, two years later, over the technical assumptions of something like AWIS also seems reasonable to expect. If AWIS isn't integrated under IAMs, that's pretty telling.
Yes, we all know that storing date/time data is hard; they should use a standard timezone (without DST), and use time-locales to localize the timestamps to the user. We get it.
The reason this knowledge is so well known is because previous programmers (without this helpful guidance) made many different technical decisions for many reasons, and learned the hard way. They then publicized their failures for our benefit.
What was the last technical decision you made that turned out to be the wrong one? Did you make the wrong decision maliciously? Or were you doing the best you could with the information available? Would you really find it helpful for somebody to come along a decade later with the attitude "these guys have no idea what they were doing, I learned about these antipatterns years ago!" (usually followed by "better rewrite everything")
I'll go out on a limb and say most people here haven't started from a single datacenter and grew to a global service. These engineers deserve the benefit of doubt unless and until more information is available. They certainly don't need armchair analysis from some randos on hackernews.
/rant
But... it's not hard. That's the point. This isn't a hard decision and has nothing to do with regional vs global. The lesson was learned by the entire industry decades before the company existed so a modern engineering team making that mistake, and never fixing it, is definitely open for criticism.
Oh dear. This kind of personal attack is not ok on HN and we ban accounts that do this. Could you please not do it again, regardless of how un-self-aware another comment seems or how you feel in response?
Your comment would be fine without the first and last sentences. It's hard never to write such things but it's always possible to edit them out. That's what I do when something like that slips out.
https://news.ycombinator.com/newsguidelines.html
Edit: unfortunately it looks like you've been uncivil before, e.g. https://news.ycombinator.com/item?id=16928159. Please make sure not to in the future.
It's easy to be unnecessarily rude in comments, but I would have said the same were we talking in person. I don't believe in hiding behind anonymity to say something you wouldn't say otherwise.
All the same, I appreciate the input.
1. You MUST use a single, consistent timezone. Using the local timezone is a mess: is it the user's local? The host's local? What about server logs, which are typically text files where the timezone can't be displayed dynamically? What if you aggregate server logs across timezones? What if you ssh into a machine in a different timezone?
2. The logical timezone to use is the one in which the vast majority of your employees work, since having people subtracting 7/8 all the time is annoying.
You could argue that for customer-facing status updates like this, Google should use a dynamic timezone. That's fair, but I'm sure Google internally uses that status dashboard, so it could be very confusing and complicate coordination to mitigate the problem. I'd argue that customers would prefer that the problem get fixed slightly sooner over having to do a once-a-year-ish timezone conversion.
UTC was adopted 51 years ago.
Also, if your office happens to be in a timezone which observes DST, you are still screwed. Now you think your times are in localtime, but in fact they are offset by one hour. This can lead to very "fun" debugging sessions and time wasted.
It can be a minor annoyance, but you know what can be an even greater annoyance? Undoing a bad decision which has percolated across several data stores.
Sometimes we need to speak about events (verbally), or share screenshots of graphs, dashboards, or other data. We can make a habit of always stating the time zone when we talk, and make sure every dashboard/graph contains a time offset.
Or we can just pick a single arbitrary time zone and always use that.
There are a few internal tools at Google that assume your local time zone (I'm in EDT) -- those are far more confusing.
> Also, if your office happens to be in a timezone which observes DST, you are still screwed.
Has this really been a problem? A date+time is unambiguous. I can only see this being confusing if California follows through with abolishing DST.
...so pick UTC then?
Why is that any different than how it is right now for anyone not in the Pacific timezone?
Converting old logs isn't hard either since ISO timestamp formats include timezones, or just pick a date for the switchover. The only scenario that gets tricky is with scheduling where users expect local times that carry over daylight-savings boundaries, but it doesn't apply to most apps.
There was some talk of changing to UTC everywhere, but I'm not sure if that happened. The next two big companies I went to at least managed to run US/Pacific everywhere. But the problem is always, by the time someone who knows better comes along, it's a PITA to change.
Pain/annoyance is a recurring theme in this sub-thread as far as reasons go for not switching or not using UTC in the first place.
I strongly suspect, however, that this is one of those situation where the negative aspect is over-estimated.
Other times, merely debating/discussing it draws attention to the pain and serves to amplify it (or its perception). Just quietly implementing it risks offending key players, but the vast majority won't even know the difference.
Unfortunately this isn't the first time it's broken and it's starting to look like a bad choice. We use a CDN in front so were able to switch traffic around but it seems it's better to do the load balancing ourselves too instead of using GLB.
You should also be looking at multi-CDN.
GLB is also distributed but clearly is a single "service" within GCP with weaknesses that can take down the entire thing, probably because of the complexity and integration involved in its architecture. It's fast and convenient but the reliability just isn't there yet.
The only challenge is that you need global geographic load balancers and that means F5 and at least a million dollar.
Also, you will find out later that some dependencies were only running in a single location and services failed with the datacenter.
Everyone always forgets about DNS, until Dyn dies.
Yes, but Google takes this to another level, all their load balancers are one single giant point of failure.
> This is why lots of folks go multi-cloud.
Or, just use AWS, where each region is almost entirely independent (some specific APIs are global by their nature, like creating a new S3 bucket), but you shouldn't be using those APIs on your run-path.
Yes, it is less convenient (and it must be less convenient for AWS to build things this way than for Google who seem to prioritise their development experience over customer availability), but new features from AWS are making it easier to deploy to multiple regions and use cross-region failover.
What is going to be faster? Updating DNS records with TTL 3600 to point to a single data center or Google fixing their problem.
We host DNS at AWS, but servers in GCP. Should we use AWS's automatic DNS failover feature to cover for such a case?
We generally use 60 second TTLs, and as low as 10 seconds is very common. There's a lot of myth out there about upstream DNS resolvers not honoring low TTLs, but we find that it's very reliable. We actually see faster convergence times with DNS failover than using BGP/IP Anycast. That's probably because DNS TTLs decrement concurrently on every resolver with the record, but BGP advertisements have to propagate serially network-by-network. The way DNS failover works is that the health checks are integrated directly with the Route 53 name servers. In fact every name server is checking the latest healthiness status every single time it gets a query. Those statuses are basically a bitset, being updated /all/ of the time. The system doesn't "care" or "know" how many health status change each time, it's not delta-based. That's made it very very reliable over the years. We use it ourselves for everything.
Of course the downside of low TTLs is more queries, and we charge by the query unless you ALIAS to an ELB, S3, or CloudFront (then the cost of the queries is on us).
I was diagnosing a networking issue from one of our service providers last Friday. For whatever indeterminate reason DNS responses from R53 took upwards of 10-15 seconds to return. While I appreciate the non-configurable default TTL of 60 seconds for ELB is not plucked out of thin air and that actual issue seemed to be on the service providers side, the lower limit seems far too low for medium/high latency networks. I wish it was configurable.
What's worse is it looks like it's our site that is the issue, so we get the complaints and I have to dig through wireshark logs.
- Route 53 failover record * primary record: Google global load balancer IP * secondary record: Route 53 Geolocation set (really need that latency) - Elastic Load balancer record per region * routes to mirror region GCP IP address (ELB's application load balancer seems to able to point to AWS external IPs) * optionally spin up mirror infrastructure in AWS
Seems brittle. Does Azure support global load balancing with external IPs?
Does anyone have such (or similar) setup actually in production? How did it work today?
IP addresses as Targets You can load balance any application hosted in AWS or on-premises using IP addresses of the application backends as targets. This allows load balancing to an application backend hosted on any IP address and any interface on an instance. You can also use IP addresses as targets to load balance applications hosted in on-premises locations (over a Direct Connect or VPN connection), peered VPCs and EC2-Classic (using ClassicLink). The ability to load balance across AWS and on-prem resources helps you migrate-to-cloud, burst-to-cloud or failover-to-cloud.
Looks like you need an active VPN connection to access external IPs.
"The IP addresses that you register must be from the subnets of the VPC for the target group, the RFC 1918 range (10.0.0.0/8, 172.16.0.0/12, and 192.168.0.0/16), and the RFC 6598 range (100.64.0.0/10). You cannot register publicly routable IP addresses."
[1] https://docs.aws.amazon.com/elasticloadbalancing/latest/netw...
;; ANSWER SECTION:
s3.us-east-1.amazonaws.com. 5 IN A 52.216.165.117
One of the biggest, highest traffic, systems on the internet!I've done a few unplanned DNS failovers, and I agree with this. What can be real trouble though is if you're running a B2B app, and your customers corporate networks can be configured in any strange way. I've met real network admins who think they need to have high TTLs everywhere in order to protect themselves from root DNS DDoSes.
For example, the public wifi in the last Hackspace in Munich I visited did not honour my 10 second TTL.
But in my opinion there aren't enough of them to justify not using short TTLs. It's their problem after all if they don't honour websites' settings: Then they will see downtime when nobody else does.
Well, I would avoid any of GCP's 'Global' features, they are an availability risk.
AWS's approach is to rather have inter-region replication, and there are lots of new features that support this.
EDIT: App Engine and Kubernetes environments in our EU region appeared to go down, Compute Engine was okay.
This was the first article I saw after I woke up, somewhat refreshed. /feeling-vindicated
AWS doesn't have a comparable service to the global version of GCP's load balancer service, which is what went down. Other GCP services were only affected to the extent that they use global load balancers for ingress, which varies by service.
For most GCP services, the ingress method is under the user's control and a switch to regional load balancers in one or more regions (whether split through GeoDNS or through round-robin) would have been a workaround.
Admittedly one point of global load balancers is to be able to mitigate a lot of other outages... I guess the secondary lesson here is to keep a short TTL on top-level DNS entries which point to the load balancers, and ideally have two DNS providers in the mix too.
Intentionally, because then customers relying on a service like this, would have regular global outages.
This is the 3rd global outage GCP has had in less than a year. As far as I remember, AWS hasn't had one in many years (but, too many people put everything in us-east-1, so a short single-region outage - which happens very seldom - seems to take down half the internet).
Its like with airplane crashes vs car crashes. The chance that you get an accident with your car is significantly higher, but usually the airplane accident ends up on "BREAKING NEWS". Its just a matter of impact.
There's a difference between fatalities and accidents. Fatalities are surprisingly low. Plus, Air travel is only significantly safer per-km, per-journey it's not at all safer [1].
> Its just a matter of impact.
Poor choice of words.
[1]: https://en.wikipedia.org/wiki/Aviation_safety#Transport_comp...
Of course there are far more car journeys taken than plane ones, so the GP's point stands (and since the point was just to make an analogy, unless you disagree that plane crashes get far more news coverage than car crashes, I'm not sure what the point of disagreement was).
(note those numbers also come from the UK, 1990-2000. If you took the numbers from the US, the car fatalities would be significantly higher (per capita or per km) and if they were from the last decade the air fatalities far lower)
This might give you a clue, PushBullet seems to run behind CloudFlare though.
https://chrome.google.com/webstore/detail/ipvfoo/ecanpcehffn...
(There's a right-click option to look up each address on bgp.he.net, but that doesn't happen automatically, for privacy.)
Let's turn this around: could a typical data center + server uptime + service uptime equal that of a major cloud provider?