Google Cloud Is Having IO Issues in US-EAST1
status.cloud.google.com
status.cloud.google.com
I think that's pretty clear where the issue lies
If I were on GCP and had paying customers, I would use Cloud Spanner. It's expensive, but it's a good piece of technology.
The key is to choose your entity wiseley. I wrote an app inside Google that did on the order of 10,000 writes per second and it was no problem; each entity group only showed up once every minute or so.
However, it has been my personal experience that Cassandra handles instances going down much better than it handles instances getting very slow but staying online. In that case Cassandra will work extra hard to just spin its wheels. If you know that is what is happening you can be better off manually bringing those nodes offline.
That said, it seems like GCP has a real, significant problem with overall reliability. While I don't have numbers to compare, seems like I see frequent outage notifications on HN for GCP. They really need to focus more on reliability than new features.
The major disagreement here from me on this is that GCP does not support IPv6 (outside of IPv6 termination on load balancers) and I do not agree with startups launching without IPv6 support in 2019.
Looking at Cloud Compute, the one responsible for the outage today and I think the big one a few months ago (13 hours!!!), they still do achieve 99% uptime (having been down 81 hours total this year).
While 99% uptime "sounds" great, this is paltry in comparison to AWS, which has set the standard for 99.99% (also known as "4 9s") in reliability, translating to about an hour of downtime per year.
Just gives you some more perspective and concrete numbers. Perhaps one of the reasons you find GCP "faster" is due to more devotion of resources to "speed" over "reliability".
Can't seem to find them, only "selected" outages. The recent, several hours long, service degradation in Frankfurt EC2 isn't in the historical list, seems like it is only in the rolling status history?
AWS and GCloud seem to also be reporting disruptions completely differently, AWS reports on tons of smaller pieces which makes any "aggregate" uptime become something completely different than if you report on larger aggregates as GCloud seems to do.
Unless I'm missing something, I don't see how one could compare comparing service availability reasonably without running large numbers of "canary" instances on both providers to actually measure aggregate availability?
"We are experiencing an issue with Cloud SQL instances hosted in the us-east1 region, beginning at Saturday, 2019-12-07 10:00 US/Pacific. Symptoms: Some Cloud SQL instances hosted in this region are becoming unavailable, and are refusing connections. Self-diagnosis: Connections to the Cloud SQL instance are rejected. Workaround: None at this time Our engineering team continues to investigate the issue. We will provide an update by Saturday, 2019-12-07 12:04 US/Pacific with current details."
Discord's description of the issue sounds like issues with either zonal or regional SAN storage.
The closest thing to public literature about Colossus and D that we've ever published is the Procella paper, which describes the abstractions that Colossus provides for it. In some ways it's similar to GFS (RPC interface, writes are generally append/overwrite) but many things are completely different now.
You need something that provides the block store abstraction on top of the primitives exposed by Colossus/D. Think of something like what modern SSD do in order to work efficiently with the underlying flash memory.
Then you have to hook that adapter in your virtualization stack (e.g. kvm) so you can boot from the volume and mount it from inside the VM. You could implement a kernel module or do it internally in kvm/qemu somehow, but iSCSI provides a straight-forward way to implement this in user-space: you have a process on your physical machine that speaks iSCSI upstream, and speaks Colossus/D RPC downstream.
(I don't know if they still do this but I have a vague memory of somebody describing the stack of an early version of GCP while I was working there long time ago)
This design has only one hop to the storage node. Low-latency workloads benefit from this design, high-bandwidth workloads sometimes actually benefit from off-loading PD to another host. To do iSCSI with one hop, you need to implement iSCSI interceptor, and basically you would have same design with less flexibility for guest OSes.
The irony, of course, that all this is a lot of legacy technologies needlessly wasting computer power: guest file-system trying to communicate with 4K blocks with “block device”, which goes through multiple layers of queues, then is re-maps to another abstraction, which goes over network to multiple hosts, etc. Not a single cloud customer ever said “we are so excited to manage volume sizes and bandwidth quotas for PD”. Better design would be to implement true data center-level filesystem to better support container workloads and leave PD for legacy cases, but Google’s storage management is so detached from reality, that it’s impossible to do cross-organizational project like this.
2) Then, for VMs you can do FS driver, jump to VMM and booms, you are done, multiple legacy levels of re-packing and redirection are gone. For shared access cases you do NFS/Samba interceptor and then same code path as above.
This system would be highly beneficial not only for Cloud, but for other Google properties as well: it would provide normal posix FS to Borg jobs. Amounts of equilibristics required to use any open-source package is enormous and by this time exceeded costs of developing FS multiple times, ask YouTube, MySQL, package management, etc groups.
Another example of Google’s storage craziness is cross-dc storage. This should be low-level Colossus responsibility. Instead godzillion of teams implement their own, GCS, PD, Spanner, Placer, etc. Crazy.
I will say though, that most of the time I don't want to use open source stuff. Observability is pretty crap, I don't want nor need software that uses write() without fsync(), and the assumptions that most OSS makes about FSs gives me nightmares on Borg.
Some stacks are mostly normal, some are really odd... and then there's storage.
But that’s just a distraction. We hear about those outages once in a blue moon, because many rely on them. What we don’t hear about is that any given colo, managed service, or CSP customer’s apps go down on their own, all the time, not because of the colo or CSP.
Such outages are banal, so we forget how much more likely they are, and fail to risk-weight our engineering efforts accordingly.
Many distributed systems try to be reliable by retrying failed nodes' work, but designs aren't always sound.
and they start this move with an outage during the busiest time of the year
why not?
Rackspace is trying to get out of the public cloud game. We had $20k/mo account with them and the rep called me every month asking if we wanted their help (professional services) migrating to AWS. I was baffled but they admitted they don’t want to be in the Cloud game anymore and wanted to get customers off their platform.
Rackspace announced that it has completed the acquisition of Onica, an Amazon Web Services (AWS) Partner Network (APN) Premier Consulting Partner and AWS Managed Service Provider.
A mere infrastructure issue, that occurs with all providers from time to time, would hardly constitute a service "wind-down". Based on a very cursory review, GCP's official deprecation policy [1] doesn't seem all that dissimilar from everyone else.
In short, to answer your question: No.
Now if ad revenue collection went down...