Google Cloud Is Down
https://status.cloud.google.com/incident/compute/19003
Status page reports all green, however the outage is affecting YouTube, Snapchat, and thousands of other users.
https://status.cloud.google.com/incident/compute/19003
Status page reports all green, however the outage is affecting YouTube, Snapchat, and thousands of other users.
We're having what appears to be a serious networking outage. It's disrupting everything, including unfortunately the tooling we usually use to communicate across the company about outages.
There are backup plans, of course, but I wanted to at least come here to say: you're not crazy, nothing is lost (to those concerns downthread), but there is serious packet loss at the least. You'll have to wait for someone actually involved in the incident to say more.
There's some irony in that.
I’m not in SRE so I don’t bother with all the backup modes (direct IRC channel, phone lines, “pagers” with backup numbers). I don’t think the networking SRE folks are as impacted in their direct communication, but they are (obviously) not able to get the word out as easily.
Still, it seems reasonable to me to use tooling for most outages that relies on “the network is fine overall”, to optimize for the common case.
Note: the status dashboard now correctly highlights (Edit: with a banner at the top) that multiple things are impacted because Networking. The Networking outage is the root cause.
this column of green checkmarks begs to differ: https://i.imgur.com/2TPD9e9.png
Not long after that incident, they migrated it to something that couldn't be affected by any outage. I imagine Google will probably do the same thing after this :)
Like the black box on an airplane, if it has 100% uptime why don’t they build the whole thing out of that? ;)
Reminds me of when I was working with a telecoms company. It was a large multinational company and the second largest network in the country I was in at the time.
I was surprised when I noticed all the senior execs were carrying two phones, of which the second was a mobile number on the main competitor (ie the largest network). After a while, I realised that it made sense, as when the shit really hit the fan they could still be reached even when our network had a total outage.
except time
You want velocity for your dev team? You get that. You want better uptime? Your expectations are gonna have a bad time. No need for rapid dev or bursty workloads? You’re lighting money on fire.
Disclaimer: I get paid to move clients to or from the cloud, everyone’s money is green. Opinion above is my own.
With on-prem solutions, you can at least access the physical servers and get your data out to carry on with your day while the infrastructure gets fixed.
I didn’t know the cloud-to-butt translator worked on comments too. I forgot that was even a thing.
But yeah, it's still a thing, and the message behind it isn't any less current.
https://hackaday.io/project/12985-multisite-homeofficehacker...
I made IoT using cheap (arduino, nrf24l01+, sensors/actuators) for local device telemetry, MQTT, node-red, and Tor for connecting clouds of endpoints that aren't local.
Long story short, its an IoT that is secure, consisting of a cloud of devices only you own.
Oh yeah, and GPL3 to boot.
And then “your data is in my butt” was just a play on that.
You can run your own hardware and pull in multiple power lines without establishing your own country.
I’ve ran my own hardware, maybe people have genuinely forgotten what it’s like, and granted, it takes preparation and planning and it’s harder than clicking “go” in a dashboard. But it’s not the same as establishing a country and source your own fuel and feed an army. This is absurd.
Fun related fact: My first employee's main office was in former electonics factory in Moscow's downtown powered by 2 thermal power stations (and no other alternatives), which have exact same maintenance schedule.
...etc.
If you don't like that you can order a KVM-VM with dedicated cores at similiar prices and the problem is not yours anymore.
Who are you getting this steal of a deal from?
Cloud costs roughly 4x than bare metal for sustained usage (of my workload). Even with the heavy discounts we get for being a large customer it’s still much more expensive. But I guess op-ex > cap-ex
I've never seen any of the providers listed offer "tons of ram" (unless we consider hundreds / low thousands of megabytes to be "tons") at that price point.
There’s a fine line or at least some subtlety here though. This leads to some interesting conversations when people notice how hard I push back against NIH. You don’t have to be the author to understand and be able to fiddle with tool internals. In a pinch you can tinker with things you run yourself.
There are also advantages to being part of the herd.
When you are hosted at some non-cloud data center, and they have a problem that takes them offline, your customers notice.
When you are hosted at a giant cloud provider, and they have a problem that takes them offline, your customers might not even notice because your business is just one of dozens of businesses and services they use that aren't working for them.
It's a fake trade-off, because you're choosing between lo-tech solution and bad engineering. IoT would work better if you made the "I" part stand for "Intranet", and kept the whole thing a product instead of a service. Alas, this wouldn't support user exploitation.
It's also my Plex media server, file server, VPN, I run some containers on there. I used to use it as a print server but my new printer is wireless so I never bothered
If there are not locks that work this way it sure seems like there should be. Using cloud services to enable cool features is great. But if those services are not designed from the beginning with fallback for when the internet/cloud isn't live that is something that is a weakness that often is unwise to leave in place imo.
If the cloud is down, revocations aren't going to happen instantly anyway. (Although you might be able to hack up a local WiFi or Bluetooth fallback.)
This does mean you need to setup a code in advance of people showing up, but it's an under 30 second setup that I've found simpler than unlocking once someone shows up. The cameras dropping offline are a hot mess though, since those have no local storage option.
http://www.ktvu.com/news/mistaken-identity-nest-locks-out-ho...
Sounds like Google and Amazon are hiring way too many optimists. I kinda blame the war on QA for part of this, but damn that’s some Pollyanna bullshit.
Even if you did have to self-assess, better to pay later than right away.
At least in EU services bought from overseas are subject to reverse charge, i.e. self-assessment of VAT (Article 196 of https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:02... ).
Though note that if you are an EU AWS customer, you are not buying from outside EU, you are buying from Amazon's EU branches regardless of AWS region. If Amazon has a local branch in your country, they charge you VAT as any local company does. Otherwise you buy from an Amazon branch in another EU country, and you again need to self-assess VAT (reverse charge) per Article 196.
Since AWS built a DC in Canada, I’m paying HST on my Route53 expenses, but not on my S3 charges in non-Canadian DCs.
I’m not an HST registrant (small supplier, or if you’re just using services personally), so there’s nothing to self-assess.
Even if self-assessment was required, you get some deferral on paying (unless you have to remit at time of invoice?).
I believe it works differently in EU (i.e. US DCs taxed) as per Article 44 the place of supply of services is the customer's country if the customer has no establishment in the supplier's country.
IBM/Softlayer, Rackspace, Google Cloud, Microsoft and I imagine everyone else large enough to count also does, too.
For Australian businesses, at least, being charged GST isn't a problem - they can claim it as an input and get a tax credit[1].
[0] https://aws.amazon.com/tax-help/australia/
[1] https://www.ato.gov.au/Business/GST/Claiming-GST-credits/
AWS tries to lock people in to specific services now which makes it really difficult to migrate. It also takes a while before you get to the tipping point where hosting your own is more financially viable .. and then if you trying migrating, you're stuck using so many of their services you can't even do cost comparisons.
For the downvoters, please just link here the proof if you disagree.
Here are the S3 numbers: https://aws.amazon.com/s3/sla/
This is a pretty neat and concise read on ObjectStorage in-use at BigTech, in case you're interested: https://maisonbisson.com/post/object-storage-prior-art-and-l...
16 9's and aws should easily last as long as the great pyramids without a second worth of outage.
What a joke
There's perhaps the additional asterisk of "and we haven't suffered a catastrophic event that entirely puts us out of business". (Which is maybe only terrorist attacks). Because then you're talking about losing data only when cosmic-ray bitflips happen simultaneously in data centers on different continents, which I'd expect doesn't happen too often.
It's about losing entire data centers to massive natural disasters once in a century.
I'm sure that's okay if you do bulk processing / time-independent analysis, but don't host production assets on wasabi.
Although I guess depending on how your own infrastructure is setup, even a multi cloud provider setup won't save you from a network outage like the current Google cloud one.
> Here are the S3 numbers: https://aws.amazon.com/s3/sla/
99.9%
https://azure.microsoft.com/en-au/support/legal/sla/storage/...
99.99%
> 99.9%
(single-region)
There doesn't seem to be an SLA on S3-cross-region-replication configurations, but I am not aware of a multi-region S3 (read) outage, ever.
> https://azure.microsoft.com/en-au/support/legal/sla/storage/....
> 99.99%
99.99% is for "Read Access-Geo Redundant Storage (RA-GRS)"
Their equivalent SLA is the same (99.9% for "Locally Redundant Storage (LRS), Zone Redundant Storage (ZRS), and Geo Redundant Storage (GRS) Accounts.").
Challenge Accepted... and defeated: https://blogs.dropbox.com/tech/2016/03/magic-pocket-infrastr...
but to be fair, storage is core to Dropbox's business... this is not true for most companies.
disclaimer: I work for Dropbox, though not on Magic Pocket.
"After a 2012 storm-related power outage at Amazon during which Netflix suffered through three hours of downtime, a Netflix engineer noted that the company had begun to work with Amazon to eliminate “single points of failure that cause region-wide outages.” They understood it was the company’s responsibility to ensure Netflix was available to entertain their customers no matter what. It would not suffice to blame their cloud provider when someone could not relax and watch a movie at the end of a long day."
https://www.networkworld.com/article/3178076/why-netflix-did...
I understand this is long-since resolved (I haven't tried building a service on Amazon in a couple years, so this isn't personal experience), but centralized failure modes in decentralized systems can persist longer than you might expect.
(Work for Google, not on Cloud or anything related to this outage that I'm aware of, I have no knowledge other than reading the linked outage page.)
Maybe you mean region, because there is no way that AWS tools were ever hosted out of a single zone (of which there are 4 in us-east-1). In fact, as of a few years ago, the web interface wasn’t even a single tool, so it’s unlikely that there was a global outage for all the tools.
And if this was later than 2012, even more unlikely, since Amazon retail was running on EC2 among other services at that point. Any outage would be for a few hours, at most.
"Some services, such as IAM, do not support Regions; therefore, their endpoints do not include a Region."
There was a partial outage maybe a month and a half ago where our typical AWS Console links didn't work but another region did. My understanding is that if that outage were in us-east-1 then making changes to IAM roles wouldn't have worked.
Your quote cd mean two things.
- that IAM services are hosted in one region (not one AZ)
And/Or
- that IAM is for the entire account not per region like other services (which is true)
(I will note that I was technically more right in the most obnoxiously pedantic sense since the hyphenation style you used is unique to AWS - `us-west-1` is AWS-style while `us-west1` is GCE-style :P)
Edit: ah, looks like the LB is sending LA traffic to Oregon.
So memegen is down?
Shouldn't that outage system be aware when service heartbeats stop?
Could this be a solar flare?
Can confirm with Gmail in Europe. Everything works but it's sluggish (i.e. no immediate reaction on button clicks).
Cloud services live and die by their reputation, so I'd be shocked if Google ever tried to get out of following an SLA contract based on a technicality like that. It would be business suicide, so it doesn't seem like something to be too worried about?
According to https://twitter.com/bgp4_table, we have just exceeded 768k Border Gateway Protocol routing entries, which may be causing some routers to malfunction.
I was actually surprised, as they tend to have excellent networking. Now I'm not nearly as distrusting as I was initially, knowing it was likely their ISP getting screwed by routing table overflow.
Might be a good month to rebuild all your models ;)
I would pay a premium for a cloud provider happy to give 100 percent discount for the month for 10 minutes downtime, and 100 percent discount for the year for an hour's downtime.
Minimum spends and a 50,000% markup based on adding that term to your contract.
But let's work backwards from the goal instead.
If you charge twice as much, and then 20-30% of months are refunded by the SLA, you make more money and you have a much stronger motivation to spend some of that cash on luxurious safety margins and double-extra redundancy.
So what thresholds would get us to that level of refunding?
Besides, a provider credit is the least of most company's concerns after an extended outage, it's a small fraction of their remediation costs and loss of customer goodwill.
Just take the premium that
you'd be willing to pay and
put it in the bank
In my country, when companies are hired to do overnight rail maintenance, they face very stiff fines if they over-run and delay trains the next morning.The fines are large enough that (for example) companies will have a heavy plant mechanic on site who does nothing on the vast majority of jobs - they're just standing by, to mitigate the risk of a breakdown leading to such a fine. Some business analyst with a spreadsheet has worked out the heavy plant breakdown rate, the typical resulting delays, the expected fines, and the cost of having the mechanic on standby... and they've worked out it's a good business decision.
The purpose of having an SLA isn't to get yourself money when your provider fails. The purpose is to make costly risk mitigation a rational investment for your suppliers.
> I would pay a premium for a cloud provider happy to give 100 percent discount for the month for 10 minutes downtime, and 100 percent discount for the year for an hour's downtime.
It takes a lot of effort (exponential) to reliably (I. E. Designed to fail-working) build something that is guaranteed to have this level of uptime at these penalties.
So I'm sure that I can build something that works like this, but would you pay me $100 per GB of storage per month? $100 per wall-time hour of CPU usage? $100 per GB of Ram used per hour? Because these are the premium prices for your specs.
From that linked page:
"Customer Must Request Financial Credit
In order to receive any of the Financial Credits described above, Customer must notify Google technical support within thirty days from the time Customer becomes eligible to receive a Financial Credit. Customer must also provide Google with server log files showing loss of external connectivity errors and the date and time those errors occurred. If Customer does not comply with these requirements, Customer will forfeit its right to receive a Financial Credit. If a dispute arises with respect to this SLA, Google will make a determination in good faith based on its system logs, monitoring reports, configuration records, and other available information, which Google will make available for auditing by Customer at Customer’s request."
AWS refunded me in the first reply on the same day!
GCP sales rep just copy pasted a link to a self support survey that essentially told me, after a series of YES or NO questions that they can't refund me.
So why not just tell your customers like it is? Google Cloud is super strict when it comes to billing. I have called my bank to do a chargeback and put a hold on all future billing with GCP.
I'm now back to AWS and still on a Free Tier. Apparently the $300 Trial with Google Cloud did not include some critical products, AWS Free tier makes it super clear and even still I sometimes leave something running on and discover it in my invoice....
I've yet to receive a reply from Google and its been a week now.
I do appreciate other products such as Firebase but honestly for infrastructure and for future integration with enterprise customers I feel AWS is more appropriate and mature.
The infinite money spout that is Google Ads has created a situation in which devs are at Google just to have fun - there really is no incentive to maintain anything because the money will flow regardless of quality.
Source: I interned at Google.
I really wanted to try out their new autoML but I was paranoid of entering my credit card and getting banned from Google
this is FUCKED. its aking to holding my youtube and google play accounts hostage.
That way Google won't ban your main account for non-payment.
It's the only way, especially considering Google cloud has no functionality to cap spending.
Phone and desktop OSs should grow a pair and create a virtualization protocol to randomize tracking info to keep PII anonymous.
For a ban, they need something concrete like using the same browser cookies, recovery email address or phone number.
IP isn't enough alone - you could be on shared WiFi.
Also, after account creation, you can log in from the same place without risk, or even use multi-login to log into both accounts at the same time.
With that said, If I delivered you services and then you credit card chargebacked me, I'd cut all relations with you as well.
AWS is mostly easy going.
Only some people at the partner programm can vary.
I had a guy who wanted to help me out even tho I was just a one person shop. After he left I got a woman who threw me out of the program faster than I could look.
There's too much liability. And no support.
>I have called my bank to do a chargeback
You're issuing a chargeback because you made a mistake and spent someone else's resources? And you're admitting to this on HN? I'm not a lawyer, but that sounds like fraud and / or theft to me.
It’s pretty convenient for companies like Comcast and Google that have poor customer service.
Of course, I get one free pass at that and if I did it over and over, I’m hosed. The difference is that my utility is regulated and has a phone number and a human whose job it is to talk to all customers.
OP sounds like they're just defending their selves from ambiguous draconian billing robots.
I think it's weird to say you get credit in dollars and then not be able to spend it on everything. That's not how money works. But that's the way hosting providers work and afaik it's quite well known. Especially with a large sum of "free money", even if it's not well known, it was on you to check any small print.
I didn't read it that way. I thought they were complaining about poor customer service that made it difficult to understand the bill or respond to it appropriately.
So no matter where you go for your cloud services, you're guaranteed a useless status page. Yippee.
https://www.whoishostingthis.com/#search=status.cloud.google...
And No, I don’t want to install a separate app to get push notifications about service disruptions for every service I use.
Now the web development side and I'm all "Wait a minute...are there any progress bars that are based on, anything real!?!?!"
I should have known...
Having an excel file where people enter statuses is not very useful to me as a customer. That’s more like a blog.
It's always interesting to see these outages at large cloud providers spider out across the rest of the internet, a lot of the world depends on Google to stay up.
Is there any reason to presume these statuses are correlated?
Apple's issue is
> Users may be experiencing slower than normal performance with this service.
https://techcrunch.com/2018/02/27/apple-now-relies-on-google...
Yup, I'm trying to check the Associated Press News right now and it's having trouble connecting to "storage.googleapis.com".
When the mainframe is down terminals are useless.
(For instance, I have a 500GB MicroSD card in my phone which contains a copy of my OwnCloud)
I don’t miss being on pager duty one bit. I see it looming in my headlights, sadly.
... but not for everybody now.
In Australia, many states have different dates for the queens birthday.
So not a nightmare at all.
https://www.theguardian.com/technology/2018/jul/25/big-tech-...
Nothing you or I or the pager can do will speed that up.
I am aware some bosses won't believe that and I am not trying to make light of it. But there really isn't much else to do except wait.
If you try to be heroic, you get back to 100% with a bunch of wasted effort and stress on your part.
Because it will be fixed by Google, regardless of what you do or don't do.
After the incident is over would be the time to consider alternatives.
The other case is really soft failures for multi-region companies. We degrade gracefully, but once that happens, the question becomes what other stuff can you bring back online. For example, this outage did not impact our infrastructure in GCP Frankfurt, however, it prevented internal traffic in GCP from reaching AWS in Virginia because we peer with GCP there. Also couldn't access the Google cloud API to fall back to VPN over public internet. In other cases, you might realize that your failover works, but timeouts are tuned poorly under the specific circumstances, or that disabling some feature brings the remainder of the product back online.
Additionally, you have people on standby to get everything back in order as soon as possible when the provider recover. Also, you may need to bring more of your support team online to deal with increased support calls during the outage.
I hope they come back. This is still pretty scary
So I wander over to my Firebase console, and there's no database loading. Thank god for twitter, and people also saying that they have the same issue or I would have for sure though we've been hacked.
I hope this is a good wake up call for everyone. I know that I'm going to think more about how we do backups and fail-safes
Of course, this is 2 weeks after switching everything over from AWS.
This is a networking issue, and your data is safe. Cloud SQL stores instance metadata regionally, so it shares a failure domain with the data it describes. When the region is down or inaccessible, instances are missing from the list results, but that doesn't say anything about the instance availability from within region.
Systems that fail 'open'...
Incident #19008 began at 2019-06-02 12:48. https://status.cloud.google.com/incident/cloud-networking/19...
Incident #19009 began at 2019-06-02 12:53. https://status.cloud.google.com/incident/cloud-networking/19...
Times are US/Pacific
The #19008 (networking) says there will be an update by "13:30 US/Pacific", but as of 17:05 there is no update.
Similarly, #19003 (GCE) says there will be an update by "16:00 US/Pacific", but no update as of 17:05.
All the latest updates seem to only be in the third incident #19009 (networking).
They don't want to admit fault or place blame because there can be legal and commercial ramifications, so they can only say canned responses.
So I searched for "gmail down" on bing and I got some results [1]. But searching on Google for "gmail down" does not return any results [2].
[1] https://www.bing.com/news/search?q=gmail+down&qs=n&form=QBNT
[2] https://www.google.com/search?q=gmail+down&source=lnms&tbm=n...
[21:55:19] POP< +OK send PASS
[21:55:19] POP> PASS ********
[21:55:21] POP< +OK Welcome.
[21:55:21] POP> STAT
[21:55:21] POP< -ERR [SYS/TEMP] Temporary system problem.
Please try again later.Nobody said this.
> "I care about how my providers behave when they have issues"
We all do.
As the other commenters stated, the communication is poor because the clouds are still growing rapidly and there's not much reason to be better. We might also be underestimating just how much more better service would cost and whether it's worth the revenue loss (if any). Are you really going to shift all of your spend overnight because of an outage? And where are you going to go?
The reality of these decisions is far more nuanced than it may seem and the current state of support is probably already optimized for revenue growth and customer retention.
Unless something is really fucked (like both GCP and AWS being down for us-east) incidents like these are not going to impact them at all.
The cost of either migrating to the other provider or, even worse, migrating to more traditional hosting companies is enormous and will require much more than "service was down for 2 hours in 2019". The contracts also cover cases like this and even if they don't, Google and Amazon can and will throw in some free treat as an apology.
On one hand I find this quite sad, but from a pragmatic point of view it makes sense.
Google started using a Beowulf cluster that the founders wired themselves. From the very beginning, the goal of metrics collection was to optimize costs. While today it’s seen as the cash cow, the focus has always been on cheap components strung together, relying on algorithms and code for stability and making the least possible demands of underlying hardware.
To think that they won’t try to save money any time they can seems implausible.
AWS had the S3 incident affecting all of us-east-1: “Other AWS services in the US-EAST-1 Region that rely on S3 for storage, including the S3 console, Amazon Elastic Compute Cloud (EC2) new instance launches, Amazon Elastic Block Store (EBS) volumes (when data was needed from a S3 snapshot), and AWS Lambda were also impacted while the S3 APIs were unavailable.”
The difference with Google Cloud is a lot of the core functionality (networking, storage) is multi region and consistent. The only thing thats a bit like that in AWS is IAM, however IAM is eventually consistent.
I'm overall happy with it, but if I needed to run a service with a 99.95% uptime SLA or higher, I wouldn't rely solely on GCP.
GCP has quarterly-ish global blackouts, and generally on the data plane at that which makes them significantly more severe.
The last time I looked at it (back when it showed more info for free, IIRC), AWS had the best uptime of the three big cloud providers, with Azure in 2nd and GCP in 3rd.
IIRC, the memorable thing was that, shortly afterwards, the head of Google Cloud made a big announcement that CloudHarmony showed that GCP had the best uptime when CloudHarmony showed that it actually had the worst. Google was calculating this by computing downtime = downtime per region * number of regions, but at the time, Azure had ~30 regions and AWS had ~15 vs. ~5 for Google and if you looked at average region downtime or global outage downtime, Google came out as the worst, not the best.
AWS, on the other hand, has given us very few problems. When we do have an issue with an AWS service, we're able to quickly get an engineer on the phone who, thus far, has been able to explain exactly what our issue is and how to fix it.
I'd love to know how this happens in the modern world. I've seen it myself only once (not GCP, but our own network with cisco equipment.)
Is something in the chain not checking the packet's CRC?
That's in own datacenters, not cloud.
Yeah, when it happened to me, it completely threw me for a loop. We had reports of corruption in video files, which started the debug cycle. It was shocking when we isolated the box causing the issue.
But I guess your bigger comment has to be right: About the only way to have this sort of error is at the hardware level, because basic CRC checking should otherwise raise some sort of alarm.
It wasn't just one box for us. Basically, the part number was defective (motherboard NIC), every single one that was manufactured. This affected a variety of things, since servers are bought in batch and shipped to multiple datacenters, damn impossible to root cause.
CRC can be computed by the OS (kernel driver) or offloaded to the NIC. I think it's unlikely for buggy CRC code to shipped to a finished product, it would be noticed that nothing works.
It's reasonably easy to run a router/gateway that has a 4G backup to get the ping out. Whether the ping works...
I wonder how often outages occur with other alarm monitoring companies. They certainly do occur, but customers don't have a lot of visibility into them.
No idea if this is what instagram does or not, just in general. Drives are hot and need lots of power and it’s expensive to out-S3 S3.
If you click through to something not on Google cloud, you see moderately elevated error rates (e.g., Instagram is up by 4x) but if you click through to something actually on Google, you see very highly elevated error rates (e.g., roughly 50000x for Snapchat).
If you read the "error reports", they actually report that Instagram isn't down (same for Twitter if you check Twitter). The error report detection seems to be just string matching. Here's an actual "error report" from downdetector that's the caused of allegedly elevated error rates:
> my twitter timeline: why isn’t snapchat working? anyone’s snapchat not working? snapchat’s being dumb. rip snapchat.
The Twitter "error report" is literally a report that Twitter isn't down.
i always thought downdetector was doing something a bit more clever than just reporting the rate of tweets containing the word "instagram" or "facebook" or something. but apparently not.
There appears to be some irregularities on consumer services as well that are of course certainly related, youtube was behaving a bit oddly for me.
The impact seems to be cascading down from just GCE to other services as well - that status page certainly does not reflect the reality of the situation. You can't even sign into GCP right now, and things that run on GCE, like appengine seem impacted.
It's amazing how far-reaching outages can be these days.
This is a networking issue, and your data is safe. Cloud SQL stores instance metadata regionally, so it shares a failure domain with the data it describes. When the region is down or inaccessible, instances are missing from the list results, but that doesn't say anything about the instance availability from within region.
is what I’ve heard so far. east seems to be OK, and Europe too
Why are they operating one with a different networking infrastructure from the other?
Since original Google infrastructure was developed specifically for first kind of services, cloud org still has problems adopting it to its needs.
One thing with gmail though. When it's down it's similar to a snow storm if you only do business in a city. Everyone is impacted and everyone understands a missed deadline is unavoidable.
[1] For those not old enough to know what I mean read this: https://www.ibm.com/ibm/history/ibm100/us/en/icons/personalc...
Edit: just got one email from the downtime, so perhaps my initial conclusion was incorrect
Better than the monthly outage from Azure.
As soon as I clicked that link, the client downloaded a PKG file, installed itself and launched itself without asking me if I wanted to share my camera or audio.
I uninstalled according to their instructions again, searched for all "zoom" files in my disk and rebooted.
This leads me to believe that following their uninstall instructions is insufficient, and there are hidden files left on my computer.
Sorry in advance for the off topic message
[1] https://support.zoom.us/hc/en-us/articles/201362983-How-to-u...
edit: Found this thread with details but no resolution it seems: https://apple.stackexchange.com/questions/358651/unable-to-c...
Deleting the .app file as instructed is not enough.
This StackExchange reply [1] showed me how to solve it, at least on macOS.
[1] https://apple.stackexchange.com/questions/358651/unable-to-c...
And even if I had, PKG files don't install themselves on macOS, they open the installer interface, AFAIK.
EDIT: And, as I mentioned, the file was downloaded by some Zoom client, not by my browser.
I can see my GKE clusters in one region but not in another, so I am guessing it's the former.
Looks like we'll need a cluster in each region going forward...
us-east4 is down
>We will provide an update at 16:00 US/Pacific.
it's 16:22 and no updates were posted. a bit unprofessional..
It might be something security related if it triggers a mandatory identity confirmation.
edit: I tried to send me a mail from another account and it worked but out of 4 or 5 mail checks at least two failed giving the same error.
[23:44:27] POP< -ERR [SYS/TEMP] Temporary system problem. Please try again later.
The problem seems much more complex.
I wanted to upload a video of the project to YouTube and add a link to it in the report. YouTube takes a long time to process the video, and then says it's unavailable.
I go to Vimeo: it's down.
I upload the video to Dropbox, and copy its link to the report.
But my report was a Google doc. And when I tried to export it as PDF (which I had not done yet) it couldn't do it. I never hated google more.
Eventually the video went through to YouTube, and I could export the PDF on the third try, but this really made me conscious of my dependance on Google.
WARNING: The following zones did not respond: us-west2, us-west2-a, southamerica-east1-c, us-west2-b, southamerica-east1, us-east4-b, us-east4, us-east4-a, northamerica-northeast1-c, northamerica-northeast1-b, us-west2-c, southamerica-east1-b, northamerica-northeast1, southamerica-east1-a, northamerica-northeast1-a, us-east4-c. List results may be incomplete.
Luckily for us eu-west1 seems to be working normally.
Pretty much every service is down
> Error: Download failed: server returned code 502. URL: https://storage.googleapis.com/chromium-browser-snapshots/Li...
The cloud components may be directly affected but for consumers, there's nothing which will provide info on what consumer facing services are getting some issues.
I'm not seeing anything at 12:47.
404 - Impressive
edit: wording
Unlikely the rest of AWS, a cached web page does not require much complexity.
Next update is in about 25 minutes.
The status page says GCS is fine but that's highly unlikely.
Scary stuff. What happens when Murphy's law decides to crash things even more?
The interesting thing is that a couple of minutes before everything went wrong, kubectl returned a "error: You must be logged in to the server (Unauthorized)" error
The networking incident looks like the one to follow for updates now.
> We are investigating an issue with Google Compute Engine. We will provide more information by Sunday, 2019-06-02 12:45 US/Pacific.
The next update is at 12:59. Just ... no.
GCE, GKE, BQ, Pub/Sub, GAE
asia-south1 us-west1 us-central1 us-west2
EDIT: It's been like that since at least 12h ago though. Not sure if it's connected to Google Cloud?
GitLab is no longer seeing errors and Google Cloud has resolved the issue as of 23:00 UTC yesterday. Any further information can be found on the issue at https://gitlab.com/gitlab-com/gl-infra/production/issues/862
(the problems with runners and the UI started at least at 2019-06-02 7:48 UTC, though they were hit-and-miss at the time)
Still, happy this is solved and we can use the (fantastic otherwise) service again!
especially if you use your bussiness for B2B services. Stuff like this could make you loose your bussiness, especially if some entity like google doesn't communicate and as a result, you do not have a answer for your own customers.
Medium sized private cloud providers are a lot better at this, considering the communication lines are a lot shorter.