GCP Outage
status.cloud.google.com
status.cloud.google.com
"Chemist checks the project status, activation status, abuse status, billing status, service status, location restrictions, VPC Service Controls, SuperQuota, and other policies."
-> This would totally explain the error messages "visibility check (of the API) failed" and "cannot load policy" and the wide amount of services affected.
cf. https://cloud.google.com/service-infrastructure/docs/service...
EDIT: Google says "(Google Cloud) is down due to Identity and Access Management Service Issue"
https://www.cloudflarestatus.com/
At Cloudflare it started with: "Investigating - Cloudflare engineering is investigating an issue causing Access authentication to fail.".
So this would somehow validate the theory of auth/quotas started failing right after Google, but what happened after ?! Pure snowballing ? That sounds a bit crazy.
DownDetector also reports azure and oracle cloud, I can't see then also being dependant on GCP...
I guess down detector isn't a full source of truth though.
https://ocistatus.oraclecloud.com/#/ https://azure.status.microsoft/en-gb/status
Both green
It's often things like "we got backpressure like we're supposed to, so we gave the end user an error because the processing queue had built up above threshold, but it was because waiting for the timeout from SaaS X slowed down the processing so much that the queue built up." (Have the scars from this more than once.)
My apps run on AWS, but we use third parties for logging, for auth support, billing, things like that. Some of those could well be on GCP though we didn't see any elevated error rates. Our system is resilient against those being down- after a couple of failed tries to connect it will dump what it was trying to send into a dump file for later re-sending. Most engineers will do that. But I've learned after many bad experiences that after a certain threshold of failures to connect to one of these outside system, my system should just skip calling out except for once every retryCycleTime, because all it will do is add two connectionTimeout's to every processing loop, building up messages in the processing queue, which eventually create backpressure up to the user. If you don't have that level of circuit breaker built, you can cause your own systems to give out higher error rates even if you are on an unaffected cloud.
So today a whole lot of systems that are not on GCP discovered the importance of the circuit breaker design pattern.
"Cloudflare’s critical Workers KV service went offline due to an outage of a 3rd party service that is a key dependency. As a result, certain Cloudflare products that rely on KV service to store and disseminate information are unavailable"
As such, I wouldn't be overly surprised if all of CF's non-edge compute (including, for example, their control plane) is just tossed onto a "competitor" cloud like GCP. To CF, that infra is neither a revenue center, nor a huge cost center worth OpEx-optimizing through vertical integration.
Scale is just a way to keep costs low. In addition to economies of scale, routing tons of traffic puts them in position to negotiate no-cost peering agreements with other bandwidth providers. Freemium scale is good marketing too.
So there is no strategic reason to avoid dependencies on Google or other clouds. If they can save costs that way, they will.
So long as the outages are rare, I don’t think there is much downside for Cloudflare to be tied to Google cloud. And if they can avoid the cost of a full cloud buildout (with multiple data centers and zones, etc…), even better.
Plus their past outage reports indicate they should be running their own DC: https://blog.cloudflare.com/major-data-center-power-failure-...
> Cloudflare’s critical Workers KV service went offline due to an outage of a 3rd party service that is a key dependency. As a result, certain Cloudflare products that rely on KV service to store and disseminate information are unavailable [...]
Surprising, but not entirely unplausible for a GCP outage to spread to CF.
Good to know that Cloudflare has services seemingly based on GCP with no redundancy.
Cloudflare adverises themselves as _the_ redundancy / CDN provider. Don't ask me for an "alternative" but tell them to get their backend infra shit in order.
Im not sure who is running the show there, but the whole thing seems kinda shoddy given cloudflares position as the backbone of a large portion of the internet.
I personally work at a place with less market cap than cloudflare and we were hit by the exact same instances (datacenter power went out) and had almost no downtime, whereas the entire cloudflare api was down for nearly a day.
https://www.businessinsider.com/google-return-office-buyouts...
Nooooo I'm going to have to use my brain again and write 100% of my code like a caveman from December 2024.
Devs during June 12, 2025 GCP outage: "What, no AI?! Do you think I'm a slave?!"
This is a good wake-up call on how easily (and quickly) we can all become pretty dependent on these tools.
Update - We are seeing a number of services suffer intermittent failures. We are continuing to investigate this and we will update this list as we assess the impact on a per-service level.
Impacted services: Access WARP Durable Objects (SQLite backed Durable Objects only) Workers KV Realtime Workers AI Stream Parts of the Cloudflare dashboard Jun 12, 2025 - 18:48 UTC
(Its Finnish inventor is incidentally working for Google in Stockholm, as per https://en.wikipedia.org/wiki/Jarkko_Oikarinen)
Supposedly there are plans for how to conduct a "cold" start of the system, but as far as I know it's never actually been tried.
Presumably for the infrequently read configs you could do this so the service with frequently read configs can bootstrap without the service for infrequently read configs.
For example a service discovery system periodically serializes peers to disk, and then if the whole thing falls down we have static IP addresses for a node and the service discovery system can use the last known IPs of peers to bring itself back up.
The have software running in most ISPs around the world:
https://help.speedtest.net/hc/en-us/articles/360039164793-Ho...
(OTOH, it's not always trivial to define/detect an outage.)
Cloud flare was really the gcp problem. Most of the others are going to be dependencies on cf or random Google stuff.
Discord for example was gcs for updates, etc
Was not basically everything (hyperbolically speaking, of course) practically impacted today?
How much weight really comes from those social media posts? Is there an indirect effect of people reading these posts, then flocking to hit the report button, sight unseen?
(downdetector infra also likely affected)
We are truly in gods hands.
$ prod
Fetching cluster endpoint and auth data. ERROR: (gcloud.container.clusters.get-credentials) ResponseError: code=503, message=Visibility check was unavailable. Please retry the request and contact support if the problem persists
But really not until after it's been on CNN a while.
How hard is it for the frontend to detect if the last update to the status page was made a while ago and that itself implies there is an error and should be reported ?
But why would the frontend have processing logic when all you need is to serve a static HTML document?
Even if it did, what would you do with that information? Throw up a screen with: Call us for service information at 1-HAHA-JUST-KIDDING
It’s not like it really matters if it’s accurate anyway.
My understanding is this is part of Google's internal PSD offering (Public Status Board) which uses SCS (Static Content Service) behind GFE (Google Frontend) which is hosted on Borg, and deploys other large scale apps such as Search, Drive, YouTube, etc.
Of course, I recognize that you purposefully pulled the Cunningham's Law trigger so that you, too, would gain additional knowledge that nobody would have told you about otherwise, as one logically would. And that you played it off as some kind of derision towards doing that all while doing it yourself made it especially funny. Well done!
It is what it says on the tin: choosing to lie doesn't mean you want the truth communicated.
I apologize that it comes across as aggro, its just that I'm not quite as giggly about this as you are. I think I can safely assume you're old enough to recognize some deleterious effects of lying
You had no idea what it is. Now you know thanks to you the lie you told.
> choosing to lie doesn't mean you want the truth communicated.
But you're going to get it either way, so if you do lie, expect it. If you don't want it – don't lie, I guess. It is inconceivable that someone wouldn't want to learn about the truth, though. Sadly, despite your efforts in enacting Cunningham again, I don't have more information to give you here.
> I apologize that it comes across as aggro
It doesn't. Attaching human attributes to software would be plain weird.
> I think I can safely assume you're old enough to recognize some deleterious effects of lying
Time and place. While it can be disastrous in the right context, context is significant. It makes no difference in a place of entertainment, as is the case here. Entertainment has always been rooted in tales that aren't true. No matter how old you are, even young children understand that.
then such pages should report a partial failure. Indeed the GCP outage page lists an orange "One or more regions affected" marker, but all services show the green "Available" marker, which apparently is not true.
Downdetector reports started spiking over an hour ago but there still isn't a single status that isn't a green checkmark on the status page.
E.g. it was mostly us-central1 region affected, and in there only some services (e.g. regular instances, and GKE kubernetes were not affected in any region). So if we ask "what the percentage of GCP is down", it might well be it's less than the threshold.
On the other hand, about a month ago, 2025-05-19 there was an 8-hour long incident with Spot VM instances affecting 5 regions, and which was way more important to our company, but it didn't make any headlines.
The issue is that they don't want to. (For claiming good uptime, which may even be true for average user, if most outages affect only small groups)
I remember reddit being down for like a whole day or so and they claimed 99.5% in that month.
Not trying to snark, I legit got nerdsniped by this comment.
Engineering and Operations in the Bell System is pretty great for this.
It's a lot easier to keep packets flowing than to keep non-self-contained servers serving.
SLA agreements.
How could you possibly trust them with your critical workloads? They don't even tell you whether or not their services work (despite obviously knowing)
https://www.google.com/appsstatus/dashboard/
https://status.cloud.google.com/index.html
Edit: The GCP status page got updated <1 minute after I posted this, showing affected services are Cloud Data Fusion, Cloud Memorystore, Cloud Shell, Cloud Workstations, Google Cloud Bigtable, Google Cloud Console, Google Cloud Dataproc, Google Cloud Storage, Identity and Access Management, Identity Platform, Memorystore for Memcached, Memorystore for Redis, Memorystore for Redis Cluster, Vertex AI Search
The only accurate status pages are provided by third party service checkers.
Well, yes, incentives, do big customers with wads of cash have an incentive to demand accurate reporting from their suppliers so they can react better rather than trying to identify issues? If there's systematic underreporting, then apparently not. Though in this case they did update their page.
If there's systematic underreporting, then apparently not.
You answered your own question.If you think about it from the corp’s perspective, it makes perfect sense. They weigh the risk reward. Are they going to be rewarded for the radical transparency or suffer fall out by acknowledging how bad of a dumpster fire the situation actually is? Easier for the corp to just lie, obscure and downplay to avoid having to even face that conundrum in the first place.
This is my position.
Heroku was down for _hours_ the other day before there was any mention of an incident - meanwhile there were hundreds of comments across twitter, hn, reddit etc.
… I get that PR-types probably want to massage the message, but going radio dark is not good PR.
From the end user's perspective, if the carrier didn't use Jibe RCS, it simply wouldn't work well.
RCS was designed and specced, by GSMA, as a telco run decentralized system that would replace SMS as like for like; but there were only a handful of rollouts. It's really only gotten use as Google pushed it onto Android, using their RCS server; recently iOS started using it although I don't know what server they attach to.
Since RCS is basically the 5th wave Google IM, it's no surprise when they have a major outage, RCS is pretty much broken.
According to Wikipedia, only the carrier's RCS server is used [1]
[1]: https://en.wikipedia.org/wiki/Rich_Communication_Services#So...
"So shady"
It's really, really hard to make a status page realtime.
> What makes you think it’s hard?
Being responsible (or rather, on a team of people responsible) for a status page of a big tech co made me think it’s hard.
“Is it down?” Is not a binary question.
> Cloudflare’s critical Workers KV service went offline due to an outage of a 3rd party service that is a key dependency. As a result, certain Cloudflare products that rely on KV service to store and disseminate information
I would love if anyone has any good tool recommendations!
Here are some others, although some seem to be experiencing issues due to the current outage I can only presume.
- https://atlas.ripe.net/probes/public
- https://www.ihr.live/en/global-report
Why would you think this outage is (internet) BGP related?
"Cloudflare’s critical Workers KV service went offline due to an outage of a 3rd party service that is a key dependency."
I really hope CF explains this apparent Google dependecy in detail in their post mortem.
as per their api documentation [1], it might be linked to firebase, which might explain this?
Getting 504 errors on anything from registry.npmjs.org that isn't cached on my machine.
the much more disastrous situation would have been the irm fallback.
… Proceeds to show worldwide degraded service level alerts.
Also spotify isn't working for me so I assume that's also related.
These are my most important productivity resources! Sad!
Update - Cloudflare’s critical Workers KV service went offline due to an outage of a 3rd party service that is a key dependency.
Jun 12, 2025 - 19:57 UTC
1: https://www.cloudflarestatus.com/I'm sure it's not entirely impossible, but sounds backwards to me. Sure - a lot of the internet relies on Cloudflare, but I'd be very surprised if GCP had a direct dependency on Cloudflare, for a lot of reasons. Maybe I misunderstood your comment?
Kind of nice to not be glued to AI chat prompts for a while to be honest.
> Multiple GCP products are experiencing impact due to Identity and Access Management Service Issue
IAM issue huh. The post-mortem should be interesting at least.
Even a few years ago senior management knew to stay the fuck out except for asking for more info.
"good enough" is, indeed, just good enough to make it not worthwhile to rip out all the upstreams and roll your own everything from scratch, because the cost of the occasional outages is much lower than the cost of reinventing every single wheel, nut, bolt, axle, bearing, and grease formulation.
To demonstrate my point, say someone like cloudflare opted to get "off cloud" and run their own datacenters. Half the web would still go down if they had some issue in their datacenters. If anything, the economy of scale and resiliency of a huge cloud network is far beyond what any single operator can ensure for their own service. If this wasn't the case, cloud services couldn't be as profitable as they are.
It isn't the panacea people seem to think it is. One critical service going down, regardless of whether it's cloud hosted or not, has ripple effects in the broader network.
Google Cloud Console won't load.
Example: https://hub.docker.com/layers/library/eclipse-mosquitto/late...
A contact in google mentioned to me that some bad update to Google Cloud Storage service has caused some cascading issues affecting multiple GCP services.
But this time...
EDIT: Updated link to point to the specific incident.
Seeing how everything seems to be broken everywhere, I'm very much looking forward to the post-mortem.
To give an example, our web servers connect to our GCP CloudSQL database via a Cloud SQL Auth Proxy (there's also a connection pooler in between but that also stayed up). The connection to the proxy was always available, but the proxy wasn't able to renew auth tokens it uses to tunnel to the database, regardless of where the webserver or database happened to be located. To mitigate this in the future we're planning to stop using the auth proxy and connect directly via mutual TLS but now it means we have to manage TLS certificates.
https://status.cloud.google.com/
File that in the status pages worth ~0 category.
Ask HN: Is Firebase Down? - https://news.ycombinator.com/item?id=44260669
Crossing my fingers for a quick resolution.
Thankfully we use AWS at work for everything critical
anyone know what tech stack they use and where they host
Cloud console does nothing.
They should host their support services on AWS and vice-versa.
Did someone screw up BGP again?
Facebook, Reddit and Hacker News is still up, but thats about it
Good luck out there!
Seems obvious.
but no tech bros, just keep following your ketamine addled edgelord when he did this with twitter..
sslv3 alert bad certificate:../deps/openssl/openssl/ssl/record/rec_layer_s3
"Firebase Data Connect unavailable due to a known Google Cloud global outage"
While the Google Cloud status page https://status.cloud.google.com/ says "No major incidents" and everything is green. So Google Cloud know there is an outage but just deem it not major enough to show it.
Edit to add: within 10 minutes of this post Google updated their status page. More curiously the Firebase page I linked to has been edited to remove mention of Google Cloud in the status and now says "Firebase Data Connect is currently experiencing a service disruption. Please check back for status. ".
* https://www.datacenterdynamics.com/en/news/facebook-blames-m...
https://health.aws.amazon.com/health/status
Perhaps CF is dependant on some GCP services?
> https://health.aws.amazon.com/health/status
Historically, the worst place to figure out if AWS is up/down is Amazons own status page.
I at least have no issues on their services across a few regions, and their console works fine.
At least some of the information has to be.
The weird part is that it took them almost an full hour to update it.
EDIT: Looks like it has been updated now (6:49 PM UTC)
No major incident as of “ Last updated time: 12 Jun 2025, 11:48 PDT”
say it's not so!
On the other side of this, Firebase probably doesn't have money at stake making the update
We should probably avoid punishing them based on free-associating made by a random not-anonymous not-Googler not-Xoogler account on HN. (disclaimer: xoogler)
Is that true in this case or are you speculating? My company runs a cloud platform. Our strategy is to have outages happen as rarely as possible and to proactively offer rebates based on customer-measured downtime. I don't know why people would trust vendors that do otherwise.
(n.b. as much as Google in aggregate is evil, they're smart evil. You can't avoid execs approving every outage because checks without some paper trail, and execs don't want to approve every outage, you'd have to rely on too many engineers and sales people, even as ex-employees, to keep it a secret. disclaimer: xoogler)
(EDIT: for posterity, we're discussing a "overall status" thing with a huge refresh button, right above a huge table chockful of orange triangles that indicate "One or more regions affected" - even when the "overall status" was green, the table was still full of orange and visible immediately underneath. My point being, you gotta suppose a wholeeee bunch of stuff to get to the point there was ever info suppressed, much less suppressed intentionally to avoid cutting checks)
Thanks for letting me know about emailing the mods, refreshingly explicit to send email.
* just trying to add a little humour. pretty stressfull outage. grarr!!
The PHP application I wrote as a student running on a single self-hosted server had a higher uptime than any of the cloud providers or redundant system I have seen so far. If you don’t need the cloud for scalability, do it yourself and save yourself the trouble and money. Most companies would be better off investing into some IT staff instead of giving away their systems in the hands of some proprietary and insanely complex cloud environment. You are becoming dependent on someone you don’t know, have no control over and can’t talk with directly. Also the single point of failure is just shifting: from your system to whatever system is managing the cloud. Guess one advantage is that you can shift the blame to someone else…