Cloudflare was down
cloudflarestatus.com
cloudflarestatus.com
So they depend on GCP for (some of) their services
> Workers KV is in the process of being transitioned to significantly more resilient infrastructure for its central store: regrettably, we had a gap in coverage which was exposed during this incident.
https://www.cloudflare.com/gdpr/subprocessors/cloudflare-ser...
Google denies they had any outages.
Word on the street is that there are large BGP routing issues behind all of this.
At the same time I have not noticed anything being down firsthand. I am in Europe.
[1] https://bishopfox.com/blog/bgp-hijacking-technical-post-mort...
> Multiple GCP products are experiencing impact due to Identity and Access Management Service Issue
If this wasn't widespread, what is?
Incident affecting API Gateway, Agent Assist, AlloyDB for PostgreSQL, Apigee, Apigee Edge Private Cloud, Apigee Edge Public Cloud, Apigee Hybrid, Cloud Data Fusion, Cloud Firestore, Cloud Logging, Cloud Memorystore, Cloud Monitoring, Cloud Run, Cloud Security Command Center, Cloud Shell, Cloud Spanner, Cloud Workstations, Contact Center AI Platform, Contact Center Insights, Data Catalog, Database Migration Service, Dataform, Dataplex, Dataproc Metastore, Datastream, Dialogflow CX, Dialogflow ES, Google App Engine, Google BigQuery, Google Cloud Bigtable, Google Cloud Composer, Google Cloud Console, Google Cloud DNS, Google Cloud Dataflow, Google Cloud Dataproc, Google Cloud Pub/Sub, Google Cloud SQL, Google Cloud Storage, Google Compute Engine, Identity Platform, Identity and Access Management, Looker Studio, Managed Service for Apache Kafka, Memorystore for Memcached, Memorystore for Redis, Memorystore for Redis Cluster, Persistent Disk, Personalized Service Health, Pub/Sub Lite, Speech-to-Text, Text-to-Speech, Vertex AI Search
If some of those things listed had actual widespread outages, it would have been much much worse.
As a former SRE there, is "widespread outage" a specific, special kind of classification that's not obvious to the public just by looking at the status page...? Or what do you mean?
When it's your own outage, it's all-hands-on-deck panic mode. When it's half the internet down, it's no longer your problem, lol
If your application is mission-critical, downtime is anything but a holiday.
Currently down, but reference: https://blog.cloudflare.com/the-ddos-that-almost-broke-the-i...
Edit: The CF status page has acknowledged it's a broad outage across many services: https://www.cloudflarestatus.com/incidents/25r9t0vz99rp
Yes.
I really do appreciate the transparency and ownership that comes with these. We all fuck up, but a lot of companies would rather hide their mistakes than own up to them. Cloudflare's approach makes me trust them more.
edit:
It works in the US but EU customers are still reporting our services as down.
edit:
EU customers are reporting ok
Their API is down too.
Amazing that something can impact their whole infrastructure like this given how much redundance they have.
CDN and WAF seem to be working fine. I think CF rushed a lot of newer services out without the reliability some of their older/core services enjoy
> Cloudflare’s critical Workers KV service went offline due to an outage of a 3rd party service that is a key dependency.
I bet that 3rd party service is GCP.
I would be pretty pissed if I were a CF customer that used Workers KV for redundancy because it was heavily marketed as running on CF data centers.
> The cause of this outage was due to a failure in the underlying storage infrastructure used by our Workers KV service, which is a critical dependency for many Cloudflare products and relied upon for configuration, authentication and asset delivery across the affected services. Part of this infrastructure is backed by a third-party cloud provider, which experienced an outage today and directly impacted availability of our KV service.