Cloudflare Dashboard and API Service Issues
cloudflarestatus.com
cloudflarestatus.com
At least we'll get a good blog post out of it in a few days.
Tunnels are one of Cloudflare’s best features for developers, instant NAT traversal for home-hosted demos and prototypes. It’s a shame you have to pay for Argo on your entire domain in order to use Argo Tunnels even on one subdomain.
Cloudflare, please offer origin tunnels as a separate service rather than bundling it with Argo routing on the client side.
One is called "tiered routing" and is a dropdown in settings whereas the "nat" solution nowadays is implemented as cloudflared.
> Apr 15 17:01:37 mysite cloudflared[12412]: time="2020-04-15T17:01:37Z" level=info msg="Connected to SJC" connectionID=0
> Apr 15 17:03:09 mysite cloudflared[12412]: time="2020-04-15T17:03:09Z" level=error msg="Register tunnel error from server side" connectionID=0 error="Server error: Reached maximum retry 5: dial tcp 198.41.248.96:9100: connect: connection timed out"
> Apr 15 17:03:09 mysite cloudflared[12412]: time="2020-04-15T17:03:09Z" level=info msg="Retrying in 8s seconds" connectionID=0
> hera ra | time="2020-04-15T17:08:09Z" level=error msg="Unable to dial edge" error="DialContext error: dial tcp 198.41.200.233:7844: i/o timeout"
> hera | time="2020-04-15T17:08:09Z" level=info msg="Retrying in 1s seconds" hera | time="2020-04-15T17:08:09Z" level=error msg="Unable to dial edge" error="Handshake with edge error: read tcp 172.18.0.2:36016->198.41.192.227:7844: i/o timeout"
> hera | time="2020-04-15T17:08:09Z" level=info msg="Retrying in 1s seconds" hera | time="2020-04-15T17:08:11Z" level=error msg="Unable to dial edge" error="Handshake with edge error: read tcp 172.18.0.2:44626->198.41.192.7:7844: i/o timeout"
> hera | time="2020-04-15T17:08:11Z" level=info msg="Retrying in 4s seconds"
> hera | time="2020-04-15T17:08:13Z" level=info msg="Connected to EWR"
> hera | time="2020-04-15T17:08:13Z" level=error msg="Register tunnel error from server side" connectionID=0 error="Server error: Reached maximum retry 5: dial tcp 198.41.248.96:9100: connect: connection timed out"
> hera | time="2020-04-15T17:08:13Z" level=warning msg="Tunnel disconnected due to error" error="Server error: Reached maximum retry 5: dial tcp 198.41.248.96:9100: connect: connection timed out"
> hera | time="2020-04-15T17:08:19Z" level=error msg="Unable to dial edge" error="Handshake with edge error: read tcp 172.18.0.2:51974->198.41.192.107:7844: i/o timeout"
Argo routing’s ability to improve page load times depends on how international your user base is, and how poor your origin’s transit quality is. It’s great at improving the latency and reliability of a cheap host.
Maybe some piece of core network equipment.
A small colo may have a single pair of routers, switches, or firewalls at its edge. If one had failed for some reason, and the remote hands removed the wrong one, it is possible you could knock the entire colo offline.
There's a bunch of other possible components: Storage platforms, power, maybe something like an HSM storing secrets, or even just a key database server.
Their failover to their backup facility may be impaired by the fact that well, their management plane is down. They probably rely on their own services. Avoiding chicken-and-egg issues can require careful ahead-of-time planning.
The tweet says they're failing over to their backup facility. I would've expected that fail over to happen much faster.
Seems like they have two issues going on. First, the remote hands could take down the datacenter. Second, their fail over is taking this long to come online.
I also wonder how much Covid impacted the process, if at all.
I'm looking forward to the details after they get things back online.
Good luck CF engineers!
From this kind of mess to a fully functional infrastructure you need at least 12h-16h to function minimally. Probably takes 2-3 days to have that node working as before.
When I was on call, I always jocked with my colleages the worst incident you could have in a datatenter was someone swapping two cables.
I want my most of my mid-level people (and all of the ones bucking for a promotion) to be comfortable running 80-90% of routine maintenance operations and have pretty decent guesses on what should be done for the next 5-10%. I don't often get my way.
>"During planned maintenance remote hands decommissioned some equipment that they shouldn’t have. We’re failing over to a backup facility and working to get the equipment back online."
Some questions:
When the gear was first powered off why wasn't it just powered back up again as soon as the on-call person's phone started blowing up? Why does this require a "failover" to another datacenter?
Why are remote hands decommissioning Cloudflare's gear in the first place? Isn't this supposed to be a security focused company? For context "remote hands" is the term for people that work for the colocation facility such as Eqiunix, Telehouse, etc. They are not employees of the tenant(Cloudflare.) Remote-hands are great resources for things like checking a cable or running new cross connects etc but certainly not decommissioning gear without some form of tenant supervision.
Why is this a single point of failure?
How can you plan for things occurring out of your control? CF engineers are people as well. Things like this happen and there will be learnings to take out of it (like how to fail over faster)
Circulating a MOP(method of procedure) for data center maintenance among all stakeholder is pretty standard. The purpose of the MOPS is so that everyone can vet the plan(roll forward and roll back)and identify the risks.
As Matthew said on Twitter, this isn't the kind of mistake that happens twice. But for those of us who do operational work, it's easy to see how this happens once. Bit rot in a cab, work orders from someone who didn't know about the legacy patch panel or assumed too much in their instructions and a catastrophe. From now on, photos for every smart hands will probably be a part of the prep-procedure.
That's a pretty odd and flimsy argument that I can't know what I'm talking about if I hold an opinion that differs from your own. I have years of experience with this particular space. Had a MOP been circulated and the maintenance plan properly vetted then the need for visual confirmation because of "legacy patch panel" and "bit rot in cab" would have been identified. Further you never unplug a cable unless you first verify what the ends of the cable are connected to. As I mentioned this is "datacenter operations 101" stuff. This is not some new startup, this a publicly traded company who has been in the game for a decade now.
If this failure was so easy to avoid, it would have been avoided. It was the result of a combination of failures, as all failures in highly-available systems are. It was a failure in bitrot, in planning, in execution, in communication, in reluctance to failover due to a concern about failing back, etc.
That said, like Matthew said, this is the kind of failure that happens once.
I had to debug customer issues to find out that this was down.
Even if they don't want to offer this to the general public (free customers), they should have another notice mechanism for enterprise customers.
Error 522: If you're the owner of this website: Contact your hosting provider letting them know your web server is not completing requests. An Error 522 means that the request was able to connect to your web server, but that the request didn't finish. The most likely cause is that something on your server is hogging resources.
I can't wait to see the postmortem, I wonder if it's a DDOS, network/hardware issue or a deployment error
We’re pushing an emergency config change to skip the cache invalidation, which will stop it timing out but means republished projects won’t update (because the old version will still be cached).
Godspeed to the Cloudflare engineers who are presumably scrambling to fix this!
I've trying for 10 minutes or so, it just loading (test from various location using VPN)
This implies that CF hosts its API and dashboard all in one DC, which I find an _interesting_ observation. One would expect a company like CF to host its critical infrastructure in a redundant fashion.
The comments seem to imply that having a redundant way to refresh the page cache, even if it were global/domain versus page, would be an okay backup for many.
I'm really suprised that they hosted all those non-vital but still quite critical services in just one DC, or somehow had one DC as a single point of failure. Network issues happen "regular enough" to want to protect against that, or at least have mitigations available.
Such as S3. The bucket names are globally unique, which means that their source of truth is in a single location. (Virginia, IIRC.)
Now... a small thought exercise. If I wanted to take down a Cloudflare datacenter and I had access to a few suitably careless remote hands, I'd take out the power supplies to the core routers, and while the external network is out of commission, power down the racks where they have their PXE servers. That should keep anything, within the DC, from being unable to recover on its own.
Edit: Also noticed that when generating API keys, the dropdown wouldn't list all my accounts for setting permissions. Just assumed it was all related or something.
Either way, overall super insanely reliable product/service and could not live without.
Time to get some fresh air!
See: https://www.sec.gov/Archives/edgar/data/1477333/000119312519... (RBAC)