CloudFlare was down
cloudflare.com
cloudflare.com
- There is a global problem that affects the CloudFlare proxy and DNS services.
- The problem appears to be due to bad routing.
- We are working to restore correct routes in order to bring both DNS and proxy services back online.
- The operations and networking team are all online and treating this as an emergency.
- We do not have an ETA on the response time but will continue to post updates via Twitter as we learn more.
UPDATE. Sites are being restored now. DNS is operating.
I don't know all the details as bugging the network team while they were fixing wasn't going to help. We'll get a postmortem blog post up.
--
Seems to be getting better now - intermittent 502's and 504's, the occasional load (speed is choppy though)
--
Back to near-instant load times, great job guys! I look forward to the blog entry on this (P.S. It might be worth hosting your status page elsewhere in the future, whilst this might not happen very often, it's when it does that your site needs to be working, the lack of redundancy here is startling.)
1) The CloudFlare security model for SSL basically lets them MITM all your traffic. Probably not a big deal for SSLizing a normal website, or even for accepting credit cards), since they're a decent-sized US company with legal liability, although I'd be concerned about their internal security vs. your own internal security (since you're still fully exposed on your side, too -- it doesn't improve security, and can at best not be a source of new vulnerability).
2) Their DNS doesn't appear particularly redundant; it's just anycast in one big block. Using CloudFlare for DNS seems to be bad practice; you should use something else and cname to CF. Ideally something with multiple DNS servers either individually anycast or in at least two independent (probably anycast) netblocks.
3) Performance of the proxy service seems adequate in my experience but for sites with large amounts of overseas-source traffic, I've heard of people getting lots of suspected-bad-guy path. For a free forum like 4chan that's probably fine; for an e-commerce site, probably not.
That was over a year ago though, I assume things have improved as they got more and better data (and resources). I was in Costa Rica in January/February then again in August/September last year and didn't come across it I don't think.
I'm in Turkey at the moment and for the last couple weeks and haven't seen one at all, and that's including using an EC2 proxy because of the censorship here.
(They linked to a report on a spam blacklist claiming my IP address had been used to send spam during the past week. I'd had my IP address for about twelve hours. Thanks, TalkTalk!)
As a web site operator, you have the option of switching this behaviour off.
I hope the issue is resolved soon and if a person caused it, they're not in too much trouble.
Something I always wondered - do you use Cloudflare simply for the CDN in respects to photos, or does 4chan also frequently become the target of DDoS and other external attacks?
More on Railgun: http://arstechnica.com/information-technology/2013/02/cloudf...
Edit -- I am officially faster than carrier pigeon:
me: was i the first human to notify you guys? | me: i caught it within the first 30-60 seconds i think | me: because i have no life and never sleep | CloudFlare Ops pal: yeah. you did.
If anyone wants to hire me to check their site instead of Pingdom, feel free to ping me!
Barely an hour from outage to being live again - I barely had time to fire an email to my colleagues and update our Twitter before it came back online - however, what's worrying here is that there was no fallback in place whatsoever, we will be hearing from the Cloudflare team soon hopefully to let us know what went wrong, but if they are half the company they purport to be, they will learn a lot from whatever went wrong today.
How much would you charge to ping my site 24 hours per day? ^^
This is why every piece of kid should be mirrored with a redundant backup and why many businesses even have entire duplicated standby systems for such disasters. Even if that gear cannot support the entire infrastructure, it's usually enough to at least publish an official status page. Having to use Twitter to update users is just amateurish in my opinion.
No systems are immune to failure. No matter how much redundancy you have, chances are you have interdependencies you did not anticipate, and sooner or later run into failure scenarios that violate your expectations.
It's very well possible that Cloudflare messed up here, but to claim so categorically that "human error is unacceptable" is a bit of a joke. We build systems to withstand the risks we know about, and guess at some we don't.
But the number of possible failure scenarios we don't understand properly is pretty much infinite.
But what I was addressing was the blanket claim that human error is unacceptable. Anyone who runs a setup much larger than a calculator will deal with human errors - whether actual operational errors or human inability to engineer for resilience against all possible but unlikely scenarios - on a regular basis.
Some things should be harder than others to break, and DNS are amongst them. Cloudflare no doubt have plenty of lessons to learn. But so have everyone else.
So then you basically agree with my point <_<
> But what I was addressing was the blanket claim that human error is unacceptable. Anyone who runs a setup much larger than a calculator will deal with human errors - whether actual operational errors or human inability to engineer for resilience against all possible but unlikely scenarios - on a regular basis.
You're twisting my words and taking them out of context. I was saying that human error is an unacceptable excuse for the entire stack of a company the size of Cloudflare going off line. And I was saying that because redundancy systems should act as a "safety net" so that administrators can make human errors. I've lost count of the number of dumb mistakes I've made over the years, but each time I've been able to switch to a back up system while I've worked towards undoing my cock up. And you said yourself that a complete DNS outage was unacceptable, so clearly you and I are more or less on the same page regarding this.
Human failure is unacceptable when we know it happens, and should do everything possible to guard against it. You say "guess at some we dont". Well, we do know all about human failure.
All that sympathy for that bloke who "fired himself", because everyone here agreed that his potential human error should have been anticipated, but not so in this case? Seems to me we are applying different standards.
Most of all, what I dont like is universal get outs. Its reminds me of the worst lie of all: "Sorry Sir, its a computer error".
You miss the point. We can continue to merely enumerate possible error scenarios until the heat death of the universe, and we will still miss some.
It is "human error" for an operations team to not put in place methods for ensuring their systems stay up within agreed parameters.
But the reality is that it is not even theoretically possible for us to engineer a system which can guarantee no downtime. Furthermore, no organization is willing to pay the bill to address even a relatively small fraction of the problems we can easily predict for the reason that many even relatively likely reasons are more expensive to protect against than is worth.
So to begin with, we can't prevent failure. And even if we could, what from the outside looks like human error is often internally a result of either intentional budgetary constraints, or unintended consequences of lack of resources.
It is not about lack of responsibility. It about dispelling the fantasy that there is someone who is guilty of not doing their job correctly behind every failure.
That is not to say that there might not have been unacceptable human errors in this specific case. But that is entirely besides the point.
> but not so in this case?
I thought it was pretty clear that my comment applied to the general statement that "human error is unacceptable", but perhaps not. I explicitly wrote "It's very well possible that Cloudflare messed up here" because I didn't want to speculate on whether the specific problems in this case before the causes were even known.
> Most of all, what I dont like is universal get outs.
They are not "universal get outs". Nobody is going to say it is acceptable if the error is caused by someone bringing coffee into the ops room and spilling it all over the single server, for example. But there is a vast range between someone who is grossly negligent and/or incompetent and who should bear the blame, and someone who is doing their job as best as is reasonable given the resources available to them, but who still makes mistakes or oversights or simply don't have the time or resources to address some reasonably unlikely issue that eventually happens to cause downtime.
> No systems are immune to failure. No matter how much redundancy you have
You have that backwards. No systems are immune to failure, which is why you have redundancy.
> chances are you have interdependencies you did not anticipate, and sooner or later run into failure scenarios that violate your expectations.
If the dependencies haven't been anticipated then someone isn't doing their job right. There's a reason why incident response / disaster recovery / business continuity plans are written. It should be someone's job to think up every "what if" scenario ranging from each and every bit of kit dying, all your staff winning the lottery and walking out the next day and even terrorist attacks. I've even had to account for what would happen if nukes were dropped on the city where our main data center was housed (though the answer to that was a simple one: nobody would care that our site went offline). It might sound clichéd, but people get paid to expect the unexpected and work out how to maintain business continuity.
> It's very well possible that Cloudflare messed up here, but to claim so categorically that "human error is unacceptable" is a bit of a joke. We build systems to withstand the risks we know about, and guess at some we don't.
This was their infrastructure failing. If you own and maintain the infrastructure then you have no excuse not to work out what might happen if each and every part of that infrastructure failed. (trust me, I have had to do this in my last two jobs - despite your accusations of my "lack of experience" ;) ).
> But the number of possible failure scenarios we don't understand properly is pretty much infinite.
You're confusing cause and effect. The number of different causes for failure is infinite. But the effect is finite. For example: a server could crash for any number of reasons (hardware, software, user error, and all the different ways within those categories), but the end result is the same; the server has crashed. Thus what you do is plan for situations when different services fail (staff do not turn up for work, your domain name services stop responding, etc) and plan some kind of redundancy around that, thus giving a little more breathing time for engineers to fix the issue and with the minimum possible disruption to your users. As Cloudflare had to resort to Twitter to update their users, they completely failed every possible aspect of such planning. And given the high profile sites that depend on Cloudflare, they have no excuses.
If this happened in any of the other companies I worked for, I'd genuinely be fearful for my job as a crash of that magnitude would me that I hadn't done my job properly.
If this was a human error, then change the process to make it difficult to do. If it was a software problem, then they have just found a new set of automated tests to write. If it was a vendor problem, contact the vendor to see how they plan to prevent this problem in the future. If it was a design problem, change the structure of the system to make this type of problem less likely.
Quite a few big names have had similar outages in the past but the ones that I forgive (Google, Amazon) are the ones that talk openly about the issue and outline the changes that they are making to fix them. I'm looking forward to CloudFlare's explanation.
BTW: My boss asked me this week for more information on Railgun and if we should change CDN providers. Railgun sounds incredible and I think it could really help our platform. How CloudFlare responds to this incident is critical to my decision to move forward with Railgun testing or to just forget it completely.
We are in the same boat, but guess where I have the phone number for enterprise support? Webmail on our domain. Lesson learned.
Are we sure that if CloudFlare wasn't so popular here, the word scam wouldn't be used?
When disasters like this happen, a quick DNS change can be a life-saver.
EDIT: Sorry guys. We got some issues with a gem after installing the recently updated ruby2.0.0p0. The unicorn workers were timing out. TerrificDNS is completely unaffected and the site is already running again.
"We're sorry, but something went wrong."
Details much appreciated!
EDIT: Spoke too soon: https://support.cloudflare.com/entries/22054357-how-do-i-do-... shame it's so convoluted
"2500% guarantee This extended Service Level Agreement guarantees 100% uptime, and adds a multiplier to owed service credits resulting from any lapse: 5 times any downtime minutes and 5 times customers affected = 2500% guarantee."
http://www.youtube.com/watch?v=wMRaKtydILI
AS13335 = Cloudflare
They should host this page on a third party provider.
I wish there were a business model for a third party site-monitor and site-uptime service, which let the site owner do more than just post updates, but also prevented the site owner from lying about historical data.
Basically New Relic (that actually worked) + Pingdom + Internet Archive + Twitter + status.example.com.
reliable, accurate, high-frequency
"CNAME setup is a manual process generally available to paid CloudFlare plans only. If you are interested in testing CNAME setup, please contact CloudFlare first with the domain you would like to test CNAME with. Please specifically mention CNAME Setup in the subject field for faster review. Allowing for CNAME setup is entirely at the discretion of CloudFlare."
So NO: This isn't even a features at all. They made it as hard as possible to set this up and will grant you the use of it as they like.
EDIT: Based on Twitter search, all CloudFlare sites seem to be down.
Why I have switched off services? Because once i enabled them sites got slower, not faster. Sure - I doubt that most users noticed that, but for example, if i checked Pingdom or Google Crawl Stats - 'Time spent downloading a page' situation was very clear. With Cloudflare it took Google 2x more time to download page than without. I had no time to investigate why thats so, but after switching Clouflare off i was again back to 500ms.
Edit: Sites are not back. I guess that work the ones for which i have DNS cached. Lets wait...
Edit: Looks like DNS is back. However if you use CloudFlare services then you might still have problems. Like:
504 Gateway Time-out cloudflare-nginx
News flash: CDN fallback like the one below is next to useless unless the first request times out reasonably quickly.
http://css-tricks.com/snippets/jquery/fallback-for-cdn-hoste...
It's so funny how everybody jumps on top of new companies that say they can proxy all of the interwebz for a low price. (Cloudflare, Blackberry)
And then they fail...
If you're doing mostly static data which rarely changes, you can probably get very high hit rates on CloudFlare, and it's cheaper than even crappy $750/Gbps/mo colo bandwidth then.
(EDIT: I guess they switched; I haven't kept up with Rackspace)
CDN itself is essentially a commodity; it's not too hard to keep multiple CDNs in rotation. There are probably 20+ big CDNs worth consideration and another bunch of resellers. (Amazon CloudFront, BitGravity, Level3, Limelight are probably the first ones I'd think of for smaller sites; Akamai is still the undisputed king for top performance.)
DNS is the thing which is more interesting to me.
I'd probably go with Route53 for cheap good anycast DNS right now; everyone else seems to either be a clown or super expensive (or bundled with other expensive DNS service). Ultimately I guess I'll end up doing internal DNS. (non-anycast DNS is also a total commodity, but good anycast dns not as much) DynECT also looks pretty good. Not sure what other anycast DNS providers there are in the <$500/zone/mo range.
There are also many other CDNs that exist in the massive territory between CloudFlare and Akamai (such as CDNetworks, the company I had mentioned).
> CDN itself is essentially a commodity; it's not too hard to keep multiple CDNs in rotation.
For latency-insensitive use cases in generally centralized territory, I agree that CDNs are "essentially a commodity". The correct strategy would seem to be to call a number of them, and negotiate a good deal, not to assume that the one that has a printed sticker price is somehow the right choice (as some people here seem to have been doing ;P).
However, to make the counter-point to this: the cache hit ratio that is being reported by CloudFlare for evasi0n.com (note: I do not have control over that site's hosting; that's choice was due to planetbeing and pod2g) is 81% <- this is for a static single-page information site. How various CDNs handle caching, whether they cache you on disk or in RAM, what they do with regards to hot connections or pre-fetching... these all have massive performance implications on your website.
Punishing "old-school enterprise sales tactics" which try to keep price from being transparent is a reasonable choice. If you're a big content site, yes, you should go through the effort, but for someone who just wants a small service, buy from people who publish their prices.
CloudFlare isn't the only CDN which publishes pricing -- CloudFront with AWS is very transparent. Rackspace Cloudfiles is transparent. BitGravity is fairly transparent. Cachefly. etc.
Akamai is the worst at this, but Level3, CDNetworks, and Limelight don't publish pricing either.
Offering a free service like CF does is the genius of the freemium model -- even if your service is more expensive or less suitable at the high end, people who start out because it's free and easy will often stick with you as long as you're "good enough" as they grow.
Regardless, if picking up the phone and negotiating with a CDN, someone whose opinion of you is totally irrelevant and where the worst-case outcome is "we won't do business with you", how are you going to handle support on your own product, or court investors of your company?
Report a service problem? Welcome to the indian support-center, where you will be slowly spoon-fed the next best canned response that may or may not match your problem. Escalating to anyone with half a clue is near impossible.
Need a quote? Wait weeks at a minimum for someone to get back to you.
Oh, did we break your reporting? Of course this is your fault, Akamai doesn't make mistakes. Now talk to the hand.
If you're in the market for a CDN then there's plenty candidates who will sell you a better experience for less money (Level3, EdgeCast, BitGravity, etc.).
Meanwhile none of the other CDNs has given us such a terrible experience. It's almost as if they care, imagine that!
When you consider that the performance differences are at best marginal (the only real differences exist on mobile and in emerging markets, and Akamai is not on top of that) there's really no good reason to put up with the spoiled brats of Akamai in this day and age.
The same goes for EdgeCast, Mediatemple, GoGrid, etc are resellers.
Akamai's whole web-presents is clearly aimed at enterprise class customers. I mean there are whitepapers everywhere, the site is filled with business nonsense and I cannot even get started without at least a phone call or five.
Maybe people "jump on" to new companies because those companies actually offer a product old companies are not? I can go to Cloudflare right now and sign up, I can give them my CC info and pay, and everything is up-front and easy. I can be online with Cloudflare within the hour (or whenever DNS moves).
It's not that I can't spend 20 minutes on call - but if your customer acquisition costs involve at least 20 minutes of salesman-time per a possible lead (thus, at least 200 minutes of salesman-time for even the smallest sign-up), then I'm very sure that whatever number they will quote - that number will not be something that I can or want to afford. Let them pester some 'enterprises' with that.
They don't promise to cache much, and they in fact don't: even on very simple single-page static information websites, such as evasion.com, they have an abysmally poor 81% cache hit ratio. They don't help at all with dynamic content due to having poorly located nodes and lots of heavyweight code running in their proxy. Their lack of many nodes in good positions (compared to something like CDNetworks or Akamai, they are one or two orders of magnitude smaller) also means they can't provide very good latency even for the times when they actually happen to have something in cache.
(Note: if someone is now going to say CDNs don't generally do well with dynamic content, they are wrong: normal CDNs actually improve the performance of dynamic content incredibly by maintaining large-window pre-connected HTTP sessions to customer origin servers, often over private networks that already provide better bandwidth: you can easily see 2x latency improvements with a normal CDN even for fully dynamic content).
So, they really shouldn't be compared with a "CDN": they have an interesting service that actually provides something valuable for many key use cases (4chan comes to my mind: in essence, something that is actually likely to experience a true DDoS attack sufficiently often and with sufficiently little provocation that it makes sense to add an external system to your infrastructure), but if you need a "CDN" there are many more reasonable alternatives that don't have as many moving parts and are thereby going to break much less often (and, if they actually do, should break only in localized regions).
Also, randomly useful filters like adding the user's country to the request header, tweaking outbound images, and auto-injecting google Analytics.
Edit: 5:01 EST ..seems to be up and down according to mass pingdom messages.
I have no idea what is wrong, or how long they will take to fix it, but I'd imagine that CloudFlare has significantly better network engineers than the average company, and so they will fix it in far less time than the average company would fix the same problem.
Also, even if you are using them for free: they aren't replacing people you have in house... they are an additional component that can independently fail, in addition to any of the things that would have caused your average company's network engineers to fail. They don't promise to cache enough content to replace much of your infrastructure.
“While we have not completed our investigation, we believe this incident was triggered by a product issue that Juniper identified last October, when a patch was also made available"
Good network engineers tend to apply newly release patches. This vulnerability was documented for almost half a year...