Cloudflare was down
cloudflarestatus.com
cloudflarestatus.com
Lots of other big sites that are down: Patreon, npmjs, DigitalOcean, Coinbase, Zendesk, Medium, GitLab (502), Fiverr, Upwork, Udemy
Edit: 15 min later, looks like things are starting to come back up
The problem seems to have been resolved now, you might have made the change when they fixed it.
Cloudflare is also the authoritative DNS server for many services. If Cloudflare is down, then for those services Google's DNS has nowhere to get the authoritative answers from.
Again, it may not have been the only issue -- and there are a number of possible reasons why 1.0.0.1 wasn't reachable -- but it certainly was an issue.
Server: 8.8.8.8
Address: 8.8.8.8#53
** server can't find status.discord.com: SERVFAIL
So it's either not just cloudflare, or all those sites use cloudflare to host their DNS. $ dig +short discord.com ns
sima.ns.cloudflare.com.
gabe.ns.cloudflare.com.Their IP belongs to AS21581; which is registered with a company called 'M5 Computer hosting' out of west-coast USA.
m5hosting.com
The last hop is Santa Barbera.
Definitely does not fall in the AWS ranges.
Holy crap I am thinking either there is some magic or everything we are doing in the modern web are wrong.
Edit: The number is from Dang [1]
>These days around 5.5M page views daily and something like 5M unique readers a month, depending on how you try to count them.
Just need good stable code and server side caching.
Spin up an apache installation and see how many requests you can serve per second if you're just serving static files off of an SSD. It's a lot.
edit: I see that there are already a bunch of other comments to this effect. I think you're comment is really going to bring out the old timers, haha. From my perspective, the "modern web" is absolutely insane.
That is great ! :D
>It's a lot.
Well yes, but HN isn't really static though. Fairly Dynamics with Huge number of users and comments. But still, I think I need to rethink lots of assumption in terms of speed, scale and complexity.
I've been at my work laptop (not logged in) and found something I wanted to reply to, so I pulled out my phone and did so. For a good 10 seconds afterwards, I could refresh my phone and see my comment, but refresh the laptop and not see it.
Serving the same content several times in a row requires very few resources - remember, reads far outnumber writes, so even dynamic comment pages will be served many times in between changes. 5.5 million page views a day is only 64 views a second, which isn't that hard to serve.
As for the writes, as long as significant serialization is avoided, it is a non-issue.
(The vast majority of websites could easily be designed to be as efficient.)
Agreed.
I was brought up as a computer systems engineer... So, not a scientist, but I always worked with the basic premise of keep it simple. I've worked on projects where we built all the fangled clustering and master/slave (sorry to the PC crowd, but that's what it was called) stuff but never once needed it in practice. Our stuff could easily handle saturated gigabit networks as the 2 core cpu only running at 40%. We had cpu spare and could always add more network cards before we needed to split the server. It was less maintenance, for sure. It also had self healing so that some packets could be dropped if the client config allowed it, if the server decided it wanted to (but only ever did on the odd dodgey client connection)
That said, I was always impressed by the map-reduce of for search results (yes, I know they've moved on) which showed how massive systems can be fast too. It seemed that the rest of the world wanted to become like Google, and the complexity grew for the std software shop, when it didn't need to imho.
I jumped ship at that point and went embedded, which was a whole lot more fun for me.
Sincerely, old timer
You know, it should be even better than it was in the past, because a lot of heavy lifting is now done on the client. If we properly optimized our stuff, we could potentially request tiny pieces of information from servers, as opposed to rendering the whole thing.
Kinda like native apps can do(if the backend protocols are not too bloated)
We were serving around that traffic off a single dual pentium 3 in 2002 quite happily off IIS/SQL Server/ASP. The amount of information presented has not grown either.
That little box had some top tier brand main corporate web sites on it too and was pumping out 30-40 requests a second peak. There was no CDN.
Not a one trick marketoid pony.
90 to 99% of those are logged-out users, so fully cacheable.
Only a handful of dynamic requests each second remain.
It's ergonomic in a very lispy way but perfectly reasonably so from the POV of that aesthetic.
Even a heavyweight and badly written web server can hit 100 QPS per core, and cores are a dime a dozen these days, and storage devices that can hit a million ops per second don't cost anything anymore, either.
HN is written in a lisp variant and most the stack is built in-house , it is not difficult to imagine efficiency improvements when many abstraction layers have been removed from your stack .
What I do remember, is that it was a social data collection experiment for a couple of social scientists, that never originally expected that many people would actually find ways to find each other and hook up using it.
I miss their old findings reports about how weird humans are and what they lie about. Now, it's just plain boring with no insights released to the public.
Nick carver from SO also one mentioned that they could run SO if a single server , while it wasn’t fun it was doable and had happened some time .
I don't know what the server's specs are but I'm sure it must be quite beefy and have quite a few cores, so let's say that it runs about 10 billions instructions per second. That means a budget of about one million instructions per page load in this pessimistic estimate.
The original PlayStation's CPU ran at 33MHz and most games ran at 30fps, so about 1million cycles per fully rendered frame. The CPU was also MIPS and had 4KiB of cache, so it did a lot less with a single cycle than a modern server would. Meanwhile the HN servers has the same instruction budget to generate some HTML (most of which can be cached) and send it to the client.
A middle of the line modern desktop CPU can nowadays emulate the entire PlayStation console on a single core in real time, CPU, GPU and everything else, without even breaking a sweat.
>Holy crap I am thinking either there is some magic or everything we are doing in the modern web are wrong.
Magic, clearly.
It doesn't need to be crazy.
A static site on DigitalOcean's $5 / month plan using nginx will happily serve that type of traffic.
The author of https://gorails.com hosts his entire Rails video platform app on a $20 / month DO server. The average CPU load is 2% and half the memory on the server is free.
The idea that you need some globally redundant Kubernetes cluster with auto fail-over capabilities seems to be popular but in practice it's totally not necessary in so many cases. This outage is also an unfortunate reminder that you can have the fanciest infrastructure set up ever and you're still going down due to DNS.
True, but this is why it shouldn't be bashed either. When you need it, you need it (cue very complex enterprise applications with SLA requirements).
To support this, look at how many people criticize Kubernetes as being too focused on what huge companies need instead of what their small company needs. Kubernetes still has its place, but some peoples expectations may be misplaced.
For a side project, or anything low traffic with low reliability requirements, a simple VPS or share hosting suffices. Wordpress and PHP are still massively popular despite React and Node.js existing. Someone who runs a site off of shared hosting with Wordpress can have a very different vision about what their business/sideproject/etc will accomplish compared to someone who writes a custom application with a "modern" stack.
It's easily served by a simple server.
i used to host a wordpress site that has 5M pageviews a month on a $10 (and later $20) digitalocean instance.
that's wordpress and a shared vps. I imagine it could be a lot higher if I have dedicated server and use self-written software.
I've been working on something where the DB is also part of the application layer. The performance you can get on one machine is insane, since you spend minimal time on marshalling structures and moving things around.
They are still using Cloudflare. Unlike CF, M5 does not require SNI.
curl --resolve news.ycombinator.com:443:104.20.43.44 https://news.ycombinator.comI host a dedicated server there (running https://www.circuitlab.com/) and when I traceroute/ping news.ycombinator.com, it's two hops (and 0.175 ms) away :)
I even checked to see if an AWS region was down once I realised it wasn't on my side (I thought it might have been my ISP's DNS servers or something).
The next move was to check Hacker News - thankfully it's not also hosted on Cloudflare, ha!
(Visit at your own risk.)
Hack?
If I type 'e' I get 'en.wikipedia.org'.
I was redirected.
That is why if you have this question, you should go to google.com
My guess is that there are more resources invested in making sure google.com stays up than for any other site on the internet.
My stuff is on Netlify (for the next week or so) and the rest is on a VPS bought from a local business who isn't reselling cloud resources. I'm kinda glad I moved all my stuff from cloudflare.
Ah, well. This too shall pass.
It will also be especially noticeable to end-users, because sites using Cloudflare are typically high-traffic sites, and so a 'minor' issue that affects only a handful of sites is still going to be noticed by millions of people.
I loathe Discord, and I can barely contain myself with schadenfreude at this news.
On my home network I use Google as a backup DNS provider so the whole internet didn't go dark for me, but I don't have a backup DNS host for my company's DNS records.
This is a very common pattern and falls into the 'nobody got fired for buying cisco/microsoft/intel' trap.
I have two issues with it;
1) It entrenches the largest provider.
You would not extend the same leniency's of outages to the third best cloud provider, this means that people will just keep pushing the monopoly forward. Even if the uptime or service is actually better on another provider.
2) You create a tight coupling of monocultures;
Simply put: You slowly erode the internet. Your site becomes an application in a distributed mainframe operated by a tiny minority of tech companies universally based in the US.
Why is this a problem? I could give moral answers here but I thing pragmatic ones are more convincing..
Giving ownership of the internet to the few gives them the ability to set the rules.
If you're on Amazon's AWS, what's to say they don't inspect your e-commerce systems and incorporate your business logic into their amazon.com shopping experience. They do this to their marketplace and create competing products already[0].
If you're doing really well, why not just drop a few packets here and there? I mean, they wont.. you're paying, right?
Hell, if you do super well they can just change the rules and make it so your services get expensive in the exact way you use them, or even legislate you off the platform entirely.
Probably wont happen, but it's a lot of trust you have to admit, and people shit on Apple for having that kind of power, and Apple is not even competing in the same market as most people on this site.
If you're on google cloud (which, I'm a fan of btw) and you feed an ML model, well, you paid for it but why shouldn't they also have a copy.. after all, it's green to do so! Their bigdata platform? Google loves data! Feed the beast.
[0]: https://fortune.com/2016/04/20/amazon-copies-merchants/
Amazon can copy your business model just fine without looking at your servers. Most of what’s on your servers is probably irrelevant from their perspective.
We REALLY need a truly decentralized, distributed DNS system that is not owned by private entities.
https://ieeexplore.ieee.org/document/7530014/authors#authors
> Unlike previous DNS replacement proposals, D 3 NS is reverse compatible with DNS and allows for incremental implementation within the current system.
You can also also register your domain on multiple TLDs.
The original idea was that with the barrier to entry being so low, anyone and everyone could set up their own websites, mail servers, etc.
But with it being so easy to compare and contrast service (i.e. the market being so open), it means that the competitive forces naturally consolidate to a winner-take-all model. If when starting out Cloudflare was just 5% better than the competition, it could have easily taken the vast majority of the mindshare on the internet. Couple that with the fact that there are huge advantages with scale to a business like Cloudflare's, and it's not hard to see how so much of the internet has become dependent on it.
I know a lot of big companies do, but I am always surprised when you see ones that don't.
What's the point of a status page if it doesn't reflect the real status...
It's either the status page goes down with everything else or the status page is wrong. Great.
EDIT: Looks like it's accurate now, 20 minutes later.
(To give an update, I'm seeing from my monitoring systems (about 15 points around the globe) sporadic outages for Microsoft, Apple, Reddit, Bing, Node.js, Twitter, Yahoo, and YouTube. And my own servers (not behind CF at all) are also flipping up and down. It started around 21:14 UTC.)
And changing DNS servers often takes many hours (or days, if .net is involved apparently)
The crazier thing is that I tried to login to our CloudFlare account, it never sent me the 2FA code... I still haven't been able to login (Enterprise account)
We had a surge of people checking if Discord was down on our site, then I noticed everything went down shortly after. Discord is still the top check right now.
I can't ever remember hitting these kind of traffic numbers before.
And I'd like to be able to have our site communicate outages like this Cloudflare one, where more than one site might be affected by a larger provider. Automating that is difficult.
This is still a side project, though, so I mostly work on it when I get the urge :)
How comfortable are you with open source? If you were willing to release your stack on Gitlab/Github, it might be worth your while.
Maybe the problem was somewhere on the AMS-IX
I actually have a WAN2 configured but not plugged in and it was set to "Load Balancing: Failover Only" ... I wonder if all of my 'connection issues' were software assuming my network link is down and switching interfaces to an unavailable one.
root@USG-PRO-4:~# show load-balance watchdog
Group wan_failover
eth2
status: Running
pings: 2
fails: 0
run fails: 0/3
route drops: 0
ping gateway: ping.ubnt.com - REACHABLE
eth3
status: Waiting on recovery (0/3)
failover-only mode
pings: 1
fails: 1
run fails: 3/3
route drops: 1
ping gateway: ping.ubnt.com - DOWN
last route drop : Fri Jul 17 17:32:58 2020I definitely agree in concept with you, but then i think back to how frequently script kiddies took down sites ~10 years ago, or w/e. I feel like what has changes is the massive CDNs in front of so many sites.
So while i do want a better solution, i'm not sure what it looks like. Thoughts?
Reddit/HN/etc will send all users to the same URL. Almost all of those users will come without any pre-existing cookies. Serving the same content to all those users should not be impossible for most sites without CF or a CDN.
One question is how to do DDoS protection without somebody like Cloudflare. Some new protocol for edge caching, perhaps?
a) complexity: trick your servers into doing something hard
b) volumetric: overwelm your servers with a lot of traffic
c) volumetric part two: overwelm your servers with a lot of requests, so you respond with a lot of traffic
A and C are things you can work on your self --- try to limit the amount of work your server does in response to requests, and/or make resource consuming responses require resource consuming requests; and monitor and fix hotspots as they're found.
B is tricky, there's two ways to solve volumentric attack; either have enough bandwidth to drop the packets on your end, or convince the other end to drop the packets (usually called null routing). Null routes work great, but usually drop all packets to a particular destination IP, which means you need to move your service to another IP if you want it to stay online; that's hard to do if your IP needs to stay fixed for a meaningful time (TTL for glue records at TLDs is usually at least a day); and IP space is limited, so if your attackers are quick at moving attacks, you could run out of IPs to use. Some attacks are going above 1 Tbps though, so that's a lot of bandwidth if you need to accept and drop; and of course, the more bandwidth people get so they can weather attacks, the more bandwidth that can be used to attack others if it's not well secured.
That sort of thing.
That could be the slogan for 2020
Outages happen no matter what the infrastructure is. There's no solution, they're just something you need to recognize and handle, which Cloudflare seemingly did relatively quickly here.
Level 3 or Telia going offline is perfectly survivable for any customer who has multiple upstreams.
It may perhaps be an exaggeration to say that there are not other providers that are similarly critical for a significant percentage of the internet.
> Appears that the router in Atlanta announced bad routes (effectively a route leak). Only impacted our backbone. Not all of our PoPs are connected to our backbone, so some would not have seen an issue. Appears to have impacted about 50% of our traffic for a bit over 20 min.
Only last in my mental list was the possibility that Cloudflare would be down.
Hope they publish a detailed post-mortem. It's always fun to read (but certainly very painful for those directly involved in writing it).
Interestingly the same domains don't show up on google's (8.8.8.8) DNS at all.
Don't use Namecheap and Cloudflare at the same time.
Namecheap is using cloudflare. So if cloudflare is down, you can't change DNS settings on Namecheap as well!
Seems cloudflare took out a good chunk of the internet temporarily.
Doesn’t HN use cloudflare? Why did it survive? (I haven’t looked for about a year, but I seem to remember HN being proxied behind CF at one point.)
HN is so reliable that’s it’s almost never needed one. I’m extremely curious how HN survived this; almost positive they used cloudflare at one point.
I think the official status page is @hnstatus on Twitter, or something like that.
; <<>> DiG 9.10.6 <<>> discordapp.com ;; global options: +cmd ;; Got answer: ;; ->>HEADER<<- opcode: QUERY, status: NOERROR, id: 8092 ;; flags: qr rd ra; QUERY: 1, ANSWER: 5, AUTHORITY: 0, ADDITIONAL: 1
;; OPT PSEUDOSECTION: ; EDNS: version: 0, flags:; udp: 1232 ;; QUESTION SECTION: ;discordapp.com. IN A
;; ANSWER SECTION: discordapp.com. 140 IN A 162.159.135.233 discordapp.com. 140 IN A 162.159.129.233 discordapp.com. 140 IN A 162.159.130.233 discordapp.com. 140 IN A 162.159.134.233 discordapp.com. 140 IN A 162.159.133.233
;; Query time: 69 msec ;; SERVER: 1.1.1.1#53(1.1.1.1) ;; WHEN: Fri Jul 17 14:37:40 PDT 2020 ;; MSG SIZE rcvd: 137
I noticed a lot of packet loss to 1.1.1.1, not an outright "outage", maybe they were rolling a deployment?
Edit: Looks like a deployment to me (looking at the logs I could see cascading traces, so it took down one DC and the other started responding - increased latency - and then down, etc..), gonna be an interesting post-mortem!
https://www.digitalattackmap.com/#anim=1&color=0&country=ALL...
DNS is completely broken.
"All systems operational" in nice soothing green.
No, not so much.
Time to go back to the drawing board, for a lot of us, to re-assess points of failure.
Edit: many websites are failing to DNS resolve but the services they provide continue to function fine behind the curtain.
Half the damn internet is not currently resolving.
This allows my internet to "work" during this time, but adds about 1s latency to resolutions. Presumably that's the time it takes my internal DNS resolver to try the secondary.
https://status.digitalocean.com/incidents/6wtmldty17g1
As big as this is, any chance a major hub/backbone went down?
Right now, I can't get to my own website (hosted on DigitalOcean, not through Cloudflare), but Oh Dear claims it's up. So I suspect that the problem is closer to me than it is to DigitalOcean (or Cloudflare).
Hidden dependency revealed.
Ancedotally the interenet seems faster.
Dang, I'm pretty disappointed in CF. I've never experienced this much an effectful DNS outage.
I wonder if that includes the roots that Cloudflare operate.
Changing back to an alternative (such as 8.8.8.8 from google) restored my access to the areas of the internet not using Cloudflare.
I should probably run a blend of 1.1.1.1 with 8.8.8.8 instead.
Shameful that so much of our decentralised web is so centralised and breakable in one place.
at least i am glad hn exists, it is the only thing that loads everywhere
Looks like DNS issues, their nameservers aren't reachable.
Basically all Riot Games (League, Valorant, TFT) are down, dunno about LoR.
This affects so many things it's scary, and Cloudflare status page has still not updated. HN got there first.
Actually the site is running but not accessible due to the issue. Glad you got the heads up, after a while I had to pause monitoring to prevent side effects.
All my DigitalOcean instances are down.
to withstand a nuclear war
brought down by cloud flares
isitdownrightnow.com down