Dyn Analysis Summary of Friday October 21 Attack
hub.dyn.com
hub.dyn.com
* us-east-1.amazonaws.com: split between internal, UltraDNS, DYN
* spotify.com: all internal nameservers now
* reddit.com: all Route53 now
* github.com: all Route53 now
* netflix.com: all Route53 now
* paypal.com: split between UltraDNS and DYN
No changes made:
* twitter.com: 100% with DYN
I don't think Dyn showed any incompetence; the parent-poster was merely remarking on relying entirely on a single provider, who, if they get DDoS'd, causes your site goes down. (There was some previous discussion about splitting between providers, but some commenters noted that it was difficult, or at least non-trivial, to replicate records between two providers.)
Lucky for Route53 users, Route53 DNS surface is really large and there is a really good chance that not even is attack could hurt it.
Aside: I came out of college as a sys admin with a CS degree and writing tools like this was par for the course. If devops folks aren't writing tools like this today, what are they doing?
At least that's my understanding anyway.
We are currently evaluating the Netflix denominator tools to spread DNS and sync our alternate providers.
My biggest problem during this outage was that I could not login to my registrar and make change to DNS directly - I had to login to DYN and ADD Route 53, it was impossible to remove DYN completely. And that's how we landed up with a split view.
NOW if anyone can tell me of a competitor for the Traffic Director product that works on Port 25 I'll be happy to consider a migration. Cloudflare has something in the works, but I’d really just like a DNS provider with a virtual load balancer that can handle my 250qps at a reasonable price.
There is an RFC to pass information about the client subnet: https://tools.ietf.org/html/draft-ietf-dnsop-edns-client-sub...
Which Google DNS uses to tell your name servers more information about where the client is located. This allows you then direct them to the nearest server.
Shame on reddit, github, and netflix for learning literally fuck-all from this.
https://status.heroku.com/incidents/965
"This outage exposed a critical weakness in our DNS hosting configuration. We are taking immediate steps to add additional DNS providers. This should allow us to avoid impact in the future, provided that at least one of our DNS providers is operational."
I don't know much about running nameservers but moving to all internally hosted seems like an odd choice to me, can anyone explain whey that's a good move?
The question you should ask is why did these companies used an external DNS in the first place?
You can't outsource DNS. It's one of the critical piece of networking that must be in every infrastructure.
The common DNS server is BIND. It's been there for 30 years, it's well known, well manageable and well understood. Sysadmins have to know it and manage it. It's especially critical for worldwide multi-site tech organizations.
There is no need for anything else. BIND can do everything and is the most flexible. Some of the alternatives lack some or most of the features (e.g. some type of DNS records).
You should assume that any organization is running it's own DNS servers. (ignore the edge cases).
---
In practise for large scale operations, the DNS tree will get very complex.
What the websites changed was only the public DNS server for reddit.com or airbnb.com. It's only the top of the iceberg. There is likely a very complex DNS setup underneath including public domains, private domains, special internal domains, CDN, per datacenter, per continent, etc... which could imply 10 different DNS services.
Who serves the top level public domain is a details. We should assume that the companies put whatever they could in little time to fix the ongoing issue.
This is simply not true. For resolvers, you can use your ISPs DNS servers or use a public resolver like Google DNS, OpenDNS, etc. For authoritative DNS there are plenty of hosted (outsourced) offerings like Route53, Dyn, Google Cloud DNS, etc.
This may not work for sufficiently complex organizations, but in my ~20 person SaaS company we have zero DNS servers and it works just fine. We use our ISP's resolvers for client lookups, and Google Cloud DNS for authoritative DNS.
Thing is. You gotta to run your own DNS since the moment you want your own DNS names. Good for you if a simple external DNS service is enough for you, a single 20 people office is not comparable to what the websites mentioned are operating.
isn't that basically the definition of DoS?
If someone is targeting them directly it doesn't matter much that DNS is up and running, their site is still down.
Spotify simultaneously has large resources and offers a non-essential infrastructure service (music to listen to while you're doing something else). The V gained in DoSing them is very small. They got attacked anyway because they shared infrastructure with other companies, which pools the V together to create something much larger. Some attacker saw a case where V >> X and attacked it to great success until Dyn was able to bring up X again. During the interim, Spotify was down despite having V << X.
In short: Spotify probably can't do DNS better than Dyn, but they can do DNS better than the sort of people who have reason to attack them (presumably trolls, maybe some future hacktivist who doesn't like some business decisions they make, unscrupulous competitors). This attack was a wake-up call for them, "oh, if we're pooling with these other folks then we'll become targets of larger hacktivist attacks and state actors, who are not directly targeting us per se." Those attackers could presumably still take out Spotify's home-rolled DNS, but they have no real motivation to target Spotify in particular any more.
(And I know it's not that simple, but that's probably the basic reasoning behind it.)
Email would still work. You can't receive email if the sending server can't look up your MX records. Since spotify.com uses Google Apps, their email would survive a total network outage if they used third-party DNS.
On previous 3 Saturdays they lost between 40 and 60 domains.
Indeed, if you want the HTTP/HTTPS traffic to go through Cloudflare, the DNS must go through Cloudflare. There are generally two ways to set it up:
a) You move your DNS auth to Cloudflare and allow it to manage it.
b) You keep managing your domain yourself, and CNAME to Cloudflare. See: https://support.cloudflare.com/hc/en-us/articles/200168706-H...
What you should do depends on your setup and threat model. Do you fear DNS auth going down? Do you think your DNS will be a target? Do you use Cloudflare to hide your HTTP origin IP addresses?
For example, if you fear DNS auth going down, but you must use Cloudflare for HTTPS (say: for caching and SSL certs), then changing DNS off CF makes little sense. You already assume stability by expecting it to work HTTP layer.
If you think you can be a target of DNS attack, I'd say having multiple auth is unlikely to give you more mileage.
If you can afford disabling CF on HTTP layer, exposing your HTTP origin IP and want to have two different DNS auth providers, fine, you can do CNAME. But then you have three vendors to worry about, and problems with each can lead to trouble.
It comes a bit as gloating in the face of the attack on Dyn and there's no reason to believe that Cloudflare's DNS would fare any better.
Answering DNS is not very costly, so if you have enough capacity to the servers, answering shouldn't be the bottleneck.
I agree that it's very bold to do that, but I'd trust them with handling DDOS more than most other providers.
$ dig +short ns nflxvideo.net
ns1.p19.dynect.net.
ns4.p19.dynect.net.
ns2.p19.dynect.net.
ns3.p19.dynect.net.Virtually everyone is behind NAT these days, often multiple layers of NAT. So how does the botnet manage to telnet or ssh into these set top boxes or lightbulbs or whatever? When I want to SSH into my home computer I have to go through elaborate maneuvers to get it to work.
So are you saying that these IoT device makers not only hard-coded root usernames and passwords into their devices, but deliberately set up UPnP mappings to those ports?
That looks malicious rather than negligent. Am I missing something?
(Well, not surprised as much as disappointed.)
One particular method is DNS rebinding. They attack your webbrowser, then your browser passes the attack to the device inside the network.
Another means is 'the weakest link'. An insecure and compromised device (even if it just a user account on that device) scans your local network behind the firewall , passes information to a command and control server, which downloads and passes an exploits to the devices it discovers on your network.
The characteristics for how these surveillance devices were hacked are that the devices are using older firmware with an exploitable telnet feature and that the default credentials are still intact. There are a few vectors (using UPnP, HTTP API, directory traversal) that can be applied to bypass authorization. Throw in a global directory of these devices (shodan.io) and you have the ability to search for these devices with specific firmwares, run every attack vector to compromise these systems, and, once connected, have them do _whatever_you_want_.
And that's only on the default port.
Infected windows desktop can scan local network and infect IoT devices.
Expect more of this in the near future, single-source infrastructure is becoming a huge liability (not that it wasn't before). I wonder what impact on SLAs it will have when cloud services providers are taken down - will they honor their SLAs or inject DDOS clauses into them to shield themselves. You won't see many standing up to multi-Tbps attacks, at least for the moment.
Myself and a lot of friends had servers killed both when renting VPS/Dedicated server or a dedicated "Game Server".
And all and all with considerably smaller botnets like the ones you rent for a few $ per hour.
If you are running a public server you learn quite quickly that if you permaban a cheater or just some annoying kid you should expect to be DDoSed these days.
Am I missing something here. It wasn't an L7 attack (or was it?) Why keep referring to it as complex?
There also isn't a lot of details here on the exact nature of the traffic. They say it was hard to distinguish between legitimate traffic and this malicious traffic. So the botnet is at least rotating their requests through lists of customers hosted with them (though that isn't complex, but it is forward thinking. If the botnet was all making non-stop requests for just a few domains, that would be a strong signal to start filtering traffic, first internally, then pushing ISPs to block it upstream).
What's described in this incident report is totally within the capabilities of a single individual with public knowledge, though. If they could have proven otherwise, they probably would have (unless that somehow conflicted with their criminal investigation).
1. Device backdoor open Port 23 (telnet), used to take over loT devices.
2. The loT devices attacked through Port 53 (DNS).
DNS is normally a low-bandwidth protocol so if you only provide DNS services, needing to purchase 1000x your normal bandwidth to handle these bursts would be miserable. If a DNS provider were also providing video services (Vimeo/Twitch/etc), then a 1.2Tbps increase in traffic could be easily absorbed.
[p.s.] The question is precisely this: is it possible that a DDoS attack on DNS can be used to affect/mask a MITM attack.
I doesn't have to be. An attack can be highly asymmetric such amplification/reflection attack. I am not saying this is what happened in the Dyn case but rather addressing your average comment. Just an example having hosts in the botnet send queries to non-Dyn open resolvers on the internet and in those queries they spoof the source IPs of Dyn DNS servers. Now all of those open resolvers on internet start sending return traffic to Dyn. If in those spoof request they request EDNS0 those responses could be up to 4K. Compare that to the size of query packet and you have a lot leverage.
Don't think they said they solved the problem. They just made it less bad. Some DNS queries got through, but not all.
I'd think at some point if the attacker isn't getting the same results from the attack he'll wait for them to scale back then strike again.
Is my thinking right here?
Not necessarily. They could configure the C&C to run off of something like pastebin/gist where they just input a list of IPs to attack and the botnet checks a specific url every X minutes for new instructions.
The average broadband user has at the least, 10Mbps/UP. If your smart appliances all started sending at least 10mbps up... you only need a million smart TV's to start causing damage.
I'm guessing there's 100M Smart TV's in the USA? Each sending 10Mb/s up? There's 1GB/s of traffic. Multiply this by the next 50 nations who have smart TVs and bandwidth to spare...
Then make it multiuple devices in a home. Then make it every smart device on the planet, using SOME bandwidth... it gets painful fast, and its free for the attackers, since the poor sap with a hacked smart TV is doing the work.
Unless they are using a setup like @tomschlick mentioned or some P2P thing.
For an attack of 1.2Tbps, you only need 120k devices at 10Mbps.
Makes me wonder when we'll start seeing reports of people getting hit with data overage charges because of hacked smart devices. I only have 4 Mbps up, but if an infected device used all of that 24/7 it would chew through my Comcast 1TB monthly cap.
I think the font-weight of 300 is the bigger culprit here. If you pop open the Chrome dev tools and change font-weight to normal, everything becomes much easier to read.
For Chrome there is a "High Contrast" add-in to normalize shiat font color selection.
Interesting. Much of the reporting on the day of suggested that the attack was felt exclusively in the United States, but this says otherwise.