AWS Best Practices for DDoS Resiliency [pdf]
d0.awsstatic.com
d0.awsstatic.com
I'm still a believer in the value of dial-up, leased lines, satellite, or radio for aiding security. You still have to apply protection to them but don't have whole Internet coming after you with protocols that aid attackers more than defenders. My method is typically to obfuscate identifiers for Internet services and use methods like authentication at packet level (eg port-knocking or VPN). The configuration details are sent over the non-Internet medium. Even dial-up can move basic credentials and some I.P. addresses quickly. Don't need to do it often, either. If you hide it (eg SILENTKNOCK), attackers start getting pretty pissed and desperate wondering why not a single packet gets through.
This method is primarily for intranet sites, though. Web sites or apps facing the public naturally are at high risk. Best to just use Cloudfare or a similar service along with hiring good security folks.
Trying to outscale a large DDOS doesn't often work. Don't worry though, amazon's happy to help let you try to pay for it!
The client's app almost always goes down in the midst of the fury. The point of failure typically comes down to amazon's load balancers or an auto-scaling failure. In the end our friends at amazon tell us to add more, bigger servers to 'outscale' the traffic and put the blame on us when everything blows up. sigh
Attempting to outscale a DDOS (the primary mitigation method presented by Amazon) is going to DDOS your bank account. Personally, I'd rather see some more recommendations along the lines of the "VPC can minimize potential attack surfaces".
Remotely triggered black holes for VPC? Elastic Firewall?
Not crazy about firewalls in general, but they would help in the case that you are paying for data-out.
So, maybe it's what's worked for them, their thinking hasn't really changed, and now they're just offering others the same thing? And upselling them in the process? Thoughts?
[1] http://money.cnn.com/2010/12/09/technology/amazon_wikileaks_...
https://support.cloudflare.com/hc/en-us/articles/200168916-C...
"the system currently does not have the functionality to automatically select the next available server if one of the servers in the group goes down"
The only big downside is that on AWS you can't have an elastic IP associated with an elastic load balancer, so you either have to run your own HA haproxy/nginx/whatever cluster in EC2 in order to have a single IP to point CloudFlare to.
If you can live with a subdomain you can point that cname to an ELB.
Alternatively, CloudFlare's API is pretty reasonable, so you could home-brew health checks that de-register dead nodes from CloudFlare. Even a simple nagios check handler could do that.
https://support.cloudflare.com/hc/en-us/articles/200169056-C...
There's no way to ensure the rest of the internet will handle it correctly though with all the proxies and DNS caches in the middle and low TTLs can also add latency to end-users who might have to constantly do a DNS lookup on new connections.
If you're using CloudFlare's full service (instead of just DNS), then it'll be seamless because their IPs don't change.
[1] https://blog.cloudflare.com/introducing-cname-flattening-rfc...
Edit: Considering that Jeff from BlackLotus is now PM of DDoS at AWS, I'm sure they are working on something.
https://twitter.com/olesovhcom/status/386563685805617152/pho...
https://www.ovh.com/ca/en/anti-ddos/ddos-attack-management.x...
Also everyone here talk about L7, cloudflare ect.. but a lot of application are pure TCP/UDP based so you can't cache anything.
Most dedicated server providers won't go that far if you have a handful of servers.
I know from personal experience that, Digital Ocean, the largest VPS provider null routes your VPS IP for 3 hours minimium for even the tiniest of DDoS's.
I doubt most of the smaller VPS providers can afford to absorb DDoS's even if they don't have overly restrictive policies like DO.
buyvm.net for 3$ per month (100Gbit apparently)
iwstack.com, 8Gbit protection for free
ramnode.com, 20Gbit
Most of AWS advices (like autoscaling) will help only a bit, but can cost a lot (lots of ec2 machines serving bogus requests).
This helps in 99% of cases, and where it doesn't it is simply because there is a resource that cannot be cached and that the edge must revisit the origin for. This is especially true whenever that resource is expensive for the origin to provide (involves database lookups and cannot be cached: shopping carts, login pages, search results), these are the ones which require you to rethink your application design.
If you're an application developer and wondering how to design your application to withstand a DDoS attack, then instead shift to just thinking: How can I make everything that this application does be cached by an edge server?
When you're not under attack using CloudFlare makes sense and saves you money anyway. At least... it does for me. On one of my web applications I use Amazon S3 for user attachment storage within a forum CMS, and my bill used to be upwards of $200 per month for just one of the sites I run. I changed the application so that it proxies the S3 request/response, and then set a CloudFlare Page Rule to sit in front of that path, and configured it to "Cache Everything". The effect of this was to reduce my AWS S3 bill down to $20 per month. After that I did it for every site.
There's a hell of a lot of benefit to using CloudFlare in conjunction with AWS, and not just when you're facing an L7 DDoS.
Disclosure: I work for CloudFlare (last 9 months) and have been a CloudFlare customer for 3 years and I was offered a job by AWS and also been an AWS customer for 3 years.
I can't imagine trying to survive a volumetric attack in AWS. Must be a nightmare. Luckily volumetric attacks are on the out and layer 7 attacks are all the rage these days. They're easier to handle in AWS with a WAF or filter.
tl;dr keep the TTLs on your DNS A records to a maximum of 10 minutes.
What if you can't absorb the cost that is attached with scaling ?
You need to add some additional logic to smooth out the rate of scaling. Most deployments fall down when the rate of scaling can't keep up with the demand.
In general the white paper provides some solid AWS-specific & AWS-centric guidance on how to buy yourself some time. It's not the end-all, be-all but a good start
They're not saying "scale up and just pay for it", they're saying use autoscaling as a tool to give you time to respond, Without first going down.