Mitigating DoS Attacks with Nginx
nginx.com
nginx.com
Also the DDoS attacks that we've been hit with actually target our uplinks by saturating them with traffic, not our services. We have a 1 Gbps port and the last DDoS we were hit with was over 20 Gbps, which is a relatively small one. The mitigation we used was to have our hosting facility get their upstream provider to route the traffic through a layer 7 DDoS mitigation filter provided by an external company. It worked wonderfully.
These are cool features, but when your link is saturated it doesn't matter what a daemon listening on a port does.
That'll be a while, of course, but already we see attackers with access to a tremendous number of unique IP addresses in the IPV4 space.. they'll have many orders of magnitude more soon.
... in the same /64 range for the most part, so as easy to block/filter/limit as one IPv4 address.
You risk inconveniencing people who are assigned just a few addresses because you potentially end up blocking many of them due to the actions of a few on the same subnet, but you can't be held responsible for hosts/ISPs doing IPv6 wrong.
Possibly even easier, because of IPv4 deaggregation.
(Because of IPv4 address scarcity, many providers have discontinous IPv4 address space. This is mostly a problem for the core, because it leads to much larger BGP routing tables.)
Not to mention that some ISPs do carrier-grade NAT specifically due to the limitations of IPv4, so blocking a single IP(v4) might affect multiple people as well.
When you think about IPv6, start with /64 networks. That's the basic unit. What a DSL customer gets from the ISP is a /64 network, not some number of individual addresses. The customer may two, ten or 2047 of then, it doesn't matter. The point is "one owned DSL subscriber = one /64 network".
Just like with IPv4, some people have larger allocations. But the basic allocation unit is a /64 network.
location = / {
if ($http_user_agent ~* foo|bar) {
# return non-standard (NGINX only) 444 code
# closes the connection without sending a response header
return 444;
}
}This will get you nothing in a typical DDoS scenario, but thanks for sharing as it may come handy for other situations.
To mitigate a DDOS, you need to go upstream to your network providers and filter out the traffic before it reaches you.
I haven't read this all the way through, but on a cursory glance it looks reasonable:
I'm starting to think that we need some agreement where instead of logs, we just get apps to emit a stream of protocol buffers and a format string for the messages and data.
Which does make me wonder if you couldn't LD_PRELOAD something which replaced fprintf and the like...
But often times you don't even know what domain is being connected to at the network layer. You need output from the process holding the connection key. And you want very clear separation from that task...
"So much" is pretty imprecise. How much waste do you believe string parsing incurs in this case?
For example, systemd finally provides a logging system that allows structured logging with key/value fields.
I'm not opposed to other ways of solving it, but I think the belief about how much resources it takes to process a text log file that people are expressing in this thread is at least a few orders of magnitude off.
This blog post is a good starting point for the kinds of strategies you need to fill that gap in protection.
GA would cost you more than 50ms too. More so than a CDN controlled analytics. But obviously that cost with CDN is an upfront latency rather than the more hidden cost with background loading of GA. So arguably GA's cost is less "bad" than the CDN's cost.
Personally speaking, I prefer the CDN approach as it produces web pages with a lower browser footprint which I think does improve the user experience (though I'm not implying that GA give a bad user experience!).
GA does give a greater breadth of information than CDN analytics though. Often that's the real deal breaker since analytics is usually driven by project managers / clients rather than by the developers.
> Charts that show how much static traffic you saved are nice, but with bandwidth close to free, it's not a big deal.
Oh it's definitely a big deal if you serve high traffic websites ;) I've spent hours working against those kind of reports on projects that were seeing 100k concurrent users. I will say that these graphs aren't so much about judging what bandwidth can be saved but more about judging what requests can be offloaded. The idea being the fewer calls to your origin servers you need to make, the more resources you have available in your farm for generating the dynamic content (dynamic content you cannot cache!). This also has the potential to save you money in server costs (depending on how they're licenced) as well as improving site performance at peak times.
> CDN analytics need to be better than GA at which point I will not only trade off latency but convert to premium all the way.
Indeed. GA will likely always be better from an account management perspective. But as a devops engineer, CDN analytics fulfils my needs. The great thing is that we have a multitude of options we have available :)
In all seriousness though, it might help to look beyond the very specific set up of your present company when asking about why other people opt for other CDN services. But for what it's worth, I've not experiences the same degree of latency issues with either Cloudflare nor Akamai that you've described. And I have done extensive load tests.
I remember listening to your talk at dotGo 2014 :-)
I tried CloudFlare in November 2012 (3 years ago, and not 1 year ago as I wrote in my previous comment). At that time, the origin server was hosted by Typhon in France. I remember that after having enabled CloudFlare, the latency was significantly increased. I haven't kept the specific timings, but to give you an idea, the response time was like 100 ms without CloudFlare and 500 ms through CloudFlare.
That said, it was a long time ago and I can guess things have changed a lot since. So I did a new test today. The origin server is hosted by DigitalOcean in Amsterdam. The median response time from my machine is around 100 ms. After enabling CloudFlare, I cannot see a significant difference in response time. The median response time, and the distribution of response time, looks very similar.
I guess that during the last few years you have expanded your network and your connections with the major hosting providers (Amazon, Google Cloud, DigitalOcean, Linode, etc.). Maybe it explains the difference between today's test and 3 years ago?
In general, is it useful and/or recommended to use CloudFlare in front a fully dynamic service, for example a HTTP-JSON API, with no static content (no images, no stylesheets, no scripts), and thus no need for the CDN feature?
In general, is it useful and/or recommended to use CloudFlare in front a fully dynamic service, for example a HTTP-JSON API, with no static content (no images, no stylesheets, no scripts), and thus no need for the CDN feature?
We do have lots of customers who do that. Two reasons: Railgun and Security. Railgun gives speedups for the JSON because of the ability to diff the boilerplate JSON. Security for APIs is of course important and clearly attackers like to go after APIs.
About security, what are the specific security features you're thinking of?
Just hoping you can give a concrete example, as I'm not super familiar with this stuff.
About DDoS I don't think there's any cheap solution for that...
That's a dealbreaker for me at their prices.
Edit: Linked to Tengine in [1].