> privacy-first analytics solution [...] The only downside of this is that we need to keep access logs (IP & User-Agent, no browsing history) for 24 hours
Keeping IPs don't make you super privacy friendly.
> privacy-first analytics solution [...] The only downside of this is that we need to keep access logs (IP & User-Agent, no browsing history) for 24 hours
Keeping IPs don't make you super privacy friendly.
* We're getting hit with a huge DDoS attack, repeatedly over 3 weeks, with no sign of stopping
* With zero access logs, there was no way to find patterns in the attack, and we had no way to block it
* Our service was going offline during these attacks
* We introduced access logs that are auto-deleted after 24 hours. We redacted all information about the site/page/activity etc. but keep IP & User Agent for pattern matching
* We were then able to identify a pattern and block the attack on Saturday
* Without access logs (even redacted ones), this wasn't possible
I was hoping a more senior engineer on Hacker News would comment and I can't wait to hear how you'd do it. I have no experience in DDoS protection at all, and this seems like the only possible way. Even rate limiting requires storing IP addresses. But if you know a more privacy-focused way to block these attacks, I'm sure I'll buy you a few beers when we hang out.
I've read the full blog post, I am not convinced it's a DDoS attack. Traffic patterns for web analytics will come from over the place and will look like a DDoS when it's not. For example, a customer misplacing their analytics in a JS loop and having a moderate traffic blog will generate billions of requests from all over the globe. Event tracking can be billions of requests as well by themselves. Never attribute to malice that which is adequately explained by simpler means.
> * With zero access logs, there was no way to find patterns in the attack, and we had no way to block it
It's a loss of time and energy. Your system should be able to handle these billions of requests.
> * We were then able to identify a pattern and block the attack on Saturday
Was it specific accounts?
> * Without access logs (even redacted ones), this wasn't possible
You can one-way hash the IP. So you can still look for pattern but you've lost the actual IP. And same one-way hash IP can block whichever IP seems devious in your firewall. (like md5 can be enough.)
> But if you know a more privacy-focused way to block these attacks, I'm sure I'll buy you a few beers when we hang out.
Haha no need to. But come say hi if you ever in Austin. julien _at_ serpapi.com.
I like the idea of an MD5 hash. Although I'm not too certain why an IP would be a bad thing to log for 24 hours. From a privacy law perspective, the MD5 hash is considered PII. And if we see an IP address in an access log, we know that an IP visited one of the tens of thousands of websites Fathom runs on, but we don't know which one.
Edit: Something else, with MD5 there's no way of finding patterns with the IPs, so you'd have to play whack a mole. Whereas raw IPs allow no real privacy invasion whilst allowing pattern detection
Layer 7 just means they are making ton of HTTP requests from lot of IPS however this is the nature of a web analytics.
> we know that an IP visited one of the tens of thousands of websites Fathom runs on, but we don't know which one.
How can the attackers know which websites have your analytics on? Crawling the web is super hard.
You should be able to get back which accounts are affected from your analytics call:
https://starman.fathomdns.com/?p=%2F&h=https%3A%2F%2Fusefathom.com&r=&sid=BIABKBRK&res=1440x900
Isn't sid=BIABKBRK the account? The access log should store this `sid` allowing you to just block the account. (or ask them for ton of money if it's a legitimate use :) )Exactly! Which is why it was so hard. There's no path pattern. Everything hits "/". So the only way to fight back is to match IP / header patterns (but even then, we have to redact sensitive headers).
> How can the attackers know which websites have your analytics on? Crawling the web is super hard.
The attacker went after some of our more high profile customers. They're known via testimonials or from Twitter.
> The access log should store this `sid` allowing you to just block the account. (or ask them for ton of money if it's a legitimate use :) )
We could certainly temporarily block traffic to a site. The problem is, without some kind of firewall (e.g. WAF), our application has to absorb so much traffic, and that's the issue. We need to block it at the edge.
I think you need some high performance ingestion code.
One day somebody is going to load test or misconfigure their site. Or you'll get a really big client. Or somebody will get Reddit frontpaged or slashdotted. Blocking huge volumes of traffic isn't always going to be an option. You should be able to handle it.
I would get off all this "sexy" stuff like lamdba and SQS and build some old school battle boxes at a co-lo. Run some extreme performance framework like Actix or Vert.X built to handle giant piles of traffic. A single box with those frameworks and a 40 gig line can handle millions of requests a second.
Use counting bloom filters synced between boxes for IP blocking. This avoids saving IP's and is needed to prevent ram exhaustion attacks anyways. Use two layers, one for individual IP's and a second for blocking subnets with lots of bad actors in them. Block for ~24h by using dual sets of bloom filters in a sliding window with 12h overlap where IP's get added to both. When a filter is 24h old, discard it. This needs to be done because you can't remove items from Bloom filters, they saturate over time.
On these super boxes you can do filter, blocking, batching, and dedup. Then send the data off to wherever for further processing.
You mentioned not wanting to use Cloudflare because they're a competitor. That's fine but I would take a look at their blog. They go over the tech they use for mass data ingestion and filtering, and it's basically what I just described.
The plaintext space (amount of possible IPs, even more so for IP blocks) is so small that that you can try all the possible plaintexts within seconds, essentially reversing the hash.
Could you salt the hash with something way less predictable that gets discarded when the next one is generated? Assuming you only kept logs for a day for example you could regenerate it daily and store $salt somewhere only the most senior of senior techs can access it (if anyone at all)
By using Bloom filter you avoid storing IP's, and save tons of space. They're basically required for rate limiting in IPV6 where the address space is so large an attacker can just run you out of ram by using so many different IP's.
It's also pretty common to have several thresholds of blocking based on individual IP vs subnets. So if more than 2 ips get blocked in a small subnet, you block the whole subnet.
Potentially, but then customers don’t see a ton of spam on their dashboard. Just a lot of repeated requests.
They don’t come in random waves either. Valid traffic is pretty well spaced out.
> Your system should be sble to handle these billions of requests
Potentially, but you really don’t want to pay for them. The only real way to go about stopping it is blocking offending ips.
I wouldn't worry about short-term IP caching. AWS's upstream load balancers and your own servers are probably doing it anyways to maintain TCP state tables. Linux kernel's "conntrack". If you don't want to cache IP's you can you a probabilistic data structure like a bloom filter synched between instances. If has a small false positive rate but is very fast and doesn't store whole IP. Bloom filter based IP filtering is used in every big DDOS prevention system I know of.
As much as you like lambda, I would ditch it. And the queues. My general advice is that any time you need to add a work queue to something, it's not fast enough.
Your analytics endpoint data ingestion should be something lightning fast like Go, Rust, or an async Java back end. Analytics is a lossy process, you lose traces because of browser behavior and plugins all the time anyways so I wouldn't prioritize 100% accuracy.
I would focus on power/dollar over reliability. If I was you, my ingest boxes would be load balanced with DNS round robin and sitting at various Colocation providers. Get a fat 40 gig unlimited data pipe. Build some stupid fast Rust/Go/Java backend that can saturate that pipe. And do all your filtering/spam analysis here.
I don't think lambda, SQS queues, PHP are the best technologies for this kind of mass data ingestion. I don't even think your ingestion layer should be on AWS. I would follow the lead of other companies doing mass data ingestion and build your own machines. That's how CloudFlare, Netflix etc are able to handle so much traffic without going bankrupt.
I would consider yourself lucky that your first DDOS was so small. 10k requests/sec is tiny. ~400k/sec can be generated on a regular desktop with fiber internet connection. Right now, a single user could knock you offline by messing around with JMeter. I think it's a wake up call that you're in the infrastructure business whether you like it or not, and you need to massively beef up your data ingestion layer. Realistically, you should be able to handle ~50 gigabit attack with 10 million requests/sec. I think that's achievable with a couple boxes colocated on 40 gig lines running fast software
I've had a similar ddos attack which manged to bypass the cloudflare protection and our app handled it without $0 increase in bill.
Please be respectful of OP, he put a few weeks working on the issue, it is not fair to tell him that you think everything he tells is shit, only because you don't believe what he says
* I found AWS Shield Advanced organically
* It was a DDoS attack
* I am incredibly happy with the personalized service. It’s like hiring someone to handle it, except you have people on call 24x7
This is what confuses me out of a lot of these replies, saying you're being swindled for paying for this "absurd service" when you should just "own your infra yourself and hire people to manage this"
Do they think owning all this yourself and hiring specialized DDoS people is gonna cost less than $36k/yr? Because I don't think it will.