There have been DDoS attacks recently on Stack Overflow and Stack Exchange, I think it's quite plausible those have something to do with this. They still seem to be under active attack intermittently.
There have been DDoS attacks recently on Stack Overflow and Stack Exchange, I think it's quite plausible those have something to do with this. They still seem to be under active attack intermittently.
The new blocking completely blocks access to the site, not just posting. See for yourself; https://www.torproject.org/
Tor is next to useless for DDoS attacks, as it doesn't offer any amplification. For every byte you send in via TCP, you get one byte out. For attackers with large botnets, it doesn't make sense to DDOS over Tor, as the number of exiting IP addresses are limited and it's easy to block them all. It makes more sense to use your thousands of available botnet IP's that aren't on any lists.
> Tor is next to useless for DDoS attacks, as it doesn't offer any amplification.
It seems that these attacks are not being carried out to take down the Stack Exchange network. It seems that these attacks are to take down Tor as a legitimate technology. Here me out.In order to carry out these attacks, the attacker already controls enough machines to DDoS one of the largest websites on the internet. So, huge botnet or nation state. Now, we have one of the largest websites on the internet telling it's tech audience:
> An immediate solution for users who find themselves blocked is to access our site
> from other IP addresses, via home internet, work internet, or other VPN services
https://meta.stackexchange.com/questions/376060/update-on-th...They are normalizing the "workaround" of using an insecure IP address when Tor is inaccessible. This will lead to all non-secretive and non-illicit Tor usage to go back to the open insecure internet. Thus, everyone still using Tor "has something to hide" (as if that wasn't the case already). By forcing all but the most desperate users off Tor, Tor can be discredited as a nefarious tool.
I think this is a reach - after all, if you wanted to discredit a tool for people in oppressive regimes, people who care about their privacy, and people doing illegal things, why start with the programming help site? (I know SE has other sites, but the biggest ones are for tech help)
...That said, I could see this being a way LE could try and unmask a very high-value darkweb programmer. Still a reach.
My guess is they'd target websites that are important to TOR users, and tor.stackexchange.com is one of them. I also don't think they started with SE. TOR IPs are continually filtered by Google, Cloudflare, and other automated firewalls.
Stackoverflow pages have many dynamic components like vote counts, reputation points, sidebar related-questions-links, new comments, etc. An excerpt from their blog explains they can't cache the output : https://nickcraver.com/blog/2019/08/06/stack-overflow-how-we...
Even though Stackoverflow's website io access pattern has higher reads than writes, the resultant generated html is still not as static as Wikipedia pages. Even if they're efficiently using cpu to generate the dynamic elements, the cost of egress traffic may also be a factor.
All that said, I don't have any insight into what heuristics they use to block certain ip addresses.
The egress traffic is also trivial - a page seems to clock in at below 50k. If they're paying $0.02/GB, that's $1 / million requests.
More importantly, if it really were a DoS attack, there's way less obtrusive methods, such as a CAPTCHA or similar verification screen.
> variants of cache... anonymous, or not? mobile, or not? deflate, gzip, or no compression?
I don't understand why the markup would vary between mobile/desktop, or why they would even consider compression a varying factor to account for in cache. Maybe that's because their backends produce such specific variants that they have a hard time caching in the first place?
If 80% of pages are only requested every two weeks, then it doesn't make sense to cache those. There's still probably lots of stuff you can cache, such user profile/stats and question/answer scores for several minutes. You can also cache many parts of the markup that are not going to change often. I mean it's always even better when you cache the entire markup, but there can still be lots to gain if you cache only small bits that are expensive to acquire/template.
> But the cost of memory to store those strings (most large enough to go directly on the large object heap) is very non-trivial. And the cost of the garbage collector cleaning them up is also non-trivial.
It looks like their cache implementation could (should?) have been based on better foundations. First, i think the cache doesn't have to reside in memory: disk accesses are fast, certainly much faster than running a database query over the network which will need to access several files and cross-reference data with extra latency on top. Then, because if you're gonna store long-lived stuff in memory you should probably use a garbage collector based that's tailored for this usecase, not your language's (.Net) default GC... maybe Redis? Don't get me wrong, i find it pretty cool if some engineers want to develop a homebrew cache, but that sounds like a huge project in itself.
> the cost of egress traffic
I'm not aware of SO tech stack, but i'd be surprised if they have much egress fees. They're a very big site so they probably run off unmetered dedicated servers if not their own hardware on cheap transit. Who knows, they may even have their own AS and peer with other providers in some locations? From a quick request, it looks like stackoverflow.com is served from Fastly AS but i personally don't understand why: i don't remember seeing any heavy content (video/images) on SO so in that kind of situation a CDN would hurt more than help on slow links. Maybe that's because they bundle megabytes of javascript crap? Now that they're blocking tor users i can't check for myself :-)
Things like this https://news.ycombinator.com/item?id=26072025
Fixes, updating the app, blocking the image, blocking all requests with empty user agents.
------
1 person abused enough SO, they dig, find the culprit being someone using Tor network.
Fixes, identify the user and ask them to stop, block all traffic from Tor.
-------
Do you propose an alternative fix?