https://www.brentozar.com/archive/2015/10/how-to-download-th...
Though that is a bit conspiracy-theory-y and I expect blocking Tor is more an issue of a small but very active number of people using it to create accounts to post anonymous spam or abuse. For that, which can be a very real problem, I would suggest a more fine-grained block: stop people from those addresses creating accounts, or posting from relatively new (or recently inactive) accounts, or perhaps somewhere in the middle: prevent connection from Tor addresses posting at all, but still let people use the site. That way you still block the abuse, but have less impact on others accessing via Tor.
StackExchange is composed of contributions from the Internet at large. I've appreciated over the years how StackExchange took this stewardship very seriously without hoarding it behind paywalls - the ethical thing to do.
One of the reasons I was motivated to contribute is because 10-12 years ago, putting tech questions in Google often led to Experts Exchange, a site that takes contributions from the public and makes you pay to see them. (Really need a StackExchange-type site to take on Pinterest now ...)
So ... no, they should not stop publishing updates. People need to bookmark sites and go directly to them instead of relying on an increasingly broken search engine that's past its heyday.
Well, if we are whopping them out on the table... 75.1K over the three I've used most, and a few thousand over a couple of others. Been around since the start, even got sent free shirts & bits in the early days for being in the top [what-ever-the-cutoff-was] on SF, SO, & DBA.SE.
> So ... no, they should not stop publishing updates.
I wouldn't mind if they stopped providing those dumps, I was just passing on that others suggest that they could (and by rights, they could, but I doubt they will, as my post (I though) said).
As long as the data stays available (just on the main sites is fine by me) under a CC BY-SA I'll keep contributing. They can't revoke the license for existing content, if they change it for future content (for the avoidance of doubt: I have no reason to believe they have any plans to) or the sites otherwise devolve (like others have in the past with rampant tracking and ads) I'll probably bugger off.
Not to be too cynical about it, but who serves the ads on those sites?
If it's Google Ads, then I think you have your answer as to why they are higher ranked.
This at least makes sense, why run the JavaScript in your scraper? It's easy to identify. What do your webserver access logs say?
I don't know how much Ye Olde PageRank still factors into Google results, but I'm assuming it's still a significant factor.
I mean anyone who'd like to crawl/scrap a website can do it from home, with the help of friends, or by renting "resential VPN" services for a few bucks a month. Blocking Tor does not prevent scrapping.
Google should definitely resolve this problem on their side, but language models got huge and complex. It's hard to tell (especially with their scale) if a website if composed from paraphrased content of 20 other websites.
I'm kinda shocked how successful those "doorway pages" are at reaching the top and how long they live there.
There are 322 SO copycats alone.
The "data" is under creative commons. It does not belong to them.