Apparently, they expanded their pool of available IPs they pull data from and now they seem to be endless (so some of the scraping domains actually work now).
I'm investigating what I can do about it. I'd appreciate any advice!
Apparently, they expanded their pool of available IPs they pull data from and now they seem to be endless (so some of the scraping domains actually work now).
I'm investigating what I can do about it. I'd appreciate any advice!
I wonder how banning so many IPs affects CloudFlare performance and if I should optimize it to block whole IP ranges instead ...
You might also be able to find common user agent headers through the CF firewall page and block them based on their UA. That'd not work if their scraper tool was randomizing the UA string but quite a few of them don't
As for user agent - they're using a very common, real browser user agent that's impossible to distinct from legit users.
Russia has no GDPR or something. So you could put (special key in) a cookie? They probably do not process it so subsequent requests without a cookie are to be discarded?
It is entirely permissible under the GDPR to use cookies for security purposes.
I blocked what I could and then I blocked the whole country (Russia) behind a CloudFlare javasript challenge (so that any legit traffic that wanted to pass it - could).
Everything stayed blocked for about a day and now it seems they gave up:
- all domains that pulled dynamic content from my site now show some other content and do not try to scrape Next Episode
- all static domains that delivered the cached version of the Next Episode homepage now deliver some other website(s)
- CloudFlare shows no further traffic on the firewall rules I've set (so the bots from those IPs are, for now, gone)
As this post is relatively old now (more than 3 days) I doubt many people will go back and check for new developments, but I wanted to give this the proper closure and let you know of the apparent happy end.
You could look at the http request headers and perhaps identify the scrapper script. You could also put a javascript challenge that is required to solve before pulling more data, and disable it for Google and Bing ips, so it's more work for them to pull data for some time.
Instead of simply blocking, you could detect them and do some kind of http slowloris response.
I'll try and find out and also I'll have to learn exactly what "slowloris" is. It may be helpful indeed!
It's usually targeted to webservers, before most HTTP servers got fixed you could DDOS a server with a tiny connection, but some HTTP clients can also be vulnerable.
But you may find better usage of your time than implementing this.
In this case though, to defeat scrapers maybe create some link which only the scraper sees and leave a "gzip-bomb"' like described here https://blog.haschek.at/2017/how-to-defend-your-website-with... and see how their scraper handle that :-)
Personally I just used a html-fuzzer to generate 5 MiB of junk html and named it wp-login.php :-) And a ssh-tarpit
As an aside, I’ve fought credential stuffers by returning real looking but actually false data, and initiating password resets... start serving different data on each hit, you may need to be annoying enough that they give up.
Problem is - right now I'm over 250 (new) IPs and they keep piling up (their domains now rarely use an IP more than once).
I may have to block entire ranges of IPs or whole ASNs.
Then, setup a script on your laptop or whatever to search this string on their domains every half hour or so.
It even prepares the expression snippet for me to paste directly into a CloudFlare firewall rule.
That's how I got to quickly identify and ban almost 2000 different IPs.
If they continue to expand the IP pool I may need to automate it though.