How to block semalt.com referrer traffic using .htaccess
logorrhoea.net
logorrhoea.net
Clicky recently tweeted about banning them from the stats https://twitter.com/clicky/status/445704464750501890
As usual I cannot penetrate the marketing-ese explanation provided by their homepage.
Any ethical crawler will respect exclusions defined in that file. The article doesn't mention trying this method before jumping to mod_rewrite, though.
So -- does Semalt's crawler look for and follow robots.txt? As long as it does, they're doing what they should be doing to let site owners opt out of crawling.
EDIT: Found a page on their web site where you can enter a domain to opt it out of crawling: http://semalt.com/project_crawler.php
No instructions for how to opt out via robots.txt, though. That's a big omission. Anyone who's going to do mass crawling needs to support robots.txt.
semalt can not fix this by offering an 'opt-out' while ignoring robot.txt
very stupid of them because they are shooting themselves in the foot by this very simple but MASSIVE breach of standards of conduct
(If you're using Apache and do want to start fighting bots, I'd suggest taking a look at mod_security. Very powerful, but beware that the default rules can be touchy)
Way to often have I seen systems getting overrun with rogue traffic, usually spam-bots and vulnerability scanners, and lots of small sites get into serious trouble because of this.
SetEnvIfNoCase Referer crawler.semalt.com spammer=yes SetEnvIfNoCase Referer semalt.com spammer=yes
Order allow,deny Allow from all Deny from env=spammer
location / {
valid_referers none blocked *.semalt.com semalt.com;
if ($invalid_referer) {
return 403; }
}