How are the batches of URLs to be crawled generated/discovered and posted at your API?
How do you deal with duplicate crawls?
How are the batches of URLs to be crawled generated/discovered and posted at your API?
How do you deal with duplicate crawls?
It changed a bit in the implementation.
And since you control what URLs need to be crawled, you protect yourself against rogue clients sending arbitrary URLs.
There certainly are a lot of elegant ways to reduce spam for this particular problem imo.
I'm not worried about the URLs, but the content of the URLs sent back.
Say the server tells a client to crawl a CNN article. The "hacked" client sends a fake CNN article back.
Some might, because of A/B testing or news updating, but even updating news will get a positive similar page and those that don't should probably fall into an exceptions category until it can be determined what can be done about it. Maybe a flag in the URL to give you a static page or just accept that it changes often enough that even faked pages won't last long?
There is a challenge for sites that serve different content based on GeoIP, A/B testing, dynamic content, etc. So some human review of the diff may help check for malice. If there's literally spam, human review would clearly detect this and that bot is distrusted.
Plus I can now cause you to have to run your own crawler anyway and either slow progress or cost you a lot of money.
You can buy a trustworthy residential IP for low cost, you can buy them in bulk in the thousands. All of them are real residential IPs from any ISP of your choosing in any country. You can rent Chrome browsers running over those IPs, directed via remote desktop and accessibility protocols (good luck banning that without running awful of anti-discrimination laws). You can do all that for under 1k$ a month for like 1 million clients.
My workplace has been at the other end of DDoS attacks directed by such services, best you can do is ban specific Chrome versions they use but that lasts until they update.
It's an uphill battle that you will loose in the long term if you rely on client trust.
I think in this specific case, the spammer is on poor footing. The spammer wants to inject specific content, ideally many times. With double processing of URLs and the spammer controls 50% of the clients then there's a 50% chance that a simple diff would show the injected spam. The problem is that the spammer needs to do this many times, so their injection becomes statistically apparent. If the spammer can only inject a small number of messages before they are detected, then the cost per injected spam will be quite high. Long running spam campaigns could eventually be detected by content analysis, so the spammer also needs to rotate content.
Obviously you can play with the numbers, the attacker could try to control >>50% of the clients. The project could process URLs >2x. The project could re-process N% of URLs on trusted hardware, etc. It's not easy by any means, but you can tune the knobs to increase the cost for spammers.
You make it sound easy. ;)