+1. Also according to some swedish friends, aftonbladet is not a very high quality news service.
68 karma · joined December 6, 2021
And since you control what URLs need to be crawled, you protect yourself against rogue clients sending arbitrary URLs.
There certainly are a lot of elegant ways to reduce spam for this particular problem imo.
How are the batches of URLs to be crawled generated/discovered and posted at your API?
How do you deal with duplicate crawls?
Would be nice to compile a list of them!