I went ahead and polished up the announcement for an early release:
https://sourcehut.org/blog/2022-07-15-searchhut/
Let me know if you have any questions!
I went ahead and polished up the announcement for an early release:
https://sourcehut.org/blog/2022-07-15-searchhut/
Let me know if you have any questions!
If you happen to be cloud hosting this, and if you do not have a global rate limit, implement one ASAP!
Several independent search engines have been hit hard by a botnet soon after they got attention, both mine and wiby.me, and I think a few others. I've had 10-12 QPS, sustained load for weeks after weeks from a rotating set of mostly eastern european IPs.
It's fine if this is on your own infrastructure, but on the cloud, you'll be racking up bills like crazy from something like that :-/
e.g. "Any websites engaging in SEO spam are rejected from the index" - how is determined whether something is SEO spam or not? More clarification of whats allowed/not allowed would be nice!
https://searchhut.org/docs/docs/webadmins/requirements/
And there's some advice for web masters on ways to improve your site's ranking without running afoul of this rule:
https://searchhut.org/docs/docs/webadmins/recommendations/
But ultimately, it's subjective, and a judgement call will be made. If it's minor you might get a warning, if it's blatant then you'll just get de-listed.
I don't recall if it supported SourceForge and GitHub (2008) but it certainly included gzipped tarballs which were popular and prevalent at the time.
e.g opencrawl, internet-archive, archiveteam
It strikes me the resources to crawl, update, and manage/index data is a common problem.