Kagi's internal index seems to focus on sites that wouldn't necessarily block spiders.
https://help.kagi.com/kagi/search-details/search-sources.htm...
I've crawled small web sites (even hosted on major providers) to archive them and never hit a wall.