Do you have a cost estimate? Also could you be more selective in indexing, e.g. by having users requests sites to be crawled.
Do you have a cost estimate? Also could you be more selective in indexing, e.g. by having users requests sites to be crawled.
Obviously failing first requests isn't ideal but for popular requests it quickly becomes insignificant. Wikipedia might (if they don't already) want to make a similar suggestion for users to contribute when finding a low content/missing page.
The first request can also be called asynchronously, and display a message to the user that it is 'processing....'.
If I search for a news event it's a news site.
If I search an error message, I know the result is going to likely be stackoverflow, github issues or the forum of the library.
etc.
I don't think this strategy will get you all the way there, but it could be combined with other ways of curating sites to crawl.