Sites have to make exceptions for Google, but likely wouldn’t care enough to allow other search engines in.
Sites have to make exceptions for Google, but likely wouldn’t care enough to allow other search engines in.
https://help.kagi.com/kagi/search-details/search-sources.htm...
I've crawled small web sites (even hosted on major providers) to archive them and never hit a wall.
[1] https://help.kagi.com/kagi/search-details/search-sources.htm...
My take (as a subscriber) is still that the majority of information I'm seeing is assimilated and distilled from the more mainstream sources, but I'm possibly underestimating how much comes from their own indexes.
I will admit that for news I tend to use the !gn bang to jump straight to Google News as it tends to be a more complete feed than Kagi's news tab, so wasn't thinking about TinyGem. But I'd also forgotten completely about Kagi's "Small Web" initiative for indexing non-commercial content that's not interlinked enough for other engines to prioritize it. I believe that's what Teclis (and possibly TinyGem as well) concentrates on.
IIRC, the few times I tried routing searches into that index I didn't find much, so I set it aside, but that's probably because I'm usually searching for more widely-consumed content.
> Kagi Search includes anonymized requests to traditional search indexes including Brave, as well our own non-commercial index (Teclis), news index (TinyGem), and an AI for instant answers. Teclis and TinyGem are a result of our crawl through millions of domains, focusing primarily on non-commercial, high-quality content.
[Edit] I guess it is possible that Google or Bing are hiding under "traditional search indexes including Brave".
(Microsoft famously sells theirs to DuckDuckGo, right?)
I don’t think Google has a publicly known licensing deal ,
Ecosia have also been relying on Bing, but recently announced that they've now added Google as well.