Last I saw Kagi was highly dependent on Google for their search results and seemed much more interested in LLMs and other side features than in replacing that core part of their search stack.
Additionally, I don’t think it’s fair to say it’s more interested in LLMs than focusing on search. I think it’s fair to say they’re interested in ensuring they’re offering a better, non ad-based search replacement.
Disclaimer: Not affiliated with Kagi in any way, just a long time happy user.
By now, crawling and indexing is a herculean task, and also quite expensive, due to the sheer size. There is Common Crawl [1]; at 400 TiB it is huge, but at 60 days refresh interval it's far from being very comprehensive or very fresh. Good for research, but likely not good for a commercial search engine.