ignoring the problem of robots.txt inaccessibility, is it feasible to have Kagi-style "private google" with a more limited number of high-signal-to-noise sites, especially if you drop the concept of e-commerce and some other low-SNR feeds?
perhaps one interesting thing is that a decent number of the highest-SNR feeds don't actually need to be crawled at all - wikipedia, reddit, etc are available as dumps and you can ingest their content directly. And the sources in which I am most interested in for my hobbies (technical data around cameras, computer parts, aircraft, etc) tend to be mostly static web-1.0 sites that basically never change. There's some stuff that falls inbetween, like I'm not sure if random other wikis necessarily have takeout dumps, but again, fandom-wiki and a couple other mega-wikis probably contain a majority of the interesting content, or at least a large enough amount of content you could get meaningful results.
Another interesting one would be if you could get the Internet Archive to give you "slices" of sites in a google takeout-style format. Like they already have scraped a great deal of content, so, if I want site X and the most recent non-404 versions of all pages in a given domain, it would be fantastic if they could just build that as a zip and dump it over in bulk. In fact a lot of the best technical content is no longer available on the live web unfortunately...
(did fh-reddit ever update again? or is there a way to get pushift to give you a bulk dump of everything? they stopped back in like 2019 and I'm not sure if they ever got back into it, it wasn't on bigquery last time I checked. Kind of a bummer too.)
I say exclude e-commerce because there's not a lot of informational value in knowing the 27 sites selling a video card (especially as a few megaretailers crush all the competition anyway), but there is lots of informational value in say having a copy of the sites of asus, asrock, gigabyte, MSI, etc for searching (probably don't want full binaries cached though).
But basically I think there's probably like, sub-100 TB of content that would even be useful to me if stored in some kind of relatively dense representation (reddit post/comment dumps, not pages, same for other forum content, etc, stored on a gzip level5 filesystem or something). That's easily within reach of a small server, not sure if pagerank would work as well without all the "noise" linking into it and telling you where the signal is, but I think that's well within typical r/datahoarder level builds. And you could dynamically augment that from live internet and internet archive as needed - just treat it as an ever-growing cache and index your hoard.