GitHub doesn't seem to be either. I get that it's a competitor but not being able to search GitHub is probably a deal breaker for most devs that aren't Drew.
I'm not opposed to indexing GitHub, but the signal to noise ratio on GitHub is poor. Nearly all GitHub repositories are useless, so we'd have to filter most of it out. I think instead I'll have to set it up where people can request that specific interesting repositories are added to the index, and maybe crawl /explore to fill in a decent base set.
GitHub is hella tricky to crawl too due to its sheer size and single entry point (meaning slow crawl speed). I've been looking at the problem as well, and so far just ignored it as un-crawlable, but I might do something like crawl only the about pages for repos that are linked to externally some time in the future.
There's an asterisk to that: they serve the underlying content through two different APIs so one can side-step the HTML wrapper around the bytes: the discovery phase has a formal API (both REST and GraphQL) for finding repos, and then the in-repo content can be git-cloned and one can locally index every branch, commit, and blob, without issuing hundreds of thousands of http requests to GH. One would still need to hit GH for the issues, if that's in scope, but it'd be way less http requests unless your repo is named kubernetes or terraform.
We're still talking about git clone:ing a hundred thousand github repos. Git Repos get big very fast. That's a lot of data when you're realistically only interested in is a few markdown files per repo.
Perhaps all repo's that have a published package is a good heuristic. Then you'll at least get all the repos of npm, python and other packages.
Some interesting repos have no published packages. A combination of number of commits, stars and forks would be probably more relevant.
And likewise, some uninteresting repos do have published packages.