The technology behind GitHub’s new code search
github.blog
github.blog
Because I pushed such a repo yesterday and it’s still not been indexed.
I thought it would be nice to be able to search through them, but I guess the file limit was reached. That’s a shame.
if it works, which I really hope it does, we should have a repo containing every file (sub-25mb) published to pypi.
it's not really useful to clone this, but the git packfile + index seems to be pretty ideal for storing + querying this data, reducing 14tb of compressed packages down to like 100gb or so.
this should enable large scale exploration of the contents of pypi from only your laptop, which is useful if you want to look at how the python language is evolving (how many packages use f-strings over time?).
the rust tool i've built to do the heavy lifting is here: https://github.com/orf/pypi-import-test. i'm learning a lot about git internals, it's quite fun
Someone less experienced than you may take any of those results as gospel. Gospel the same way StackOverflow could be, but this time by a source who will do it’s best to tell you what you want.
my personal rule about AI is about the same as the rule about being a lawyer. Never ask a question you don’t know the answer to.
Personally, I'm happy with the new code search so far. I stopped using Sourcegraph because I could never get the deeper results I wanted - it would just return the top five repositories including some common code snippet and I couldn't explore further than that. GitHub Code Search doesn't seem to have this problem to such a degree, since I can use negation more naturally, and since my query is not limited to some shallow subset of the corpus before refining it.
Our new ranking (https://about.sourcegraph.com/blog/new-search-ranking) should help a lot here, and it's live on https://sourcegraph.com. Can you share some of the queries you tried so we can see how much ranking helps and how to handle them better?
Thanks for the link to your blog post. I didn't realize Steve Yegge had joined your team - congratulations, that's quite an endorsement!
I've always liked Sourcegraph just because our company names are so similar (founder of Splitgraph here)... we might have even gotten a few candidates because of that initial confusion. :)
How has GitHub Code Search impacted your product direction? Do you see it as an opportunity to focus more on the internal use case, or do you have plans for some other differentiation? It's always unfortunate when a big company introduces a product so similar to the core product of a startup, but I'm sure there is a silver lining there, especially when you have a talented team and a mature codebase (for example, Fly.io has been able to carve out a niche for itself despite Cloudflare moving to compete in the same areas). Either way, best of luck to you from a fellow S-grapher!
If your language's package manager puts your deps within the repo directory by default, people will commit the vendored code. See: Node, Go (since go.mod).
As long as we can all agree that committing compiled code is a crime punishable by 24 hours in the shame cube.
But you don't NEED to do this do you? I'm ALREADY in a repository, I just don't want to check out, say all of WebKit, I just need to find where a specific reference is defined.
Maybe, maybe on a really serious day do I need to search an entire organization. But hardly ever.
I have never, in over a decade ever, wanted sophisticated symbolic searching from GitHub code search, I just need remote grep.
Why is the code search not feature bisected into this 99% use case, and then the occasional global repository search, which can behave entirely differently?
If you check out our prior blog post, "A brief history of code search at GitHub" (https://github.blog/2021-12-15-a-brief-history-of-code-searc...), you can learn a bit about the evolution of this feature. And, in fact, we used to use git grep to search repositories.
This doesn't work well at GitHub's scale. We have 100M users and over 200M repositories in a multi-tenant environment. Your git grep is going to be competing for resources with other user's pushes and clones.
> when the search doesn’t show me what I’m looking for on the first page, this is exactly what I do
This is pretty much always.
If you're in some repository that uses a webkit api, and you want to know how that api is defined, how do you do that without global cross references or a global lookup?
Even for local lookups, indicies are useful (as any ctags user will tell you!), but for any kind of cross repo xrefs they're ncecessary.
What you use github search for doesn't require all this engineering, but what I use it for does. Why wouldn't they build something that satisfies both our necessities well?
On grep.app I regularly search all repos. It's very useful for finding out how to use APIs or where APIs from dependencies are defined.
So I suspect you don't want it because subconsciously you know that Github's "search all" feature won't return you useful results.
Hell they still don't provide a way to filter out test directories which makes the code search inside a single repo useless a lot of the time.
I used to search my org's code on a daily basis and all of GitHub a few times a week at least, back when I used git for work.
I think you may be overestimating how common your specific use-case is.
That seems very slow by today's standards. There's a rather... eccentric guy who easily beat that almost 10 years ago with his implementations of string search: https://news.ycombinator.com/item?id=6954298
The indexing they talk about in that article seems like rearranging the deck chairs on the Titanic so far as that is concerned.
It’s possible GitHub could also leverage those language servers (which would be super cool) but doing it at GitHub scale is certainly not trivial.
Poor man’s version using the new GitHub search would be to construct a regex that matches one but not the other.
To be specific, I was looking for the definition of one method in Highcharts so I could understand what it does and override it, GitHub gave 6 pages of results. I was able to find the function immediately in my IDE once I checked out the 100mb+ repository and it indexed it. If I’d been able to the same w/ the search on GitHub it would have saved me considerable time and hassle.
This search could be implemented by something that compiles and indexes like the IDE (sourcegraph) or maybe some kind of shallower parsing. Highcharts is in typescript which I’d don’t know well but in JavaScript the later might be a little tricky because there are so many ways to define a function (one hell of a regency.). I’d contrast to Java where is would be very easy to write a rule that would turn up a class definition if not a method definition, in my case finding the class would have solved my problem.
https://github.com/highcharts/highcharts
for Series.drawPoint and expecting a direct hit for
https://github.com/highcharts/highcharts/blob/29d2a83a5a997b...
practically I tried "Series" and "drawPoint" also.
/func.+MyFunc/
vs a search for where it is used: /\.MyFunc/
I agree that it would be nice to have more language aware IDE-style features but I’m just happy to have regex support which is almost always powerful enough to express what I am looking for.I also worked on something similar to the search engine that is described here for the purposes of making auto-complete fast for C++ in Clangd. That was my intern project back in 2018 and it was very successful in reducing the delays and latencies in the auto-complete pipeline. That project was a lot of fun and was also based on Russ Cox's original Google Code Search trigram index. My implementation of the index is still largely untouched and is a hot path of Clangd. I made a huge effort to document it as much as I can and the code is, I believe, very readable (although I'm obviously very biased because I spent a loot of time with it).
Here is the implementation:
https://github.com/llvm/llvm-project/tree/main/clang-tools-e...
I also wrote a... very long design document about how exactly this works, so if you're interested in understanding the internals of a code search engine, you can check it out:
https://docs.google.com/document/d/1C-A6PGT6TynyaX4PXyExNMiG...
Github shards and indexes individual files according to their hashes. It also uses variable length ngrams (neat!). This makes horizontal scaling simpler, but also means more of the index needs to be scanned for org/repo-scoped queries ("Due to our sharding strategy, a query request must be sent to each shard in the cluster.").
They mention in a blog post linked to this one (https://github.blog/2021-12-15-a-brief-history-of-code-searc...) that Zoekt indexes much bigger than the corpus size, so they don't work at GitHub scale.
The github code search looks like an impressive piece work (congrats!).
That said, I'm curious about the nuances regarding corpus size. Their blog post claims they have 115 Tb of source code, but that a positional index is "too expensive". A positional index is a 3.5x blow-up, which is ~500 Tb of data. A 1 Tb SSD retails for $50, so that's $25,000 for storing a positional index. 500T of GCP local SSD is also ~25 k$/year. Even if you factor in replication/redundancy, the resource cost is far less than hiring a software engineer. I guess they think machines with local SSD are too much overhead to manage?
I’d love to see more discussion on how they are dealing with the false positives though. It looks like a positional index is being used to achieve this, but that usually blows out your index size.
Additional information about deduplication would be especially interesting to me as well. It seems to solve this quite well. I usually try a search of Jquery to test this and it does not return multiple copies of different versions of it which is a good indicator that it’s slightly fuzzy.
What I find really interesting about all the code search engines I know of is that each one implemented its own index. Nobody is using off the shelf software for this. I suspect that might be down to no off the shelf software providing a decent enough solution, and none providing a solution that scales. At least none that scales with decent costs.
I did a small comparison of GitHub code search a while ago https://twitter.com/boyter/status/1480667185475244036?s=61&t... But I should note a lot has improved since then, and it looks like sourcegraph now also does default AND of terms rather than exact match, so my complaints there are resolved.
Impressive work by GitHub. I am sure some of the people behind it will read this comment, let me say well done to you all. I am very impressed. Also please post more information like this. There is so little out there.
I started writing my own very simple index and search engine, but quickly decided to just use ClickHouse via https://tinybird.co as my backend (Serverless SQL with automatic APIs is pretty sweet) because I wanted to build out the product side of things and my data is really small, so I felt like it was going to be a lot of effort for little reward.
Maybe one day I will need to write a custom index or search engine that actually scales though :).
An early version of Blackbird experimented with trigrams plus a bitmask of the next character, but it didn't work well because it wasn't selective enough. This is mentioned in the blog post:
We tried a number of strategies to fix this like adding follow masks, which use bitmasks for the character following the trigram (basically halfway to quad grams), but they saturate too quickly to be useful.
Thanks for the clarification. Looking forward to see what else you and your team end up writing about. Which reminds me to publish some other posts I have about searchcode.
Seems a lot of others in this thread are interested as well.
I mean GH got a long way using ElasticSearch until now.
I'm not sure that's true. It used ES for a long time, but the search was also terrible, at the edge of complete uselessness.
The current search sucks ass, you can’t find anything.
I was trying to search for something in the WebKit source the other day and I had to use Sourcegraph because the GitHub search gave me zero results.
Obvious disclaimer because I work on this (and I also worked on the legacy search): It is way better.
So it's possible that a document containing `aaa` might match our ngram search, but we double check after retrieving them and exclude them from the result set.
Actually, our search engine is so fast that syntax highlighting the search results is often slower than finding them... so if we store the language tokens directly in the index, we'll be able to directly emit syntax highlighted snippets and make it even faster.
It may also enable some interesting search capabilities in the future, like searching within comments or by code structure.
I'd always wondered how they implemented that: it turns out they add extra internal filters to their searches along the lines of "RepoIDs(...) or PublicRepo".
Question for the team: Do you have an additional permission check in the view layer before the results are shown to the end-user? I worry that if I switch a repo from public to private it may take a while for the code search index to catch up to the new permissions.
Another fun example is that your SSO session might have expired. While you technically have access to view the result, we can't show it until you go through the refresh dance to get another valid token.
We read all the feedback on the forum here: https://github.com/orgs/community/discussions/38692, so please keep providing it. Videos and screenshots are super helpful too. Thanks for bearing with us as we continue to polish the UX!
It feels like github code browsing is a step between a full editor with lsp and a static site. I Hope they work out the Kinks and make it more smooth
We do have a C++ code indexer in beta, https://github.com/sourcegraph/lsif-clang - it is based on clang but C++ indexing is notably harder to do automatically/without-setup due to the varying build systems that need to be understood in order to invoke the compiler.
As my colleague mentioned in a sibling comment, we have an existing indexer lsif-clang which supports C++. I just added a Chromium example to the lsif-clang README right now: (direct link) https://sourcegraph.com/github.com/chromium/chromium@cab0660...
We are also actively working on a new SCIP indexer which should support features like cross-repo references in the future. https://github.com/sourcegraph/scip-clang
Right now, Abseil doesn't have precise code navigation because no one has uploaded an index for it. In an ideal world, we would automatically have precise indexes for all the C++ code on Sourcegraph, but that's a hard problem because of the large variety in build systems, build configurations, and system dependencies that are often specified outside the build system.
But this could be tied into CI, especially for projects utilizing runners for building the code.
So the barrier to entry now is orders of magnitude less and with github and others pushing for codespaces it could be the final piece to tie everything together.
One often overlooked drawback of generating this data during CI is that you, the project owner, are now paying for the compute. Of course, maybe you qualify for a free tier. And if not, if you're already running a CI job for testing/linting/etc, the marginal cost of also generating code nav symbols might not be too bad. But one benefit of the stack graphs approach mentioned up-thread (and the code search indexing described in OP) is that all of the analysis and extraction work is done in dedicated server-side jobs that GitHub is footing the bill for.
I don't think what you are saying is actually true for stack-graphs[0][1].
[0]: https://github.com/github/stack-graphs
[1]: https://github.blog/2021-12-09-introducing-stack-graphs/
This metadata file would be generated by a language-specific tool. For instance, Cargo could generate it for Rust projects, ctags/cmake for C / C++, Sorbet for Ruby.
Then a service like Github / Gitlab / your own homegrown code viewer could provide things like "Show type", "Jump to source" without ever needing to build a language-specific parser or interpreter (which seems arbitrarily difficult, most build systems provide escape hatches, so you can't assume much about project structure).
Basically, LSP as a static file in a standard format that tools can read to understand a codebase, without needing to model the language's semantics.
I think a language-agnostic semantic metadata format is a good idea, but requires a lot of compromise. ctags partially does this, but only to a very coarse level (mostly definitions and references). I think some ctags implementations also define 'extension fields' that could be used to give type information, but I don't know how/if these are used in practice. SemanticDB is extremely fine-grained, but highly specialized to JVM languages and type systems that are designed to work with the JVM. Finding a common set of semantic features that can be used across languages and type systems that is fine-grained enough to be more useful than ctags sounds very difficult to me.
We took that torch and carried it forward, building the spiritual successor called SCIP[0]. It's language agnostic, we have indexers for quite a few languages already, and we genuinely intend for it to be vendor neutral / a proper OSS project[1].
There are also some community-maintained OSS SCIP indexers. - rust-analyzer: https://sourcegraph.com/github.com/rust-lang/rust-analyzer/-... - scip-zig: https://github.com/zigtools/scip-zig
We also have Precise Code Navigation — still currently only for Python [1], but we're working on other languages as fast as we can. As mentioned down-thread, that is built on our new stack graphs framework [2]. We're often asked why we're not leaning on something LSP-shaped for this. I go into a built of detail about that in a FOSDEM talk from a couple years back [3], and also in the Related Work section of a paper that we just published [4].
[1] https://docs.github.com/en/repositories/working-with-files/u...
[2] https://dcreager.net/talks/stack-graphs/
It's a great 101-level exercise to write an inverted index implementation you can do it in an afternoon , and then expand to a leaf /aggregator in follow-up exercises.
If I wanted the kind of search engine I can get a teenager to write in 16 weeks why would I expect my org to be paying $$$ for the service?
Symbol extraction for C and C++ is currently disabled because we were having problems with the performance of the tree-sitter queries we were using, but we are planning to bring that back.
I wonder if it would be possible to leverage LSP as a kind of tokenization generalization framework, or even piggyback off of the existing effort by incorporating search-friendly/-helpful metadata into future versions of the protocol spec.
This is what a 101-level inverted index implementation looks like: https://github.com/BurntSushi/imdb-rename
In other words, absolutely nothing like what GitHub built. Nowhere close.
It seems to me that you know what you want from such a service, but focusing on making C++ exceptionally great in this service would come at the cost of, say, general quality across all languages, or frontend usability. A very reasonable tradeoff for a beta-quality product.
Why must you be so harsh on them?
Even now I find myself using it despite the index being a few months out of date.
[0] https://github.com/django/django/search?q=DeleteView&type=co...
[1] https://github.com/django/django/blob/main/django/views/gene...
The underlying data may be limited (I have no idea how large it is, I doubt it has indexed every public repository out there), but I never failed to find examples of what I was looking for.
It's the number one way I research and understand new libraries/API's and programming languages.
There's a lot more you can learn from usage in the wild than tutorial posts sometimes.
Most often I end up using code search for figuring out where a piece of code originated, just to find thousands of random projects that have also copied the same code verbatim. Sorting for "relevance" or "latest/oldest indexed" are equally useless.
1) I never want to search all repos globally. At worst I want to search all of my org's repos.
2) the search UI is a little clunky, in a way I'd need to be using it again to remember.
Between those two I think there's loads of progress to be made outside of raw search power. Of course it's nice to have that, but that's what I'm really after.
[1] https://docs.github.com/en/repositories/working-with-files/u...
Is it inverse frequency, so common bigrams get split last? And the goal is to be able to search on a larger gram that covers the more common trigrams as often as possible?
One nit I have about current search: I’ll look something up and find I’m getting results for some obtuse commit in some old branch somewhere. I’d like to be able to optionally say “latest commit on branches only please” or “main branch only please.”
Another thing, which might betray that I don’t understand search all that well: language aware searching that knows, for example, that a single or a double quote are syntactically interchangeable. Don’t omit half the results because I used one quote over the other when looking up `interpolation = ‘nearest’`
Like searching for "OPTION" and getting "-DOPTION=TRUE" among the results. Very commonly needed to find all usages of a flag, even instances where the flag is being passed to (at least, that I know of) CMake and Meson.
[0]: https://stackoverflow.com/questions/43891605/search-partial-...
SourceGraph likely has challenging times ahead considering the valuation.
But to be clear, as a company we are doing well and growing nicely inside customers, with a ton of cash in the bank, an awesome team, and a huge opportunity ahead of us. GitHub’s new code search has been out for 14 months now, so this is nothing new.
It’s a big market and there’s way more room for differentiation and dev choice in code search/intelligence than in CI. There’s a lot of code intelligence that GitHub won’t support (precise code nav for more languages, comprehensive code ownership, metadata from other dev tools that know things about code outside the GitHub/Microsoft suite, etc.), there’s a need for the ability to fix (with our Batch Changes) not just find, and even in the core search workflow there’s so much room for improvement with AI fine-tuned on your own code, etc.
But talk is cheap and only shipping matters. So, watch what we ship, and send any feedback and requests our way!
The article is a top-notch technical write-up, the devs on GitHub code search should be proud of what they've achieved so far!
Honestly, we're rooting for GitHub to improve their code search, viewing them as a close peer-not a competitor. We also maintain OSS projects like Zoekt, which IIRC GitLab is maybe looking at using for their own. The more devs that 'get' code search, the better off Sourcegraph is frankly!
GitHub has a nice intuitive/simple UX, we could learn a thing or two there (though, easier to do with less features.)
Still, Sourcegraph search tech is quite a bit more powerful:
* Searching over commit messages, diffs, filename, etc. are super nice for tracking down regressions / finding 'that PR I swear my coworker made'
* Expressiveness like "find this regexp in repositories, but only if the repo has had a commit in the last month AND has a file named package.json in its root"
* Since Steve Yegge joined us, we've started thinking about ranking of search results, a notoriously difficult thing to do well in code search unless you have great factors to rank on (e.g. a semantic understanding of code): https://about.sourcegraph.com/blog/new-search-ranking
* We stream results back, so you can get a comprehensive set of results - not just a few pages, from our API.
* Works in GitHub Enterprise, not just GitHub.com. Plus on all your code hosts, think BitBucket, GitLab, Azure DevOps, Gerrit, Phabricator, etc. and even non-Git VCS like Perforce.
* Respects permissions of all your code hosts (a very difficult problem, as there are no official APIs to query this info from code hosts in general)
Having code search is one thing, but using it is another:
* Code Insights (we use search as an API to gather statistics about code, track code quality, keywords, etc. both over time and retroactively and let you build dashboards)
* Batch changes (find+replace, but over thousands of repositories. Run a Docker container per repo, run your custom linter script etc. and then draft or send PRs to thousands of repos, manage/track campaigns with thousands of PRs like that over time, etc.)
* Precise code intel / semantic awareness of code, we use SCIP indexers for this (spiritual successor to Microsoft's LSIF format for indexing LSP servers.)
I am super happy GitHub continues to push their code search effort, and genuinely believe it's a great thing for all developers and us over at Sourcegraph. Also excited to see when they do their public rollout of this :)
Anyway, that's just my take as someone who works there-other Sourcegraphers will chime in later if anything I said above feels off to them I'm sure :)
We are looking at Zoekt for code search: https://gitlab.com/groups/gitlab-org/-/epics/9404
Overall, we’re building what our customers need, and our product goes way beyond what GitHub can offer. Sourcegraph indexes all the code and increasingly all the code intelligence (including code nav but also code ownership and other metadata in the future from your other dev tools). We charge based on active usage, so we make money when devs at customers /choose/ to use us over the alternatives. We’re trying to do this the right way, and tons of customers agree. (If anyone reading this disagrees, please let me know!)
Re: your comment about our sales team, I’m really sorry to hear that and want to understand more so I can fix the problem. Can you please email me at sqs@sourcegraph.com?
We don’t think any of today’s code host vendors with their current strategies can make truly great code search and intelligence because they’ll be biased toward their own bundled tools and limited to the subset of code hosted on that instance. It’d be kind of like Encyclopedia Britannica or The NY Times building a web search engine: helpful, but so much more limited compared to what the independent Google became.
And yes, none of this was a surprise. GitHub’s new code search has been out for 14 months now.
OK, hope this puts an internet rumor to rest!
What exactly do they mean by "special repositories" here?
You can write your own search engine that will perform very well on a surprisingly large amount of data, even doing naive full-text search. A search tool I came across a while back is a great example of something at that scale: https://pagefind.app/.
For anyone who doesn't know anything about search I highly recommend reading this (It's mentioned in the blog post as well): https://swtch.com/~rsc/regexp/regexp4.html.
Algolia also has a series of blog posts describing how their search engine works: https://www.algolia.com/blog/engineering/inside-the-algolia-....
---
It's interesting that GitHub seems to have quite a few shards. Algolia basically has a monolithic architecture with 3 different hosts which replicate data and they embed their search engine in Nginx:
"Our search engine is a C++ module which is directly embedded inside Nginx. So when the query enters Nginx, we directly run it through the search engine and send it back to the client."
I'm guessing GitHub probably doesn't store repos in a custom binary format like Algolia does though:
"Each index is a binary file in our own format. We put the information in a specific order so that it is very fast to perform queries on it."
"Our Nginx C++ module will directly open the index file in memory-mapped mode in order to share memory between the different Nginx processes and will apply the query on the memory-mapped data structure."
https://stackshare.io/posts/how-algolia-built-their-realtime...
100ms p99 seems pretty good, but I'm curious what the p50 is and how much time is spent searching vs ranking. I've seen Dan Luu say that majority of time should be spent ranking rather than searching and when I've snooped on https://hn.algolia.com I've seen single digit millisecond search times in the responses, which seems to corroborate this.
I'm curious why they chose to optimize ingestion when it only took 36hrs to re-index the entire corpus without optimizations. A 50% speedup is nice, but 36hrs and 18hrs are the same order of magnitude and it sounds like there was a fair amount of engineering effort put into this. An index 1/5 of the size is pretty sweet though, I have to assume that's a bigger win that 50% faster ingestion.
Since they're indexing by language I wonder if they have custom indexing/searching for each language, or if their ngram strategy is generic over all languages. Perhaps their "sparse grams" naturally token different for every language. Hard to tell when they leave out the juiciest part of the strategy though: "Assume you have some function that given a bigram gives a weight".
Search is so cool. I could talk about it all day.
It's interesting that GitHub seems to have quite a few shards. Algolia basically has a monolithic architecture with 3 different hosts
I used to work at an Algolia competitor. I don't know for sure, but my guess is that Algolia shards their indices by customer. Algolia does not provide global search. GitHub code search does. That, and the desire to deduplicate data, is what led us to our current sharding strategy (notably, it is different than the old GitHub code search's sharding.).
I'm guessing GitHub probably doesn't store repos in a custom binary format like Algolia does though:
We have a custom index format, so I would say this is the same, unless you mean something different. We of course translate repos from their Git form to our index document form for indexing.
I'm curious why they chose to optimize ingestion when it only took 36hrs to re-index the entire corpus without optimizations. A 50% speedup is nice, but 36hrs and 18hrs are the same order of magnitude and it sounds like there was a fair amount of engineering effort put into this. An index 1/5 of the size is pretty sweet though, I have to assume that's a bigger win that 50% faster ingestion.
The index size is a bigger win, but being able to reindex quickly is huge for our development velocity and trying things out. We really feel it when things are slow. This is also not our final goal, we want to scale the system up considerably.
Interesting, do you keep a copy of the index document form of repos or is that done on the fly during indexing? Is your custom index format a binary format? I have no idea whether that's standard practice, or just a compressed text format is enough. I guess that non-binary formats would be enormous though, and given that an index is by definition relatively unique it probably wouldn't compress that well.
I do feel the development velocity thing. I've felt something similar on my smaller scale projects. Being able to fully re-index the corpus in less than a day definitely seems like it would provide a lot of opportunities to experiment and try stuff out without it being too costly.
Scale up in terms of what? Is the current system not indexing all of GitHub, or you mean you want to index on more things (E.g. commits, PRs, etc)?
My guess is that Algolia's indices are sharded by customer and each cluster probably has multiple customer indices.
do you keep a copy of the index document form of repos or is that done on the fly during indexing?
As mentioned in the post, the index contains the full content. Our ingest process essentially flattens git repos (which are stored as DAGs) into a list of documents to index (the prior state is diffed for changes).
Is your custom index format a binary format? I have no idea whether that's standard practice, or just a compressed text format is enough. I guess that non-binary formats would be enormous though, and given that an index is by definition relatively unique it probably wouldn't compress that well.
Binary formats are normal, posting lists are giant sorted blocks of numbers, so there are a lot of techniques that can be used to compress them. Lucene's index format is pretty well documented if you're interested in learning more (interestingly, Lucene has a text format for debugging: https://lucene.apache.org/core/8_6_3/codecs/org/apache/lucen...).
Scale up in terms of what? Is the current system not indexing all of GitHub, or you mean you want to index on more things (E.g. commits, PRs, etc)?
It's not indexing all of GitHub yet, nor do all users have access yet. Those are the things we are focusing on now. In the future, we want to support indexing branches.
Ever since then, I've exclusively used sourcegraph.
Our Code Search team is currently working on moving to Zoekt[0] which is expected to be a significant improvement as it is purpose-built for code search.
We also shipped an improvement[1] to our existing search functionality at the end of last year. If you haven't used it recently, I'd encourage you to check out code search again to see if the quality has been improved for you.