Try the powerful search, including free text search, operators like f: or symbol: (the full list of operators can be found at https://developers.google.com/code-search/reference)
Then try the references feature by clicking on symbols.
Try the powerful search, including free text search, operators like f: or symbol: (the full list of operators can be found at https://developers.google.com/code-search/reference)
Then try the references feature by clicking on symbols.
You'd need to do a lot of educated guesses to properly build the cross-references, and within Google you can just build the code or whatever analyzer you need.
But you're half right here, using blaze makes it way easier.
The cross references underneath Google codesearch (both internal and external) are created by, roughly speaking, partially compiling everything and groveling around in the AST.
One of our teammates Luke Zarko has a decent overview of this: https://www.youtube.com/watch?v=VYI3ji8aSM0
> Bizarre, especially when you consider that the whole underlying infrastructure of google codesearch has been open-sourced.
Some of the code is also open sourced at kythe.io, though it lacks a considerable amount of the internal logic necessary to scale it out to all of google3. So it's not exactly plug & play. But, it should be enough for a company the size of Microsoft/Github to get started using it.
Admittedly, we need to do a better job of articulating this difference, but if anyone would like help getting spun up with precise code intelligence on Sourcegraph, please DM me: https://twitter.com/beyang.
cs.chromium et al experiences instantaneous dropoff from "wow this is kind of amazing" the moment you start wading through JavaScript or Mojo glue code. Here's the JS source of the dino game; notice how you can't click anything: https://source.chromium.org/chromium/chromium/src/+/main:com...
A quick Google for "most popular languages on github" just found https://madnight.github.io/githut/, which (if it's correct) reveals that GitHub's most popular language (by commit activity) is Python (17%), very closely followed by JavaScript (14%). Then you have Java (12%), TypeScript (8%), Go (8%), C++/Ruby (both 6%), and PHP (5%). So on GitHub C++ apparently represents 6% of PR activity, while the top two languages (Python/JS) which represent 31% of PR activity are not only interpreted but also dynamically typed.
Which tells a very interesting story about the benefits of being able to generate an AST with type information: you can scale insight that much more broadly and deeply. :(
You need to be able to build the code to even have a chance at that level of understanding of the source. And really you need to be able to build the code with a specialized toolchain. This is tractable with a huge investment if you get to choose or dictate your tools. If folks are bringing their own tools it is basically impossible.
Pushing for standardization of fact output from toolchains could be a path to getting somewhere when you can't dictate toolchains. Or really just having some kind of fact output from default toolchains would get you a long way.
> why don't you have decent crossrefs
We do have compiler-accurate cross references for many repos. Some examples:
- TypeScript: https://sourcegraph.com/github.com/sindresorhus/got/-/blob/s...
- C: https://sourcegraph.com/github.com/neovim/neovim/-/blob/src/...
- C++: https://sourcegraph.com/github.com/KhronosGroup/Vulkan-Sampl...
Why are not all repos covered?
Because different languages have different build systems, so inferring the right build commands, dependencies etc. is not so straightforward; these are necessary pre-requisites for compiler-accurate cross references. We're working on fixing this with auto-indexing: https://docs.sourcegraph.com/code_intelligence/explanations/...
For C and C++ specifically, auto-indexing is challenging because of the large variety in build systems, informal specification of dependencies (such as in a README instead of a machine-readable format), and platform-specific code.
Outside of auto-indexing, we do have an indexer for C and C++ right now (https://github.com/sourcegraph/lsif-clang) which can be run in CI; that way one can generate an index and upload it to Sourcegraph on a regular basis. It is 'Partially available' (https://docs.sourcegraph.com/code_intelligence/references/in...) right now. We're keenly aware of the interest in C++, and are working our way through different languages based on usage.
https://source.chromium.org/search?q=comment:%27%5B%5E%5C%5C...
https://source.chromium.org/search?q=comment:%60%5B%5E%5C%5C...
In this case... about even unfortunately lol.
Disclosure: I work at Google.