How Google Code Search Worked (2012)
swtch.com
swtch.com
I'm imagining a "drill down" TUI with rg and fzf. fzf can be good for both filenames and other filter-downs. Thinking of breadcrumbs and easily stepping forward or backward, ability to easily bookmark/"pin" parts of search paths as presets for easy reuse later, etc.
EDIT: I recognize this would be outside of the scope of rg itself, I'm voicing it in case it sparks ideas about the functionality you're thinking of adding. I'll think more about it and see if I can explain better
Here is how I used it - I’d type in some code I was working on and the search result would show similar code and how it was used. Great for debugging and thinking by looking at similar solutions. Sigh.
A couple open source projects that I've seen are Hound and Zoekt. Hound actually uses this code search backend with a nice frontend in React. Zoekt is what I was going to use since it scales really well, is faster, and has good search operators for filtering by repo name, language, etc. Google was using Zoekt until recently for code search across all their open source repos.[0]
Power drill is fantastic for drilling down data. So much money was made thanks to this one.
With codesearch, answering those kind of questions is near instant.
Blaze is amazing (albeit a bit slow) but Bazel should be more or less the same, haven't used it. Dremel, Spanner, Tensorflow, Proto, grpc are all available outside. Abseil (https://abseil.io/) is a great library available to everyone.
I mean, we’re talking about code here. Text meant to be interpreted and understood by a compiler. Why can’t we do better?
Why can’t I say “show me everyone that’s calling this function”, like an IDE lets me do? Or “show me functions that accept <type> as one of their arguments and return <type>”, in a way that integrates with the real grammar/AST of the language(s) in question, without resorting to clunky regular expressions?
I should be able to write structured queries against a codebase, with regexes being just one part of that query language.
For example C has a preprocessor and linking step driven by a build system. And C has a bunch of different build systems available, some of which are procedural rather than declarative.
Maybe you'll need to support package management - if a function signature calls for a CopyOnWriteArrayList do you need to know what the subclasses and superclasses of that type are? Do you need to resolve all the dependencies to be able to do that?
If you're thinking "No problem, everyone compiles their programs in CI anyway" - are you happy to skip indexing unused code and uncompileable code?
And of course you'll be chasing after language and build tool changes - not only to one language, but every language.
On the other hand, a nice simple grep? Sounds much simpler to me.
You can try it out here: https://cs.chromium.org/
In typical google style, the documentation is all google-internal, but by clicking bits of source code you should figure out most of the commands.
Doesn't work well on mobile unless you have a beefy CPU - sorry!
Quoting patio11: "I intend to boot up a livegrep instance on the first day of every startup for the rest of my life. It borders on miraculous."
It is indeed very good.
> To minimize I/O and take advantage of operating system caching, csearch uses mmap to map the index into memory and in doing so read directly from the operating system's file cache. This makes csearch run quickly on repeated runs without using a server process.
Does anyone know some resources where I can read more about this technique? (how-to, pros/cons, caveats, etc.) I'm interested in figuring out the best way to have a commandline tool persist state that it can quickly access across multiple runs, but so far a background server process is the only technique I'm familiar with.
Was that a joke? Does someone have a Lion system around that can verify?
It's been there for 11 years in this repo. Along with DECnet. Suppose there's no real drive to remove it.