Code Search at Google: Han-Wen and Zoekt
sourcegraph.com
sourcegraph.com
Most major parts are open sourced at kythe.io, and there's a somewhat dated talk given by Luke here: https://youtu.be/VYI3ji8aSM0
Corollary: while we can do a lot with indexing generated code (even cross language) in Kythe, there are limits. Macros may be one, I forget atm
Unfortunately I don't know if there's any significant use of Kythe outside of Google. We get a handful of questions on the open source repo from time to time, but that's all I know about.
Docs: https://docs.sourcegraph.com/code_navigation/explanations/pr...
Is it possible at all this story helped spur the widespread adoption of git (the early implementation of this tool)?
The story mentions git because git5 got me into developer tooling. More in particular, it put me in touch with Shawn Pearce who ran the Git/Gerrit team at Google. When I went to work for him, Shawn wanted to have codesearch support in Gerrit, and Zoekt was ultimately the outcome of my explorations in this space.
IIRC, Git5 was deleted approximately 5 years ago because Fig (the Hg based replacement) had taken over all the use cases
However, I don't think it makes sense to downplay git5. Anecdotally, basically everybody knew about it, and I'd constantly run into people using it (which is by itself noteworthy since nobody was exactly talking about version control all the time).
Git5 was at the time the most robust solution to chain commits, which was tedious bordering on impossible without some tool. Without definitive data, I wouldn't say users were less productive with git5: it definitely was a useful tool that people at least recommended for chain commits. I was definitely more productive with it.
There were a lot of footguns though, and I do think the hg wrapper that superseded it was way better.
So git essentially was the "final form" that integrated all the various workflows and topped it off with a maximally-scaled use case (linux) that proved out the tool, drove innovation in integration/scripting/gadgetry, and provided a clear beacon for everyone else to adopt it. So it won.
But in 2004, if you asked around, everyone would have told you that a tool like this was coming at some point (even if they probably wouldn't have described it as very git-like!).
It was a... special time. Let's not reminisce.
After that, I think many people moved to subversion, which had a lot more functionality for distributed VC, for exmaple there was a server. svn was popular for a while but building it was painful (due to berkeley db) and it sort of never grew. I invested a lot of time in (specifically apache with mod_dav and mod_dav_svn) but lost interest in VC after fighting with subversion.
git came along and from what i can tell it mainly had "it's by linus, and the kernel uses it" and "it's fast" and "something about reflogs". I use git day-to-day but I still; can't explain how git became so ubiquitous; I find using it outright painful.
Both were definitely way better than SVN/CVS/etc.
Since I have limited brain capacity I focus my efforts on being able to use git, not hg, merely because it has so much marketshare.
As someone who has used all of them in different companies since the 90s: RCS was ancient, even in the early 90s. Most widely found in things like UNIX source trees.
CVS came along later (mid/late 90s), and was much more widely used.
SVN came on the scene around the late 90s -- it was a massive improvement on CVS, and spread across open source and most professional shops like wildfire. Major sites like sourceforge were built on SVN, but also supported CVS.
Git only became prevalent starting around 2006-2008, and adoption was actually really slow because of its inherent complexity. When Github appeared, that was really when the shift started in earnest.
There were others along the way: MS SourceSafe, a moment when everyone was toying with Bitkeeper, etc., but these were as marginal as RCS.
One of my first real accomplishments was migrating that codebase from RCS to CVS- which was relatively easy as CVS used RCS under the hood.
so, in terms of algorithms, Zoekt wasn't actually inspired by Google's internal code search.
The precise query syntax of zoekt is mostly copied from google's internal syntax, though.
> "Zoekt, en gij zult spinazie eten" - Jan Eertink
> ("seek, and ye shall eat spinach" - My primary school teacher)
I'm on the Kythe team, and I don't know off the top of my head what Zoekt is. Looking it up, I see it's some sort of trigram search, which means if it's used at all (I have no idea), it's codesearch proper, not Kythe.
The Kythe index is the semantic index of the codebase, Codesearch does all of the text/regex/etc searching.
You really should check it out if you haven't already. It's incredibly useful; I used it all the time. Not open source though.