Improving GitHub Code Search
github.blog
github.blog
Finally!
Search-for-literal is so important when you have technical users working on non-prose text.
They say this is going in a dedicated search page 'to start with', if "<literally any text>" doesn't work in the top bar eventually this is still going to be miserable.
I use Github's UI for exploring and searching codebases more often than my own environment, since I do a lot of curious browsing.
No offense, but the search is so bad for anything worse than a single word, that I've developed a sort of intuition for how to phrase things -- and then still spend a lot of time crawling pages of results haha.
This was sorely needed
The index is likely based on Roaring bitmaps, presumably https://github.com/RoaringBitmap/roaring-rs in this case.
Nice architecture, exactly how I would have done it also.
You need to support a proper query syntax, with tags, rankings, stopwords, stemming. Then you need to have a proper db backend (reverse indices). Trigrams dont help for regex. Then a templated representation. Google codesearch would do only the 2nd of 3. ElasticSearch is commercial, and only java.
Doing that from scratch is a bit silly.
Code and log search are two specialized use-cases of search that definitely warrant a non-full-text approach and as far as I am aware the trigram bitmap index would be considered state of the art for both.
That isn't to say you can't solve the problem with a full-text search engine, many people do with the solutions you alluded to. However they are drastically less efficient and probably out of the question at Githubs scale.
Microsoft has done a great job of actually improving the product and investing in it, something that seems to not happen most of the time with giant acquisitions.
GitHub today, at least for me, is definitely improved over the GitHub five or so years ago. It does feel like some things are bloated due to a push for feature after feature, but the core features have gotten better to the point I don't care. I just wish I didn't have to spend 2 minutes turning off all the annoying features I don't care about for small projects on every new repo I create.
I admittedly haven’t used Visual Studio in a while (though I use VS Code daily) but I remember it as my favorite IDE. Certainly better than XCode, and I even preferred it to IntelliJ.
That'd cover 95% of repo I've seen.
My main use case for GitHub search is identifying provenance of misc. changes in vendor source code tarballs for e.g. Android kernel releases. It's hard, but sometimes possible to rehydrate most of the existing commits through cherry-picks and careful rebases.
The biggest problem with the lack of indexing branches and forks is that sometimes vendors makes releases through branches, or that sometimes repos of interests are forks of e.g. `torvalds/linux`.
Hopefully we can see those being indexed in the future.
I'm also curious: has the plan to drop "less active" repos from the index gone through? Has anything changed?
Whaaat? I hope it doesn't go through. I use GitHub code search for clues when reverse engineering cheap Chinese IoT crap. Usually I can find some headers / SDKs accidentally uploaded and set to public by a random Chinese guy. Those repos usually have one commit and zero traffic, but they contain invaluable information about proprietary MCUs.
Also the search text field is bit messed up in Safari when the text gets longer than the field.
NOT language:html
Having to use dropdowns and multiple input fields is more cumbersome than the filter language of repo:, file:, lang: etc.
Would still be great to ignore my docs directory.
(Though it would be even better if the two options for case-sensitivity and regex search were enabled by default instead of needing me to toggle them on every time.)
"search.defaultCaseSensitive": true,
"search.defaultPatternType": "regexp",
Also see:
https://docs.sourcegraph.com/admin/config/settings#search-de...Ólafur is responsible for a lot of Scala tooling and some pretty neat original ideas.
Sourcegraph also came up with LSIF, which is useful format for building tooling for language servers:
If you want to build this sort of stuff, the work Sourcegraph has done with LSIF + SemanticDB is probably your easiest bet.
N=2 isn't great, but there's my experiences if we're tossing them out there.
Comment (and code) search is available for projects in all GitLab tiers: https://docs.gitlab.com/ee/user/search/#basic-search
Premium and Ultimate users have access to Advanced Search: https://docs.gitlab.com/ee/user/search/advanced_search.html
So it's a workaround, but a bad one.
Here's the ticket (2015): https://gitlab.com/gitlab-org/gitlab/-/issues/13891. The fact that it has so many duplicates in your own project's issue tracker is a good indicator of how bad your issue search is.
> no way to combine a text search with label/milestone/status filters, etc.
You can combine text search with field search (like label/milestone/status. Here’s an example: https://gitlab.com/gitlab-org/gitlab/-/issues?search=Visuali...
Give it a shot and let us know what you think! Where can we improve it?
Also it would be great if more repositories were indexed. How do things work behind the scenes? Maybe it is possible to build a more memory-efficient index just for exact string search, which probably make up most searches.
Anyway, this website is amazing and I use it quite often. Thank you a lot for working on this!
Some feedback:
* Copy & Paste does not work. When trying to select text from code snippets, the code is dragged like an image instead of being selected.
* The GitHub icon during search is not clickable. Clicking to the right of the GitHub icon does not go to GitHub but instead shows an embedded view of that repository. There, the GitHub icon somewhere on the right is a bit hard to find (maybe write "View on GitHub" next to it so it is discoverable with Ctrl + F?), but at least it is clickable.
We're also working hard to increase the number of repositories indexed. :)
The results are just so full of repetition
Search improvements? It is impossible to create a worse search experience than Github. Just clone and use git grep instead in most cases.
Edit: ...and the 425% price increase for SSO..
(Granted, this is largely due to a culture of titles like "Check out this thing" that provide zero searchable metadata + no tag system.)
This is on our radar! We de-duplicate exact matches now, but we'd like to do the same for near-similar documents.
1) an ability to easily express higher level concepts in a search that's aware of code semantics ("match only function names", "find call sites of a method") etc. Maybe this is possible with existing tools (probably is?) but I tend to get lazy about learning DSLs - would love to see this in a UI if it's possible
2) ability to save searches I do frequently - after a certain level of complexity in a query (I've added ignore rules, I crafted the right regex, etc), I want to be able to save the "context" of a search so that I can easily return to it later
> would love to see this in a UI if it's possible
We do have code navigation via the UI, so in a way it's possible!
> ability to save searches I do frequently
Absolutely! This is possible using "custom scopes". If you're in the technology preview, click on the scope dropdown, scroll to the bottom, and choose "custom scopes". You can make a custom scope to search a set of respositories, a particular language, within a directory, or any combination with boolean operators!
http://lemonodor.com/archives/001232.html
For example, if you're looking for a search-and-replace function you know you wrote or had somewhere on your machine, you could do
mdfind "org_lisp_defuns == '*search*replace*'"
(Or just use the regular Spotlight UI.)1) This doesn't seem to exist in quite that way, but you can prefix a literal with "def:" and the engine will return only definitions of that thing (so far as it can tell). It's not quite what you (or I!) want, but close.
2) This exists and is called "scopes". On the landing page, to the left of the search bar, click the grey pill that says "All repos". At the bottom there is a "custom scopes" option.
I've been lucky enough to have a few projects that others have found useful, and so they've ended up in Conda forge, Gentoo, and other package repos. After I make a release, the Github-wide search is just absolutely flooded with dependabot PRs, Snyk PRs, and dozens of other bots. Literally thousands (and sometimes tens of thousands for CVEs) of automated PRs and issues that make it impossible for me to see how others are using or discussing my packages (usually to see if I've just horribly broken something).
Googles codesearch tool[1] actually compiles the code and uses the compilers parse tree to make the search index. The only time it doesn't work is if you are looking for code that has been "#ifdef 0"'d, when it falls back to regular string matching, and the difference is night and day.
Please github... please try to make a search index by compiling everyones files. Plenty of projects have CI buildbots which have all the info to automatically compile millions of projects, and at the same time generate the necessary parse trees. Even for interpreted languages like javascript/npm, python/pip, you can use heuristics to make a cross-module function call graph accurately most of the time.
[1]: Example URL from it: https://source.chromium.org/chromium/chromium/src/+/main:thi...
Stay tuned for more!
It is a custom search engine, built from the ground up for code. We'll be sharing more details about it on the GitHub blog soon.
PS: How do you feel about being bought by Microsoft? Maybe some of you feel it's a good time to implement s2s inter-forge federation to plant a nail in the coffin? Sounds like Gitea is on a good way to support it based on ActivityPub/forgefed and it would be sad if Github was relegated to a for-profit walled garden.
E.g.
p: or f: instead of path: for filenames
l: instead of language:
-f: to exclude specific filenames (makes it easy to filter out tests)
You get the idea.
https://github.com/isaacs/github/issues/402
I don't even care it's very fast. Just make it work. I just hope this isn't snake oil. Weird that they claim regex support but no wildcard support.
whenever I search for code, it will say something like "Last indexed on Apr 2", but if you go to the actual file, the date will say 5 years ago or something. So currently the "Last indexed" listed date is completely useless, and you have to basically click through to every result.
Yes, sadly, that is literally when the file was _indexed_. So it's not particularly useful. It's a difficult problem to solve, but I'll bring up your feedback to the team.
It is not a friendly site. Open source projects would do better to use an open source code forge like <https://sr.ht/>.
ps: yes, enterprise server instances
Skipping well known vendoring directories would be cool too.
Time to go secrets and url hunting.
Unfortunately I can't reproduce the problem publicly because it happened while searching a private repo.
No mention of case-sensitive search...
Not ashamed to be a Rust evangelist! The reason I mentioned Rust is because we spent a lot of time making the experience really fast - which is super important for a product like this. I really think getting the performance we have would have been enormously more difficult in any other language.
In general, Rust, C, and C++ are going to be faster than languages like Ruby*. He brought up Rust while discussing the performance of the new tool. Although performance is more complex than language choice, etc., saying it's written in Rust gives the viewer an approximate lower bound as to how fast the tool should be.
*: (GH started as a Ruby shop, so I wouldn't be surprised if that's what the original tool was written in).
Benchmark - ripgrep is faster than {grep, ag, git grep, ucg, pt, sift} (2016) - https://blog.burntsushi.net/ripgrep/