Google Open Source Code Search
cs.opensource.google
cs.opensource.google
(I work at Google)
You can firewall off Sourcegraph 100% for complete confidence, and aside from the first admin's email address (so we can notify them of any security updates) we only send back aggregated anonymous usage statistics which we are extremely transparent about: https://docs.sourcegraph.com/admin/pings
We sell developer tools, not user data.
The option is documented in our config docs, though, and also appears in the config editor's autocomplete in the app if you type `telemetry`, though, so it's not really a secret https://docs.sourcegraph.com/admin/config/site_config#disabl...
What kind of telemetry is critical? How do I disable that?
[0] https://docs.sourcegraph.com/admin/config/site_config#disabl...
I'm human and screw up, frequently; this instance just happened to be on the ridiculously important topic of privacy -- hopefully you will forgive me for that, I wasn't trying to be malicious but certainly in retrospect I can see this being interpreted as such.. :/
The right option to turn it all off is just this one, since we only send ping data as part of the version update check you disable that and it's all off. And you can confirm this in the code as I just did here[2][3]: https://docs.sourcegraph.com/admin/config/site_config#update... And as I mentioned previously you can always firewall off Sourcegraph 100%.
As an aside, I can promise you that I wouldn't have continued to work at Sourcegraph for the last 5 years if I thought our business was selling or collecting identifiable user data in ANY form. We only collect just enough information to help prioritize what features we improve and (aside from the first admin's email as I noted already above) it is all 100% anonymous and aggregated numbers that we are extremely transparent about[4]. Our person running analytics is also constantly trying to make this more transparent[5] because we all are very security and privacy aware and know the #1 way to convince people to not run software is to make them think you are spying on them or using their data in ways they would not want.
It's obvious to me this should be more clear in our docs, I'm going to forward all of this conversation onto the rest of our team to make sure we improve our docs here.
[1] https://sourcegraph.com/search?q=repo:%5Egithub%5C.com/sourc...
[2] https://sourcegraph.com/github.com/sourcegraph/sourcegraph@f...
[3] https://sourcegraph.com/github.com/sourcegraph/sourcegraph@f...
[4] https://docs.sourcegraph.com/admin/pings
[5] https://github.com/sourcegraph/sourcegraph/pull/8930#issueco...
It's a different work flow, but you simply don't need cs/ when your code base is orders of magnitude smaller
If you already know how to index, this is a completely open source alternative, likely with less bells and whistles.
I worked at Google and miss Code search. But I have lots of ideas as well how one can go beyond the status quo for code reading and debugging. Join if interested.
Chromium also has its code indexed by the older version of this tool: https://cs.chromium.org/
Chromium recently switched over to a new version of code search.
Compare that to the redesign where each xref jump has me staring at a spinner for half a second, the xref bar has no visual separation between type, filename, and code snippet, all buttons visually indistinguishable with a blue on white color scheme, the "layers" dropdown is replaced with a mishmash of buttons scattered across the layout, etc.
I really hope they're not forcing this abhorrent redesign on their developers as well.
You select your binary on borg (think kubernetes/docker), and it'll fetch from the binary with which CL (think like perforce "CL") it was built, and/or additional cherrypicked CL's, then it'll somehow go back in-time and represent how the source code looked then.
later one can (I tried it in Java, but I believe it's available for other languages too), you can inject statements right around the begining of function (a way of breakpoint), and that statement can be something like - let's log how this function was called - you were able to reference nearby statements. This could be set from the command-line, and took a bit mastery (and was bit afraid first time using it, or more like had chilling effect on me), but then my task (with 10 or 11 instances) reported these log lines, and I was able to see them in the browser.
(I have no experience with GCP, or the public face of Google Cloud, so I don't know what's available there), but this was freakin cool.
cs.opensource.google is amazing
https://cs.chromium.org/ is the next best thing.
However this part of internal codesearch is the one part that is actually (partially) open sourced: kythe.io
If you install sourcegraph, you get the same btw. Sourcegraph indexed search is powered by zoekt.
As for UI, treetide/underhood I mention elsewhere is the only open option now.
But Kythe comes with command line utils and an API you can query directly as well.
What is missing from the open source is a production-ready parallel serving table builder. There is one in golang which uses Apache Beam, but last time I checked the go workers are not well supported on the Flink runner. It didn't even work properly on the GCP runner. Hope this would change.
https://cs.opensource.google/kythe/kythe/+/master:external/c...
Note that you'd only be seeing the final result, not the whole process by studying source code. Also, I'd say definition of good code varies by domain.
You may also hover mouse over a symbol to find its declaration and occurrences.
I wonder if they can open source code search itself.
But in reality you probably want something more like SourceGraph which packages everything up nicely so that you don't need to worry about it, or something more specialized.
It’s easy to self install and use, with good documentation with added bonus very fast.
Surely code is easier to explore in an IDE which understands the context and dependencies of the project... This just seems like a glorified "find"
Also, Code Search has baked in a lot of goodies. History layer, cross-references, call sites, ... and it's snappy. Moreover, is really well integrated with all the other internal tools used for coverage, code analysis, issue tracking, web text editor, ... .
I think an IDE (like IntelliJ IDEA) can't reach that level of integration with several other systems unless you fully buy into the ecosystem a company like JetBrain proposes you (their issue tracker, their code review tool, ...).
So, summarizing, it's a tool made by Googlers for Googlers' needs and it's amazing using it every day for all the above reasons.
It should probably be 'Google Code Search'. I would have expected Google to come up with a search engine for all Open Source code otherwise.
For me is as powerful as Bazel, but without the need for a JVM and all the insanity that comes with it in a desktop/dev environment.
The syntax is great, powerful (insane customization) and together with Ninja theres nothing like it.
Its in C++ and even being as powerful as Bazel, its a light, standalone library that can handle a huge amount of source code, dependencies, tools and configurations.
I was working on a big source tree and got frustrated that it kept rebuilding files that hasn't changed just because I switched git branches to look at one file, and then suddenly "Yay, another 18 hour full rebuild!".
I tried to fix it and found there is no option to ignore file timestamps, and some guy has tried to patch it to do that[1]... But the patch requires putting an option in GN files which seems to break them wherever I put it... I tried to patch GN, but it wouldn't ever seem to pass that option through... Ended up patching Ninja to always have the option on, but then random other operations broke (like simple file copies).
A day wasted, and problem not solved. Maybe my use case isn't common, or a bad workman blames his tools, but for me at least it wasn't a nice experience.
Would be interested to know how they expand the bubbles as your cursor moves closer.
Demo: https://demo.repomono.com/cs/view.php
Code is here: https://github.com/repomono/cs
If you want to get universal code search for your own (private) code on any/all code hosts, Sourcegraph is easy to set up internally (self-hosted Docker install) at https://docs.sourcegraph.com/. Or you can get code search for all OSS projects at https://sourcegraph.com/search. More general info at https://about.sourcegraph.com.
Lots of Xooglers and current Googlers use Sourcegraph, too. Just mentioning Sourcegraph because I’ve seen several other folks mention us in the comments (thanks!).
I filed an issue on sourcegraph/sourcegraph, would you mind posting more details there about the errors you ran into? https://github.com/sourcegraph/sourcegraph/issues/8970
Managed && !"@Singleton"
(I'm omitting the fully-qualified class name for brevity)
If I also wanted to look for HealthCheck classes I could update the query to:
(Managed || HealthCheck) && !"@Singleton"
I think it also helps that OpenGrok has a separate input for filtering file paths (completely splitting the "where" and the "what" parts of the query). And this file path search supports the same boolean operators. So if I want to narrow my search to two particular repositories I could put CrmSearch || AutomationPlatform into the File Path input. And because this input only handles file paths, I don't need to remember any special syntax. Whereas if you clump the entire query into a single input, then users need a way to tell you whether a search term applies to file paths or file contents.
There is another open source application for code search opengrok [1] (it's completely open source unlike sourcegraph and supports multiple version controls beside git).
Take a look. It's easy to install and operate on bare metal, cloud and containers, instead of convoluted sourcegraph way of kubernetes or docker.
Another reason not to use sourcegraph is it’s proprietary (with some open source parts), unlike opengrok fully open source.
Thus, Apache 2.0 and some custom license requiring you to accept the terms, have a correct number of seats, and does not allow you to "copy, merge, publish, distribute, sublicense, and/or sell the Software."
I am not sure what all of this means, though. Better check out the licenses yourself :)
One nice-to-have would be support for C, does the C++ extension work for C as well?
Thanks!
Check out the recent release notes: https://about.sourcegraph.com/blog/sourcegraph-3.13#basic-co... And how out of the box code intelligence works: https://docs.sourcegraph.com/user/code_intelligence/basic_co...
Following this article https://bugs.xdavidhu.me/google/2020/03/08/the-unexpected-go... recently featured on HN, I tried
https://cs.opensource.google/search?q=%2F%5E(%3F:(%5B%5E:%2F... : no result
https://sourcegraph.com/search?q=/%5E%28%3F:%28%5B%5E:/%3F%2... : 2 results !
I'm based in France, with a french IP, but my browser language header is set as ACCEPT-LANGUAGE en-US,en;q=0.9,fr-FR;q=0.8,fr;q=0.7
And this page is fully in english.
I didn't say it was a problem for most users ( aka passive users ). I said it's a problem because they make it difficult/impossible for tech/active/advanced users to switch it.
> Websites should never assume they know better which locale their users want than the users themselves.
Yes. You just restated my comment. Not allowing tech/active/advanced users the option is the problem. I love comments that appear to debunked what instead you wrote but just write it in a different way and pretend it is new.
Still?
> and since you got downvoted, evidently a few others felt the same.
This has got to be the saddest thing I've ever seen on a forum.
> Next time you'd probably better write it out explicitly.
Okay I'll give it another shot.
https://news.ycombinator.com/item?id=22563157
Let me know if that cleared things up for you.
heh
in my case I wish it was based on my IP
but for some unbeknownst reason, some components of the page are in Russian. I'm not in Russia; nothing in my browser request indicates I'd like to read Russian.
When I first stood the site up and tested it, Chrome would always break in as soon as it loaded, with a popup to translate the site into _English_ from _Romanian_!
I was able to suppress this only by turning on every single language hint in META.
Lots of people log errors to some sort of monitoring system. I can't remember seeing any localisation/translation API that would log an error rather than just silently serve English. I infer from this that just serving English is universally accepted and considering it an error is so rare that I've yet to see an API that caters to it.
English is basically the world's default language (like it or not). Sites that translate partially, but sometimes show English text instead of the language specified by the browser expect the user to understand the world's default.
A language inferred from geoip is the user's area's default language. Sites that show that language instead of that specified by the browser expect the user to understand the area's default.
These two behaviours seem really quite similar to me. Their technical backgrounds differ, but the resulting behaviour is much the same. One has been widely accepted since the dawn of the web, AFAICT, which leads me to believe that the other has been just as acceptable for just as long. And thus, my answer to "when did it become acceptable" is "it always was, you just didn't notice".
I also wonder, what is happening in places, where people traditionally where always from different language groups. I don't know if then there is always a single common language.
My 'preferences' and settings are a total disaster. I end up having to go onto the gray market to buy gift cards and prepaid credit cards as I seemingly never can buy stuff online when I want to, as I'm either in the wrong place, or in the wrong language. But I know I'm still me.
What is with this '100% of people in this location read/speak the same?'
What if I want to learn Russian, but I'm in China? Why cant I just tell my computer to show me Russian, and the browser tells the site give me Russian if you have it?
Why is this so hard?
I really dislike things that try to make it easy for me, as all they do is prevent me from being able to function.
In the first point, there's someone with a requirements document that assumes every country has one official language and everyone in that country speaks that language, and so feels successful and internationalization-ready when a geo-IP served page is automatically switched to the "correct" language (much like "Falsehoods programmers believe about names").
Second, configuring a computer's locale to set a browser's request headers correctly is beyond the technical expertise of many users. It would be better if things were consistent, but at the point where some locales were set incorrectly and some were uniquely set intentionally your analytics would have showed that you improved the situation on average by trying to guess the locale (screwing over users who knew how to use their computer) than by respecting it and eventually getting everyone to understand how to set their desired language.
i dont know IE but it was in a very good position to guess the language of the user as well.
The reasons I've heard from web developers on why they don't use this is because they believe that the user probably never set that up right, and that multiple people could be using the web browser so they need to be able to do the right thing.
What I typically do is select the best matching language from the Accept-Language HTTP header, and then override it with a session-specific value IF one is supplied. Example:
1. https://rkeene.org/projects/6to4/
2. https://rkeene.org/projects/6to4/?lang=fr
You can see PART of the problem here from the web developers perspective. This isn't a negotiation so you have no way to know which languages the server supports. If your preferences aren't totally inclusive you'll get something "wrong". This can be solved by exposing that information (as not done above) and allowing the user to override it (as above).
[0] https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Ac...
Frankly, it's been the best hourly return on investment of anything I've done in my life up to this point, by far. Assuming I wouldn't have gotten the job otherwise (which seems reasonable), each of those hours spent studying has proven to be worth several tens of thousands of dollars. I'm not exaggerating; I just did the math.
Maybe the interviewing process is broken or sub-optimal or whatever, but it is what it is, and if you can get through it by doing some additional studying, then it's absolutely worth it. Google is a good place to work on designing big systems, so if that's your interest, consider just putting in the work.
IIRC, Google Code Search was the impetus for creating the RE2 library.
People build their businesses on Google products and services like this, not realizing the 90%+ mortality rate over 5 years or so.
But at the end of the day they are an advertising company. Any product that doesn't help them sell advertising -- and lots of it -- will eventually impede the progress of the careers of the managers and employees who work on it. That's when the axe falls... not when a product "fails," necessarily, but when it's no longer "sexy."
People who work there are very well aware of that, and are OK with it. No one outside the company should allow their own business or career path to depend on Waymo, at this stage. It could vanish tomorrow at the whims of the Alphabet execs and/or directors, because their own business doesn't depend on it.
> Have fun using ddg and Firefox
And you've shown yours.