Google’s revamped Cloud Source Repositories in beta
cloud.google.com
cloud.google.com
I've been mucking around with some code in javac recently and was trying to show someone something I'd found, and browsing the openjdk mercurial repo was just so painful compared to hunting around google3.
(As a bonus, I find ddg+Firefox snappier than Google+Chrome)
That made me think, so thank you and yes I agree.
I may transition from my mail server after I've deployed automation for this built atop GCP and the Gmail API.
Even then I'll still make plenty of use of mutt for better supporting OpenPGP, bottom-posting, and true threading - all of these things are still useful in Debian, and I am a Debian developer. But I'll be more easily able to mix usage of Gmail and mutt, without some hacks I currently use to achieve that.
Well, permanent until they kill the entire service 5 years from now.
(Though honestly, the value from the internal code search was that everything was in there. So if you got an error message from a service you were a client of, you could just search for the error message in their code and see what it actually meant. "Aha, this 'optional' field that's marked 'deprecated' isn't actually optional or deprecated," most of the time ;)
Correct. Looks like it just parses and indexes the source.
The code search functionality is the same as what Google engineers use internally. Cloud Source Repositories uses the same indexing and retrieval technologies on the same type of infrastructure.
As you mention, the more code you have, the more benefit you get from having fast search tools across your entire code base that can perform complex semantic and regular expression queries. Even with smaller codebases, I find it the fastest way to find the code I need.
(Disclaimer: I work at Google and am the PM on this product)
Conversely, I think it's important for Google engineers to be able to try new things and publish products without having to worry about supporting them forever. Is there some sort of happy medium? Like a Google labs of products that are in various states of whimsical testing?
Cloud source looks nice, and for GCP projects I'm sure it will have a nice tie in, but for now I'm 100% more likely to push code to Gitlab until it makes sense to do otherwise.
[1] https://en.wikipedia.org/wiki/List_of_Google_products#Discon...
They actually had precisely that, but then they cancelled it.
> Google described Google Labs as "a playground where our more adventurous users can play around with prototypes of some of our wild and crazy ideas and offer feedback directly to the engineers who developed them."
Has anyone played with this or similar tools that can provide search by indexing at source. It’s a little surprising that Google doesn’t just index everything public in Github, gitlab, etc.
I would like a good code search solution especially because I have projects across lots of different repo servers but want developers to be able to find code regardless of repo home. A colleague indexed it with solr and that was ok, but had limited semantic ability and definitely nothing like what Google describes where your recent search history and activity is displayed to each user.
GitHub excludes a few patterns in robots.txt that means a lot of the public data is never seen by crawlers that respect robots.txt, for example:
Disallow: /*/*/tree/*
Disallow: /*/*/blob/*
and Disallow: /*/*/commits/*/*
Disallow: /*/*/commits/*?author
Disallow: /*/*/commits/*?path
Disallow: /*/*/branches
Disallow: /*/*/tags
https://github.com/robots.txtThey probably want people landing mainly on
- user profiles
- repo roots
- individual issues
- individual pull requests
And I am glad they do, because when I search on Google, the above mentioned kinds of results are exactly what I am looking for, not individual source files etc.
Thanks for taking a look. I'm glad you find Cloud Source Repositories interesting. We'd love to Cloud Source Repositories more practical to use. Can you explain what is impractical for you? Is the initial process of setting up multiple mirrors? Or the routine of having to go to another site for search?
(Disclaimer: I work at Google and am the PM on this product)
I have no easy insight or overview into code across these repos without scripting. The impractical part is in cloning mirroring all of these repos to get insight from your new function because user behavior is baked into all the workflow around the Github repos.
The go to another site isn’t so bad as this is a new function and not available in Github, so programmers will likely check it out when they need.
But manually forming everything and syncing. It would be neat if I can import an org in one fell swoop.
That being said, I do have scripting forking everything over on my list of things to do when I have cycles on the weekend.
Apparently they forgot about Flutter users.
Cloud Source Repositories currently supports very large repositories. The same backend scales to the needs of the Android open source project, which regularly checks in massive binary files. There are many APKs and VM images in the Git repositories we host. However, Cloud Source Repositories does not support LFS or the mirroring of LFS content.
LFS is not deeply integrated into Git, creating usability problems. Because the content is not part of the object graph of the repository, you have to decide in advance to use Git LFS. You don't get the benefit in existing repositories with large files. You also can't back out --- once you're using Git LFS on a medium-sized file, you can't change your mind and instruct Git to send it inline in fetches, without rewriting history, which breaks existing clients that have cloned the repository.
For the same reason, history mining commands like "git blame" and "git log -S" don't have access to the object.
In addition, it complicates migrating to another host. Usually in Git, you can take out your content by running "git clone --mirror one-url and then "git -C directory.git push --mirror another-url". With Git LFS, this copies over the pointer files but not the underlying large file content and you must remember to take extra steps to instruct the new host about where the blobs are stored.
Google's Cloud Source Repositories team believes that Git itself needs to deal better with large files. The first step of this work within the Git project has been partial clone, which is supported in Cloud Source Repositories and in the public Git 2.17 release (for best results, please use Git 2.19, released September, 2018). If you run "git clone --filter=blob:limit=512M <url>", files larger than 512M will be omitted from the initial clone and fetched automatically on demand when needed (for example during checkout operations). See https://crbug.com/git/2 for more details about this feature. We are continuing to work with the community on adding other features related to large file support into Git.
(Disclaimer: I work at Google and am the PM on this product)
We've generally heard that developers move from Bitbucket to Cloud Source Repositories because they want to use other Google Cloud services and it's simpler for them to manage their source in the same place where they debug, build and deploy that code. Developers often mention that they appreciate the unified identity, IAM permissions, and integrations that Cloud Source Repositories has with Cloud Shell, Cloud Build, Cloud Debugger and others Google Cloud services.
However, developers don't need to move from Bitbucket to take advantage of the utility provided by Cloud Source Repositories. You can mirror any number of your repositories into Cloud Source Repositories and take advantage of the code browser, code search, and various integrations with Google Cloud Platform.
(Disclaimer: I work at Google and am the PM on this product)
I'm glad you like it. Which languages would you like to see support added for?
Note that search works today across all languages but the semantic understanding of source which enhances search is limited to Java, JavaScript, Go, C++, Python, TypeScript and Proto files.
(Disclaimer: I work at Google and am the PM on this product)
Can't think of others off the top of my head, so that is two if we want to consider major upgrades "abandoned." Perception often doesn't match objective fact.
(Disclosure: I work at GCP)