An ex-Googler's guide to dev tools (2020)
about.sourcegraph.com
about.sourcegraph.com
Original discussion: https://news.ycombinator.com/item?id=25217291
1. Code search is more mainstream. It used to be most of Sourcegraph's sales came from ex-Googlers missing code search, but nowadays, most of our new revenue originates from dev teams just struggling with the challenge of big codebases.
2. Investment in dev tools has only accelerated. A lot of pandemic beneficiaries (virtual events, tele-medicine, at-home exercise equipment, crypto) are now reversing, but dev tools growth continues to be strong.
3. Bazel adoption has grown. Were I writing the post today, I would have had a category for monorepo tooling that included companies like EngFlow and Turborepo.
4. Kubernetes is still dominant for container orchestration, but serverless and managed services seem to be gaining more adoption. Classic tension between customer focus and open standards—will be interesting to see how we trade off dev-ex vs proprietary platform dependency.
Since this post was first published, we've also made a lot of advances at Sourcegraph and are looking for people who want to play around with some new experimental search and code intel features. Come say hi in our Discord! https://discord.gg/SSCBGByJeu
[1] https://admin.phacility.com/phame/post/view/11/phacility_is_...
Shameless plug but anyone interested in a code review tool on top of GitHub may also want to check out CodeApprove (https://codeapprove.com). It's got a lot of the things that people from FAANG companies, especially Xooglers, miss about the excellent internal code review tools those companies have.
> Join the waitlist
:-(
Since you seem to be in the beta, you may want to let them know their Homebrew tap is pointing to the wrong GH org: https://github.com/withgraphite/homebrew-tap/blob/master/For...
There is also a (blessed) community fork at https://we.phorge.it/
* Consistent build structure
* Rigorous code reviews
* Ginormous company-wide shared filesystem holding all code
* Point-click-drag report building
* Big-query
I'm sure I'm missing additional awesome internal aspects of being a developer at Google, and some parts are also horrendously bad.
- BigQuery is available on GCP
- Bazel is an open source Blaze
- google source formatting (for java at least) is open source
There are probably more…
There's a reason it isn't popular outside of elgoog.
IRL, every project tends to do a little bit or a lot of its own thing.
If you love fighting with your hands tied behind your back, choose Bazel.
Otherwise, be pragmatic: Learn Make, Maven, and Gradle; then you'll be well-equipped for 95-99% of cases. Thankfully pip and npm are as straightforward as it gets.
I also appreciate what Python and Javascript offer, there are some amazing libraries and tools tied to those ecosystems.
What an odd sprinkling of something entirely personal.
> Otherwise, be pragmatic: Learn Make, Maven, and Gradle; then you'll be well-equipped for 95-99% of cases.
There's a time and a place for Bazel. Very large monorepos like those at Pinterest and Uber, with cross dependencies, and written in multiple languages benefit a lot from the remote backend and distributed cache of built artifacts.
Make, Maven, and Gradle, even only for JVM based projects seem to not be entirely comparable.
Things like engineers losing time wondering why their node dependencies weren't correctly installed, or dealing with a pre-commit check that reminds them they didn't manually regenerate the generated files, or having humans write machine-friendly configuration that's not actually human-friendly because there's no easy way to introduce custom file transformations during the build.
Bazel doesn't spark joy for me and I wouldn't say I look forward to using it, but personally I would still always choose it for a codebase that's going to have multiple developers and last a long time. It's vastly easier to go with bazel from the beginning than to wish you could switch to it and realize people have already introduced a million circular dependencies and it's going to be a multi-month or multi-year process to migrate to it.
And it's clunky, and you need to teach every new hire how to use it. The juice just isn't worth the squeeze. Just write an adapter in your build infra for the well known tools and be done with it. You'll get much more value putting work into something else, like code review tools, testing, dev environments, staging...
I wasn't familiar with EngFlow; it looks like it's a startup that raised seed funding less than a year ago. I think what you're referencing with "Google RBE" is a Google Cloud project now rebranded to "Cloud Builds" — which supports the underlying native tools like npm, pip, etc without requiring you to switch to Bazel.
Bazel is better than Make for building large C/C++ projects (although it's hardly the only game in town for "better than Make"). But aside from that use case, in my experience it's not really worth the hassle. You can get most of the benefits you want without using it, and people are already going to be familiar with the tools native to the ecosystems they work in like pip, npm, etc.
It depends on what you're doing. If you're compiling pure go code that's truly only go, no go build will be faster. But if you have cgo or generated code or generated data or config files or...then well maybe you want something that is more flexible. And of course if you aren't building just go, then things get complicated fast. What if you have a backend and a frontend? `go test` probably isn't running your webdriver tests that compile your typescript somewhere in the build pipeline. Having a unified toolchain with one command (`blaze test //...`) is valuable compared to various test.sh (or make layered on top of n independent build systems or...)
And of course if you're like me and need to do things that involve reasoning or searching about code and dependencies, blaze is super necessary. "Find all of the dependencies of this file and test only those" is a question that most build tools aren't even remotely equipped to answer.
So in polyglot situations, bazel and similar prevail I think, but there's absolutely a point below which you don't care (and that point is going to be basically only hit by applications, not libraries).
Wether or not that sandbox blocks incorrectly declared dependencies is unclear. Last time I used clearcase many eons ago it surely did not. Our project had tons of classic makefiles issues like not depending on included headers. Remote builds were also magnitudes slower than local builds, our network was maybe not the best, also reading the page above how the “shopping” algorithm works you can imagine it being fairly slow anyway. Maybe that was best, imagining how incorrect dependencies mixed with remote caching would result gives me nightmares.
Here's a real-word example from last week: I spent several hours doing creative hackery trying to convince a particular compiler to compile the same file using different dependencies depending on context. That would have been a few minutes of work with bazel, but when your build process consists of "run this off the shelf compiler" and the compiler doesn't have any native support for building two different versions of the same thing with slightly different dependencies then you're in trouble.
Teaching new hires bazel is a one-time cost, and the only thing they'll need to get started on their first day is "bazel build <target>" and "bazel test <target>". When you don't have bazel, on your first day every new hire has to read through a gigantic wiki page explaining how to set up their dev environment (with different sections for mac and linux and special asides describing what to do if you're on a slightly outdated version of the OS, and then the wiki gets out of date and you have new engineers and the infra team all wasting their time debugging why the instructions suddenly stopped working for people on slightly newer machines, etc.)
Look at how much pain many companies have gone through to move to K8s from existing infrastructure, there is perceived value driving that.
Bazel lacks that perception of value.
Every real-world build system I’ve seen provides that functionality. In particular a standard make(1) can do it just fine.
If I could have a wish to upgrade file transformation with make(1), I'd probably want a widely-available, standard, simple command to make a rule-specific virtual filesystem that overlays a configurable read-only/copy-on-write view of selected existing files or directories behind a writable rule-specific output directory.
Basel is as non-standard as it gets - essentially yet another Google’s case of NIH - but apart from that, how is an ad-hoc single-use pseudo-language better than reusing standard mechanism? To me it’s just a bad design.
That's just an argument from tradition.
And you want sandboxing because that's what gets you good caching. The value of bazel is never having to run make clean because artifacts aren't correctly being built from cache. Having no distinction between clean and incremental builds is really nice.
>The value of bazel is never having to run make clean because artifacts aren't correctly being built from cache.
You can get that with make(1) too, check out FreeBSD's META_MODE for one example. And it didn't require reinventing the wheel.
I didn't do any such thing. My point is simply that make and bazel are similarly "nonstandard single use pseudo languages". In many ways, bazel is superior to make from a language perspective (it resembles other languages more closely, being a dialect of python, and avoids the loadbearing tab issue), so I think I could make the argument that bazel is in many ways less nonstandard, but make is certainly more common than bazel, so it could go either way.
> You can get that with make(1) too, check out FreeBSD's META_MODE for one example.
This suffers from the same issues that natural make does (notably the whole mtime thing). See https://apenwarr.ca/log/20181113 for a much better explanation than I can provide as to why make's entire model of "change" is irreparably broken, and why hash (+sandbox!) based approaches (which bazel and redo and nix and cargo and nearly every other modern build tool use) are far superior.
> And it didn't require reinventing the wheel.
You call inventing a new syscall to not-even-fully fix a limitation of the tool not reinventing the wheel? Like I guess its not, it's just like building a weird grand shrine around the broken wheel. It's far worse. I don't want to need to change my operating system to have make work better, but still worse than the alternatives. That's simply not a compelling argument.
If someone wanted to make a new build system to compete with bazel and have those kinds of features, it's probably a safe bet the competing system would use some kind of sandboxing as well
Even if you ignore everything else, just the security part is a big deal: supply chain attacks are an increasingly big concern for companies of all sizes. If your build system allows any script invoked during any part of build process to secretly read or modify any input or output file, hackers are going to love it.
Almost all tech companies (even the multi-billion dollars ones) that aren't doing something in the spirit of `bazel build` to generate their binaries have wide open, planet-sized security holes in their build systems where if you get one foot in the door you can pretty much do anything.
Most of industry still seems unaware that no-knobs data query and pipeline systems even exist. If I only had a dollar for every time I saw a PR tweaking the memory settings of some Spark job or hive query that stopped running as the input data grew....
I'd love to see more people write their workflows using the Apache Beam API so they'll have the option to switch to a no-knobs, scalable pipeline engine in the future even if they're not using one today.
* Automated presubmit/bisect queues
* Well built code review ui and not like github
On this topic, most companies will give up on “rigorous code reviews”.
I’d love a tool that helped with context in a larger system design way. Coming in cold on some project and expecting to know all the design decisions, and if the code structure is going away from some higher level goal is where most problems that I’ve seen come in.
I don’t see how any of the current tools help with that, and I somewhat doubt these secret tools do either - but would love to seen them if they do.
To pull an example from the paper, if your line of code says
if (a == b || b == 1)
you might get a comment that says something like> "Changing this line to
if (a != b || b == 1)
> does not cause any test to fail."Page 4 lists other mutations, like replacing logical ANDs/ORs with just `true` or `false`, switching arithmetic plus to minus, etc.
This isn't a larger system design context thing, though, just a testing one.
Disclosure: I work at Google.
I'm also guessing this ties into Google's tooling of being able to tell which line of code is hit by which tests and running those as part of the mutation test?
EDIT: The blog has some discussion about the correlation between some of these mutation failures and real bugs. I remain a little suspicious to be honest. Also
"I also looked into the developer behavior changes after using mutation testing on a project for longer periods of time, and discovered that projects that use mutation testing get more tests over time, as developers get exposed to more and more mutants. Not only do developers write more test cases, but those test cases are more effective in killing mutants: less and less mutants get reported over time. I noticed this from personal experience too: when writing unit tests, I would see where I cut some corners in the tests, and anticipated the mutant. Now I just add the missing test cases, rather than facing a mutant in my Code review, and I rarely see mutants these days, as I’ve learned to anticipate and preempt them."
Now this one is expected, if you build automated tooling that comments on your code review you expect people to try and avoid that. However there's still the question of what's the quality improvement and what's the productivity impact.
if (a == 10)
With if (true)
Then you're not testing the behavior of your application. If in production `a != 10` and you've never tested that it might be a problem.Mutation testing can also do stuff like removing entire statements, etc.
The next question would be let's say I spend my time writing the tests to resolve this (could be a lot of work) is that time better spent vs. other things I could be doing? (i.e. what's the ROI)
Even ignoring that is there data to support that the quality of software where mutation testing was added improved measurably (e.g. less bugs files against the deployed product, better uptime, etc?)
Is this method better than just looking at code coverage? Possibly none of the tests enter the if statement at all?
EDIT: where I'm coming from is that it's not a given this is an improvement to the software development process. It seems like there was some experimentation around the validity of this method which is good but like a lot of software studies somewhat limited. It also seems there's a lot of heuristics based on user feedback, which is also good I guess, but presumably also somewhat biased.
The related paper has a lot of details including: "Since the ultimate goal is not just to write tests for mutants, but to prevent real bugs, we investigated a dataset of high-priority bugs and analyzed mutants before and after the fix with an experimental rig of our mutation testing system. In 70% of cases, a bug is coupled with a mutant that, had it been reported during code review, could have prevented the introduction of that bug."
Which should imply(???) that 70% of "high priority" bugs can be eliminated during the review process by using this sort of mutation testing. Seeing data to that effect would be cool (i.e. after the fact) and if it's real that'd be pretty incredible and we should all be doing that.
It's not the lack of a test, but the fact that the existing test doesn't cover a branch.
Rarely taken paths is exactly where bugs often hide, because they likely weren't well exercised in human-driven testing either, and because they're hit rarely in general so most users don't run into them often. If the condition is not easily reproducible it can be very hard to figure out what to even put in a bug report.
Say you're working with images and have a branch for a 16 bit color (565) image. Such images are rarely used today, but they still exist. Among many users of your code it's possible only 1% ever hit that branch -- and those end up experiencing weird issues nobody else sees.
Or another example of such things is error handling. An application that works perfectly fine on a LAN can be hell to use on a flaky connection.
> is why would we suspect the lack of a test indicates a bug?
I can only speak for my experience but the code is not better because it is mutation tested. It is better because we have thought about all of the edge cases that could happen when inputting data into the system.
Mutation testing, as a tool, helps you find statements that are not being exercised when parsing certain data. For example if I write an HTML parse and I only ever provide test data that looks like `<a href=....` as an input string and a mutation testing tool replaces:
if (attrs.has("href"))
return LINK;
with: if (true)
return LINK;
It is clear to a human reader that this conditional is important but the test system doesn't have viability in this. This means in the following situations you can be screwed:1. Someone (on the team, off the team) makes a code change and doesn't fully understand the implications of their change. They see that the tests pass if they always `return LINK;`.
2. If you are writing a state machine (parser, etc) it helps you think of cases which are not being tested (no assertion that you can arrive at a state).
3. It helps you find out if your tests are Volkswagening. For example if you replace:
for (int i = 0; i < LENGTH; i++)
with: for (int i = 0; i < LENGTH; i += 10)
Then it is clear that the behavior of this for loop is either not important or not being tested. This could mean that the tests that you do have are not useful and can be deleted.> For most non-trivial software the possible state-space is enormous and we generally don't/can't test all of it. So "not testing the (full) behaviour of your application is the default for any test strategy", if we could we wouldn't have bugs... Last I checked most software (including Google's) has plenty of bugs.
I have also used (setup, fixed findings) using https://google.github.io/clusterfuzz/ which uses coverage + properties to find bugs in the way C++ code handles pointers and other things.
> The next question would be let's say I spend my time writing the tests to resolve this (could be a lot of work) is that time better spent vs. other things I could be doing? (i.e. what's the ROI)
That is something that will depend largely on the team and the code you are on. If you are in experimental code that isn't in production, is there value to this? Likely not. If you are writing code that if it fails to parse some data correctly you'll have a huge headache trying to fix it? Likely yes.
The SRE workbook goes over making these calculations.
> Even ignoring that is there data to support that the quality of software where mutation testing was added improved measurably (e.g. less bugs files against the deployed product, better uptime, etc?)
I know that there are studies that show that tests reduce bugs but I do not know of studies that say that higher test coverage reduces bugs.
The goal of mutation testing isn't to drive up coverage though. It is to find out what cases are not being exercised and evaluating if they will cause a problem. For example mutation testing tools have picked up cases like this:
if (debug) print("Got here!");
Alerting on this if statement is basically useless and it can be ignored.> Is this method better than just looking at code coverage? Possibly none of the tests enter the if statement at all?
Coverage does not tell you what the same thing as what mutation tests tell you. Coverage tells you if a line was hit. Mutation tests tell you if the conditions that got you there were appropriately exercised.
For example:
if (a.length > 10 && b.length < 2)
If your tests enter this if statement and also pass when when replaced with: if (a.length > 10 && true)
Or: if (true || b.length < 2)
You would still have the same line coverage. You would still have the same branch coverage. But, if these tests pass, it is clear that you are not exercising cases where a.length <= 10 or b.length >= 10.> where I'm coming from is that it's not a given this is an improvement to the software development process
In my experience if I didn't write a test covering it, it was likely because I didn't think of that edge case while writing the code. If I didn't think of that edge case while writing the code then I am leaning heavily on defensive programming practices I have developed but which are not bulletproof. Instead of hoping that I am a good programmer 100% of the time and never make mistakes I can instead write tests to validate assumptions.
> Seeing data to that effect would be cool (i.e. after the fact) and if it's real that'd be pretty incredible and we should all be doing that.
Getting this kind of data out of various companies might be challenging.
100% test coverage is the most easiest goal to achieve (in python). It's pretty much dumb activity. To write proper tests to check BL is a much harder task.
From my experience it gives the biggest value for dynamic languages where you must be sure that every line is touched before releasing code.
Formatting can be almost entirely automated with pre-submit checks that the code is formatted, and automatic formatting tools.
So for your two examples, this practice of adopting a style guide prevents them from arising. The “tabs vs spaces” argument is simply settled, because the style guide has an answer. The “for loop vs for-each loop” is either addressed by the style guide, in which case there is an answer, or it’s not, in which case it is not a legitimate style complaint one way or the other, and the reviewer is making a mistake by bringing it up.
I was just trying to flowerly highlight that most “in depth code reviews” I’ve seen are very focused on surface level things. I’d dig a tool that focused more on structural changes.
Discussing these deep topics is also socially very challenging and you need some quite grown up people to discuss them (that late after the damage has been done) without awkward "but I spent all this time" moments and at least some personal animosity. But the grown ups think and discuss first typically, which also might be why you don't see them discuss those topics in PRs...
Fair point, and 100% on the nose. The times when this comes in handy is when you have to PR on a project you don’t know a lot about (which could be argued is a process problem, but I am getting way off trail now).
However, with your excellent point in mind, I am even more confused on what a “culture of rigorous code reviews” is.
Google also has documentation, guidance, and training on how to review code effectively.
Basically: reviews are actually treated as important.
Shameless plug for https://trunk.io/products/check - which will handle universal enforcement of all the tooling for all of the pieces of your tech stack.
1. Search another team's code
2. Quickly jump into a Cloud IDE to make a change
3. Build and test the code, because everyone uses one build system
4. Send that code off for review and have the review tool automatically select the right reviewers and run the right tests
For a simple change (let's say a typo fix) all of the above can happen in a single browser tab in under 5 minutes even if you've never been in that part of the code before.
----
While I'm here a shameless but relevant plug: if you're looking for a Google-like code review experience on GitHub check out https://codeapprove.com
At Coherence, we agree that code search and code review, along with automated code intelligence, are a huge step forward. Tools like sourcegraph, graphite, codeapprove are all solving important problems in the cycle. But we believe that the general approach of “building” vs “buying” development infrastructure for these key steps is holding teams back.
Check us out at withcoherence.com if these sound like problems your team faces, too! (Disclosure, I’m a cofounder)
I really hate dependency on GitHub for workflows but Microsoft is nailing the dev tools people are looking for and I think taking a look at what happens inside of big tech dev tooling is always interesting to see what we can build in the wild to improve developers ergonomics.
You also don’t have problems to solve like google’s. Use the tooling appropriate to the task.
I find this problem with non-programmer ex googlers as well. They are so used to having so much infrastructure and opaque/screwed up objectives they’ve been given to meet that they often struggle in a startup environment (where the goals are merely “let’s get this button to show up 100% of the time,”, “let’s store the data for our 10k customers”, or “let’s give our phone number to these key customers so if there’s a problem they can call us right away”, or “figure out the most important thing that’s keeping our sign ups from converting and fix it — this week”).
I remember my gf working for a company of about 500 ppl (so not a startup) who, shortly after she joined brought in a new CEO, an ex googler. In fact it had been the only place he’d worked.
When I met him it was clear he 1 - was quite smart and 2 - had no idea how the company worked or why ppl bought its product. And 16 months later, he was gone, along with the ppl he’d brought into the exec team, and the company was losing even more money.
But by that point my gf was long gone.
I have heard certain large companies having very a fragmented development culture (every team does their own thing, no central repository), so maybe Google's tools being particularly refined is the result of most of the company sharing the same tools.
Amazons approach is to let each team use their own combo of tools, all centered around the Pipelines workflow, while G has one set of tools and everyone uses it the same way (comparatively).
Source: Googler but still writing lots of code outside of Google and constantly think to myself "why is this so needlessly hard? Why does nothing integrate well?"
I'm a strong apologist for Bazel, but I can absolutely see why projects avoid it. The more you deviate from what Google does the more paper cuts you'll get, but the less work you'll have to do. Bazel is great but it's certainly not a clear win in the general case.
I am looking forward to new set of bazel rules being worked on for eg. https://github.com/aspect-build/rules_js and https://github.com/jvolkman/rules_pycross which will makes it more idiomatic to work with existing language ecosystems.
- Blaze defaults to building everything remotely, on a massive farm of build servers. (This farm was sizeable enough to have its own capacity planning teams. Yes, plural.) There are a few startups trying to clone this SaaS for Bazel, but it is - predictably - not easy.
- Since all builds were remote, build _telemetry_ was also uploaded and available on the intranet. This meant that if you wanted to ask someone "hey, my build failed, can you help", you could literally send them a link "https://google-build-system/00000-11111-22222-the-rest-of-my..." with all the information about your build: the thing(s) you were trying to build, the options you were building with, all the foreign keys you needed to look up build timing data, etc. You could also query company-wide build telemetry data to see what people were building more frequently than others, if specific areas needed engineering investment, etc.
- Since builds were remote, they pretty much had to be hermetic and reproducible, which meant platform teams - e.g. C++ team, Java team, Python team, Go team, etc. - could make infra-level changes and test out the effects on various teams. This is a very Google-scale specific need: allow another team to do your infra maintenance for you, i.e. upgrade the Python runtime underneath you. (Yes, this comes with caveats. Google eng tends to live at head, so while there are "my team needs to control our rollout to the latest infra version" concerns, there is a consistent mandate for most teams that "upgrading to the latest infra that we run on" is a priority.)
The big place where problems came up was when source code needed to exist _outside_ of google3: open-source repos, for instance, very frequently have to do a delicate dance of figuring out (1) is their primary source of truth going to be inside google3, where the tools/infra are uniform and Better TM, or on github/etc, where they have to configure their own test stack, CI stack, etc., and (2) what are the downstream implications of decision (1).
So, I was trying to find a URL shortnener which allowed for pretty links similar to what the internal “go/“ site supports.
I couldn’t find anything, and it seemed simple enough so I made it:
I miss the functionality a lot for a number of not malicious use cases, but I don't see a way to avoid that problem at the moment.
I think arguably Facebook/Meta has them beat.
Blaze is cool, but do you want to do it for every small project? It's actually a bit annoying for small projects, the verbosity of it. Granted that it helps optimise a lot of downstream workflows and enables pretty good tooling, but you wouldn't need those in startups not operating at that scale.
Startups need speed over accuracy. Google needs accuracy over speed.
Source: Acquired from startup to Google.
Related, but from distant past: https://news.ycombinator.com/item?id=15722849
@Dang: Can we add a feature to auto link related threads in an expand\collapse section?
But yeah, for the day to day java I was writing IntelliJ was way better.
There’s a special place for people who use it to be misleading when advertising their products. (A very popular iOS app dev did this many years ago and I’ll never use their app out of principle.)
And I think Brittanica out ranking Wikipedia was the final straw for normal ad-hoc search.
2) For awhile Google wanted to move more new development onto GCP (an effort called "Google on Google"), but it had limited traction and has been abandoned AFAIK, because internal tooling, infrastructure, and devex was just so much better than what GCP had to offer.
2) feels like this just proves my point
Rather than force the teams who developed the original tooling inside Google to develop a similar public facing commercial product, they had GCP developers write parallel tools, or they bought third party tools and tried to integrate them. Both solutions proved not as easy as management had hoped.
A ham sandwich with a million buttons that each do different things, sure, but it’s not impossible to learn the ones you need, and they rarely break. How many times have we seen apps fail one of those criteria?
It also has the distinction of printing a massive amount of money and being a huge competitive advantage over other ML institutions — like, say, Facebook. So it seems obvious why developer experience wasn’t prioritized, just “ok.”
But using it, I feel deeply in my gut that there has to be a better way. But I've not seen anything across any of the major clouds that is much better.
Though, to be fair, I did once purchase an ex-parrot.
6. forming nouns (from titles of office, status, etc.) expressing a former state.
Former states can be current states, too!