This video was hugely influential on changing the way Google does internal tools and operations.
This video was hugely influential on changing the way Google does internal tools and operations.
The employees set up TeamViewer on their "dev" server and promised the contractors in writing that this all had sufficient permission, this was a dev server, and the credentials on the dev server were not going to get them into anything else that might be troublesome by mistake.
The last of those three statements was accurate. As you might imagine, TeamViewer on a nominally tightly controlled network was not even in the same hemisphere as acceptable...
...and while debugging, the contractors made an incompatible DB schema change on dev to see if it fixed something, only to get a nasty surprise when the employees ran to them within an hour or so asking what they had done, because by "dev" they meant "prod".
I don't really have a great moral for this story, other than maybe "netsec/infosec teams need to actually work with other teams and not just be opaque sources of fiats, or people are going to work around them to get things done instead of trying to work with them, and that's going to end poorly for everyone."
I know a Google contractor who was denied access to a lactation room at her office (she's the type of contractor who works full-time alongside regular employees). When she tried to fix that, she got sent into an infinite loop of support tickets being bounced around between different departments claiming it was the other one's responsibility. Literally no one was able to fix that for her. Her manager was on a different continent, and wasn't able to help either.
If they don't, they call it "comfy job with awesome paycheck and not a lot of pressure" :-p
I think the technical term is Golden cage :)
lots and lots of money
But dropping Borgmon readability was the most immediate and obvious. It was basically true that no one had Borgmon readability. The policy was a catch-22: you couldn't get readability for the simple/formulaic Borgmon macro invocations that were encouraged and often sufficient. You could only get it for doing something "clever". I got it by writing fancy borgmon rules to paper over a problem that (in hindsight) I should have solved elsewhere.
Another was easing quota management. IMHO the most unbelievable thing in the video was that after Broccoli Man told Panda Woman to get quota in two cells, she just said "done". Besides the hassle in transcribing what you needed into the request system [1], various types of quota were chronically unavailable where you needed them, even in tiny amounts. In 2010, I kept a critical infrastructure service running by regularly IMing major clients' on-calls asking them to donate 0.1 cpu(!) of their quota in some cell or another when I didn't have quite enough to grow. There was a "gray market" mailing list where people would trade resources they couldn't get through the primary system. But eventually, they built a system that for small services would make the quota just happen for you.
Overall, it was a kick in the pants for the most basic infrastructure teams that made them see how unnecessarily hard this is for their internal customers, prompting them to make small things just happen while keeping large things possible. In any large organization, it's healthy to get this kind of feedback regularly. The actual specific changes and technologies are pretty specific to Google in 2010...
[1] Many people managed this very very tediously with spreadsheets. I eventually wrote a tool to generate the requests based on comparing your intended production config with your current quota.
I'm only half joking about that too. The other half is that sometimes it is better to let the system collapse under its own weight when that's what is required to convince the right people the system is broken.
By the way, do you happen to be the jwb that filed a... creative BT quota request?
AAAAA+ GREAT CUSTOMER WOULD SELL TO AGAIN
Now over multiple changes, it used to require one fairly big one. It's still a pain in the languages that require it--which is all the main ones, but very few of the niche ones.
Many other things changed as well. Much of what the video complains about got automated and better documented. But the company has grown so much, and the product lines have diversified so dramatically, that there are still plenty of places to complain about the overhead.
borgmon was a truly weird system.
Putting people in charge of formatting/style is just an excuse for wasting time bikeshedding, either the code is wrong and a tool can tell you, or its not wrong.
It's more about good usage of idiomatic language constructs, which still requires good human judgement to evaluate.
Note: I'm not a Googler nor a Xoogler, but my partner worked at Google for some years and so I've heard her complain a lot about Google and readability.
FWIW, these can be the same person - if the have appropriate readabilities. The problem is that it's not easy to get readability in the first place, so occasionally teams have to look for people outside.
And that's an acceptable trade-off for a different sized company. Google arguably has the biggest centralised codebase in the world, and simply has different requirements.
As other have said, most of the overhead is from attaining readability, not the requirement itself.
AWS does not operate this way nor do many CDNs operate this way nor do ISPs operate this way. There's other high scale businesses out there. Google isn't the only one. Using Google's scale to justify business practices is a self-fulfilling prophecy.
I'm not denying that, but all the same, we're not talking about thousands of separate codebases. It's a single monorepo codebase with consistent styling, the advantage being that people can switch teams or contribute to other Google software under the same style guide.
The overhead isn't a bug, it's a core feature as consistent style makes it that much easier to switch teams.
Ever heard of COBOL?
Because readability has always been in the eye of the beholder, and codifying it makes it even worse.
It was studied for multiple programming languages Google uses.
It was done, at my request, by the engineering productivity research team, who are experts in this kind of research - you can find public papers they publish[1].
For background: I was the person responsible for production programming languages, and I did this precisely because i did not feel at the time there were recent good studies as to whether readability was really worth the cost.
The answer is "yes, it is".
There is an upfront cost, but readability has a meaningful and net positive effect on engineering velocity.
It is large enough to make the cost back quite quickly.
One could go down the rabbit hole of seeing whether you improve (or make worse) the numbers by changing various parts of the style guides, but like I said, what is there now, has in fact been studied scientifically.
Honestly, i'm not sure why you think it wouldn't be. You come off as a little immature when you sort of just assume people have no idea what they are doing and don't think these things through. You are talking about a company with 60k+ engineers. Every hour of all-engineer time you waste is equivalent to having ~30 SWE do nothing for a year.
Maybe there are companies happy to do that. I worked at IBM for a few years when i was much younger ;). All I can say is that as long as i'm involved in developer tools at Google, i'm gonna try not to waste my customers time.
[1] I say this in the hopes you don't go making assumptions about whether the studies were done properly. They were properly controlled for tenure, number of reviews, change size, etc.
I don't know of any research off hand, but I'm pretty sure the industry consensus is that good identifier names improve the quality of the code (Go style notwithstanding.) Readability is one way to training engineers to do it.
This is empirically false. Consistency, even if it is unfavorable to your preferences, is superior to inconsistency. So a codified set of best practices is better than none at all.
There are part of Google's style guides that I would change if I could, but I also prefer having a style guide (and one that goes beyond things that are lintable) than none at all, because consistency across the codebase means that I can usually understand code at a glance, or if not, know at a glance that something unusual is happening. (this is in fact precisely the argument in favor of autoformatters like gofmt/black/prettier, but extended to softer concepts that can't always be formatted: consistent style, even if it isn't your favorite, is superior to inconsistent style).
Indeed, but at what cost? Every step of added gatekeeping reduces development velocity, and Google does not have a monopoly on high scale services (cough AWS cough). In an ideal world for every full-time dev writing code, you'd have a full-time reviewer whose job it is to simply review that dev's code. But the real world has budgets and deadlines, and humans become discouraged when changes take too long to complete. So is readability worth the added gatekeeping cost is the question.
I'm a bit of a polyglot at Google (I have readability in C++, python, and Javascript/Typescript, and am part-way through the Go and Java processes, which covers basically every popular language at the company), and IMO, the readability process for each language has been a net positive, although it was painful for a while when I had no readability and none of my coworkers did (but there are procedures for getting around that today).
For example, I "donate" my readability, and review random changes from people who don't have readability but need a quick turnaround or aren't yet working to gain readability.
I'm not. I'm not saying that code should be merged without any stylistic commentary; I'm simply proposing that stylistic commentary be provided by other stakeholders in the service you're contributing to, on top of whatever linter or sanitizer is being used. That is, after all, the goal of a human reviewer. To add _another_ layer of gatekeeping is what I'm questioning.
> The impression I've gotten is that a AWS papers over their lack of development gatekeeping with unreasonable oncall loads.
I don't agree with this perception at all. AWS spends a lot of time building frameworks and libraries which attempt to minimize human gatekeeping to only the necessary areas. So think retries, backoff, timeouts, circuit-breakers, etc. Instead of adding gatekeeping at the code level (with things like code style), instead gatekeeping is added at the architecture and operations stages, which is where most of the work of keeping a production level service up, available, and fast. Operational level gatekeeping in the bounds of a given organization is a much more neatly constrained problem than stylistic gatekeeping and this gatekeeping happens on a per-service basis rather than a per-CL/review basis. This keeps overhead a lot lower.
My impression from being outside of Google (and having a Xoogler for a partner) is that Google loves bureaucracy and process, and that engineers that continue to stay at Google are fine with navigating these processes. Its answer to problems is to add more bureaucracy, another item for the checklist that a person/committee/reviewer needs to check for. That reflex is hard to shake, IMO.
I guess I'm not following the distinction. For most people, that's exactly what happens. Either you, or the someone on your direct team, will be the person with readability for your code. The readability process just ensures that you don't end up with silos where style diverges too much, and that the stakeholder who proposes stylistic commentary has a baseline knowlege of linguistic best practices.
With few exceptions, the thing people complain about is attaining readability themselves, not getting changes approved.
> I don't agree with this perception at all. AWS spends a lot of time building frameworks and libraries which attempt to minimize human gatekeeping to only the necessary areas. So think retries, backoff, timeouts, circuit-breakers, etc.
I think we're talking a bit past each other. These things are all table stakes defined by frameworks and updated by centralized processes as best practices change. Human reviewers, for the most part, are only going to say "why are you diverging from the default".
But importantly, that sidesteps my statement entirely, which, rephrasing, is that amazon carries along more technical debt than Google. Whether that's a good long-term business decision remains to be seen, but the impression I get from ex-amazon friends is that maintainability is less prioritized, and maintainability is one of the goals of the readability process. Global consistency helps anyone understand and improve anything else, and to an extend avoids haunted graveyards (though...not entirely).
> instead gatekeeping is added at the architecture and operations stages, which is where most of the work of keeping a production level service up, available, and fast. Operational level gatekeeping in the bounds of a given organization is a much more neatly constrained problem
I'm not sure what you mean, but if anything, this sounds more bureaucratic.
> My impression from being outside of Google (and having a Xoogler for a partner) is that Google loves bureaucracy and process, and that engineers that continue to stay at Google are fine with navigating these processes. Its answer to problems is to add more bureaucracy, another item for the checklist that a person/committee/reviewer needs to check for. That reflex is hard to shake, IMO.
FWIW I don't get this impression. It's honestly difficult for me to tell what you even are referring to. Like, other than readability and promo, there's very few processes or bits of bureaucracy that are consistent across the company, and those two have gotten decidedly more streamlined since I joined the company.
Well okay that's half true, I can think of a bunch of places where bureaucracy has increased, but they're all spurned by legislation.
Separating these two out, even if one person may fulfill two roles, IMO adds to bureaucracy that I don't care for. Ramping up onto a new language sounds like a pretty fraught experience if nobody on your team has readability and the service you're contributing doesn't have a member who has the time to actually work with you, especially if your work is not relevant to theirs. They even touched upon this in the linked video. The whole idea that you can enter a bureaucratic catch-22 at a company seems like an anti-pattern to me, a moment to stop and think; a "modern" engineering organization should try its utmost to accelerate development, with only as many checks as needed for safety and no more.
> With few exceptions, the thing people complain about is attaining readability themselves, not getting changes approved.
That's what I mean by bureaucracy. Even having "another eye" on the code should work. Having a certification process for readability is even more bureaucracy.
> But importantly, that sidesteps my statement entirely, which, rephrasing, is that amazon carries along more technical debt than Google. Whether that's a good long-term business decision remains to be seen, but the impression I get from ex-amazon friends is that maintainability is less prioritized, and maintainability is one of the goals of the readability process. Global consistency helps anyone understand and improve anything else, and to an extend avoids haunted graveyards (though...not entirely).
Ah, that makes a lot more sense. In my career I've never worked at a place that prioritizes code style that much, and I've worked at other high scale places in the past. This seems like a Google specific need. You're right that global consistency helps to save off "there-be-dragons" codebases, but I feel like this is a case of YAGNI. Unless you have engineers constantly cross-cutting across the company, most engineers learn the style guidelines of a team/product/service by ramping up on the team. The optimization for global consistency seems premature to me.
> FWIW I don't get this impression. It's honestly difficult for me to tell what you even are referring to. Like, other than readability and promo, there's very few processes or bits of bureaucracy that are consistent across the company, and those two have gotten decidedly more streamlined since I joined the company.
But promos drive _so much_ of the company culture, at least from what my partner tells me. And the video in the OP certainly makes it sound like Google has lots of bureaucracy, though I'm not sure how many of the requirements outlined in the video were there for hyperbole or not. But especially for services that have low SLOs, it makes no sense to put thought into failover, evacuation, or anything like that.
Keep in mind this video is ten years old. While yes, trying to write something from scratch in a language no one on your team has any experience in can be fraught (although this raises other questions: what exactly are you doing, in every case I've used a new lang, my team has been able to find people to review if needed), even in that case, the organization has help for you (see again my prior comment about "donating" readability, and my understanding is that the readability granting process actively prioritizes people who "need" the readability because they have few potential reviewers).
> That's what I mean by bureaucracy. Even having "another eye" on the code should work. Having a certification process for readability is even more bureaucracy.
I think initially the primary driver of readability is C++ (but Java + Guice also...needs it). Take a look at https://abseil.io/tips/. That's 70 tips, which is around a third of the internal ones. I have C++ readability and have internalized, some of those. The readability granting process ensures your code is reviewed by someone who knows all (or nearly all) of those tips, and will recognize when you're doing wrong things in your code and help explain how or why you can improve. When I personally went through the C++ readability process, I got mentorship on how to fix bugs in my code, some of which a peer on my team would have caught, and some of which they wouldn't have. I can now recognize those bug prone patterns (which are hard to lint for, in my case they mostly had to do with parallelism) and know about tools to debug them (msan and asan, which an experienced C++ user should know, but on my team of people who had basically no C++ experience at Google? Nah).
That follows into other languages. Being granted readability means you have a familiarity with the language that suggests you'll be able to effectively mentor others and not mislead them. You don't have that otherwise.
> Unless you have engineers constantly cross-cutting across the company, most engineers learn the style guidelines of a team/product/service by ramping up on the team.
Google uses a single repo and builds stuff from HEAD. This has up- and downsides, but one of the upsides is that its very easy to investigate unusual behavior in your dependencies. I was doing something similar just this week, and fixed obscure bugs in like 7 other teams tools that were causing downstream problems that those teams weren't aware of. No bureaucracy. I found the bug, fixed them, and sent the change to the owning team. Imagine other situations, where I'd need to maintain a fork, or file a bug to have them fix it, or coordinate an update with them.
You'll pay the cost no matter what, choosing to do so in a way that also provides value and mentorship makes a lot of sense to me.
> But promos drive _so much_ of the company culture, at least from what my partner tells me. And the video in the OP certainly makes it sound like Google has lots of bureaucracy, though I'm not sure how many of the requirements outlined in the video were there for hyperbole or not.
I mean the entire thing is tongue in cheek and again is a decade old. I joined Google ~5 years ago, and even then most of the problems with resource acquisition were solved ("flex") and borgmon and its associated readability were gone, replaced with monarch, which is centralized and used a python-based (though still admittedly arcane) query language that doesn't require readability or managing your own instance. And since then things have gotten even more turnkey for any service that's...reasonably shaped. And, well, the level to which promo drives company culture is consistently overstated (at least in some ways).
Edit: you can look at Google's C-style guide for some examples, https://google.github.io/styleguide/cppguide.html#Structs_vs...
It isn't possible to statically analyze if a class/struct is a POD or if the methods enforce invariants. But it's often very easy to do so with a human eye. And there's value in the distinction!
Similarly, forcing someone to justify using a power-feature (operator overloading, templates, metaclasses, whatever) can only be done by a human. There may be cases where the power feature is warranted and the benefits outweigh the cost, but a linter can't know that. (and ultimately all of this comes back to: things look consistent, and when things are inconsistent, that's a strong signal that something unusual is happening and you should pay close attention)
In any case, readability will comment on stuff that cannot easily be quantified, such as when to use a certain object hierarchy or dependency injection, etc...
The only one that sounds more difficult to codify is telling people of the existence of duplicate functions. But as someone who contributes to the linux kernel, I can tell you right now that the only way that works reliably is to have a very large pool of reviewers. Very experienced engineers frequently miss what people are doing in other parts of the source base, the name might not be what they expect, etc, etc, etc. In the case of linux there are a fair number of duplicates, or similar functions, and people write coccinelle patches to replace them on a fairly regular basis after they have been in the kernel for years.
So, I doubt giving someone a formal gatekeeper flag, really helps vs just having wider change review.
No tool like that tells you if returning a bool instead of an enum is appropriate here, or that a reference vs a pointer makes more sense given the rest of the code.
I'm sure a clever machine learning algorithm could figure that out with a corpus as large as Google's. Maybe. But no tool like that works today.
And not strangely at all, Google does accept "what clang-tidy does" as the canonical way of formatting text. But readability at Google is far more than just formatting.
Readability is frustrating and annoying, but more than just lint.
Naming isn't very language-specific though [*], and poor naming should be flagged in coding interviews. Any Google engineer should generally have good naming, and should be able to critique others' naming in any language.
[*] Granted there are some language-specific naming conventions, like ! and ? in Ruby methods. But I'd say the majority of languages don't have important naming conventions, and those that do often have lints for them.
Exactly! And Readability process is exactly the training Google uses to ensure someone knows how to look for this issue, and that someone actually is looking for this issue in code reviews.
There is plenty about the Readability process I don't like (particularly that one has to get it through a fairly artificial process), but checking for things like the above (which I just use as one easily understood example) isn't something that happens automatically in an organization as large as Google, with as many engineers as Google.
A common story on HackerNews is just how bad the average programmer is, how crummy code they produce. Taking steps to ensure they do it better is a good thing.
We can argue about the process, for sure, and there are many things I would love to change about it. But "yeah, that should be done in code reviews" is exactly the idea here.
I'm not going to turn a candidate away because they use shorthand names when writing code under time pressure on a whiteboard.
I mean, given the halting problem alone, and the absence of general AI, yes, I'd say that...
I'm not sure what to tell you otherwise. I've worked at a place with such strong coding style guidelines, and even though I personally did not agree with every individual point, I accepted the tradeoff because being able to count on consistency in a very large code base was extremely helpful.
Small note: readability isn't a test or quiz you take (asterisk). It's obtained by merging code in the language you want readability for. If you merge code for a language often and the reviewers have very few style-based questions for the code then you will get readability fairly quickly.
> The only one that sounds more difficult to codify is telling people of the existence of duplicate functions. But as someone who contributes to the linux kernel, I can tell you right now that the only way that works reliably is to have a very large pool of reviewers. Very experienced engineers frequently miss what people are doing in other parts of the source base, the name might not be what they expect, etc, etc, etc.
A better example would be knowing when you should use `const std::string&`, `std::string_view` or `char*`. Example: https://abseil.io/tips/1
The best readability advice I have recieved has been:
1. Direct "I was confused by X" or "The recommended way to do A is using B", etc
2. Reasoned: "std::string_view is more efficient and clearer in intention than char ptr, it also improves type safety as it is read only and clear about ownership"
3. Linked to source material where examples are given totw or other examples in the code.
Very true. Readability can't help with that, nor is it designed to. It's mostly there to help novices and new hires. Experienced engineers already have readability themselves so they don't need this extra review.
Someone with readability in a language, who keeps up with the style recommendations, will generally produce code that is easier to read by other engineers.
Let's make it specific. Read https://google.github.io/styleguide/cppguide.html for readability for a language, namely C++. All the things that can be automated, automatic tools have been written for. But, for example, you can't automate "Prefer to use a struct instead of a pair or a tuple whenever the elements can have meaningful names." Because what does it mean for a name to be meaningful?
Are you optimizing for someone who already knows all the project lingo, or someone who doesn't know any of it?
Are your engineers native English speakers?
There are a whole bunch of things which make the perfect variable name frequently less than perfect, and putting project insiders in charge likely yields the opposite result.
Take: https://elixir.bootlin.com/linux/latest/source/mm/khugepaged...
If you don't know what a vma, pte, pfn, compound_page, young pte, huge page, lru, etc your going to be unable to even begin to understand what that code is doing, despite those all being pretty reasonable variable names and actually fairly industry standard concepts. It gets worse as you move to more esoteric topics. Expanding pte to PageTableEntry might help some subset of users, but at the expense of those that work on the code daily. So who do you optimize for? Is it readable if the only people that can read it already know what it does?
Also probably the privacy review could be a bigger bottleneck these days ;-)