Real world examples of these design docs can be found at https://github.com/golang/proposal
Real world examples of these design docs can be found at https://github.com/golang/proposal
Google's design document culture is bad. Google has succeeded in spite of it.
In my experience, having worked at many large tech companies, design documents obfuscate, not enlighten. They become increasingly out-of-date as the code evolves, creating anti-documentation that makes it take longer to understand code. Yes, yes, people should update design documents as the code evolves. Everyone knows that in practice, nobody updates old design documents.
Design documents make it too easy for other developers to shoot down ideas. Sometimes the worth of code isn't apparent until it's made. It's far too easy for someone to comment "this will never work" on a proposed change. It's much harder for someone to deny benchmarks attached to a proposed change. It's too easy for reviewers to knock out functionality.
Design documents turn every feature into a half-assed, lowest-common-denominator risk-minimized barely-adequate shell of itself.
The real reason everyone at Google writes design documents is that promotion committees demand documents as "evidence of complexity". No design document, no impact. No impact, no promotion.
Code is just code. Bad changes can be backed out. It's much better to move fast and iterate quickly than to create an illusion of care and add friction to every aspect of the development process. Up-front design of software just does not work. If it did, waterfall project planning would be successful.
These questions you highlight --- Why are you making this change? What impact will it have? --- can be asked during review of actual code. There is no need to build a speedbump, not if you trust your people.
Developers should be able to choose to create design documents and solicit feedback. Most changes don't need this process. You should trust developers to know what changes require a more extensive discussion and which ones don't.
A culture that requires permissions and signoffs before work can begin is a culture that leaves products stagnant for years.
>Design documents turn every feature into a half-assed, lowest-common-denominator risk-minimized shell of itself.
Sounds like a sentence written by someone who is an engineer and not a support staff or a user, i.e the people who have to deal with the fallout of every feature change and every engineering decision. I could just as easily substitute "feature driven design" into most of your points above.
Requiring someone to put thought into the design of a product and to subject that thought to rigorous scrutiny is, on the whole, a good thing. One that leads to products end users want to use.
Do support staff and users sign off on design documents? "No" is the universal answer. Are you claiming that engineers aren't reasonable human beings who can take support staff and user concerns into account on their own? What makes you think the people reviewing design documents can do that?
Is it that you just trust a subset of engineers to understand the big picture? Are most engineers just drones? You shouldn't hire people who don't give enough of a shit to take the big picture into account.
I am an engineer. I am also a user. I do support my software. I've been programming for twenty years. I find that "rigorous scrutiny" hurts more often than it helps. Maybe this rigid scrutiny was more appropriate in a world with release cycles measured in years, but we don't live in that world anymore. When you ship every week, you can easily undo mistakes, and you're better off erring toward iteration.
At a certain point you can't. I was recently asked to implement a feature for a usecase for another engineer. It required a design doc and review by a few representatives from related teams. He and I wrote the doc, and the initial review was that any solution that would fix his usecase would break the testing infrastructure for practically every other developer in the building. He could write a few fewer loc when testing, at the cost of tests failing due to unrelated changes.
Neither he nor I had the awareness of the impact, and realistically couldn't have, it was neither of our responsibilities. And because I spent an hour filling out a template document, I didn't need to waste my time doing that.
Why wouldn't continuous integration have caught that bug? If a diff breaks the build, it shouldn't land. If a diff causes tests to fail, it shouldn't land. If a diff lands anyway and causes problems, any developer negatively affected should be able to insta-revert the diff.
How does a design document help?
So again: how does a design document requirement help?
At least where I work, no. Your team, your manager, and anyone who you think will be affected, and depending on the scale of the change, you inform everyone that uses your tool or works on your product so that anyone can provide comments and feedback.
> If you do suspect ...
And that's exactly it. If I suspect something then of course I can communicate it directly or redesign, but we're trying to assess impact on areas I might not have no awareness of.
So let's put it in another way:
Design docs is one of many formal methods of communication in any organisation. It preserves context, change history, and a good part of risk management. It protects you and potentially saves downstream rework.
Feedback from a team that worked with another tool we used in the testing process informed me of the potential for the problems. To be clear, this would have eventually happened anyway, its unlikely that a PR breaking the CI system in this way would have gotten approved, but there would have been much wasted effort on my part in the interim.
Or so they thought: http://pythonsweetness.tumblr.com/post/64740079543/how-to-lo...
Even if you care, you can't know the entire story. I work for a small company, tiny compared to Google, but often we run into someone proposing a change that backtracks on a strategy decided 2 years ago. Luckily it was encoded in a design document or else, how would anyone new find out about it?
New people want to know the big picture, but the best way to do that is to write about your decisions as you are making them, and give them some sort of searchable record.
Personally, I love Github because while you don't have formal design docs---you do have a history of the entire argument of why a feature should be implemented, followed up with a history of all the breaking changes caused by it, and the final reversion. It's great to plunge into a codebase, and see how did we get there, before deciding to propose something.
Have you ever been on a team, and proposed something only to hear "yes we tried that, it failed for X,Y,Z reasons... but no we never noted down starting this 6 month initiative down anywhere except my head".
Why should anyone be bound today by decisions made two years ago? Maybe that strategy no longer makes sense.
I've seen it go both ways. On one hand, sometimes a new developer in an old codebase does something "against the grain" of the system due to unfamiliarity or JavaScript-induced brain damage. On the other hand, sometimes circumstances change and even good old designs become obsolete.
A culture of good taste and transmission of institutional knowledge helps preserve the good aspects of design. I don't think design documents help: they're just bytes on a disk. Unless you have people to enforce them, these documents won't do a thing. If you do have people who know what the system is supposed to look like and who shepherd changes to work with the original design, you don't need the design document.
Then that'd be a great time to discuss that the strategy should be changed, and move forwards as a team---not a single developer deciding to go against the grain with side effects that may be unknown.
Here's an example:
At my company it's our strategy that all file operations happen atomically and asynchronously. Even if your function was the one to create a temporary file, and you are absolutely sure it's happening locally, deleting it must be an asyc task handled by another worker.
Why? Because historically, small tasks start off as local only operations, but get upgraded to handle remote instances. Remote operations can fail fairly often, and we don't want the entire task chain to die because you couldn't delete a temp file on another server.
Now .... to any single developer writing a small script, this feels inane because forcing it to a asyc op, using our message bus tool chain, etc. will slow down the entire function by an order of magnitude. But it saves on 10x more integration work 6 months from now that they don't anticipate.
Engineering is a process to solve problems at its core. Fundamentally, you cannot solve problems without know on what the the questions are.
Wow. I feel terribly offended as an engineer.
Could we not confuse "engineering" with "pushing random changes at any times that may not even pass the tests by any bro-ninja that just felt like it".
The poster has clearly never written or maintained any critical software.
Move fast and break something is the domain of hacking.
Whatever software is made, the [lack of] efforts put in design and reviews will have high impact on quality [or lack thereof].
Believe me, I'm the furthest thing from junior you'll ever see. I don't particularly care what you call me, but I'm produced tons of value.
There's a certain type of mid-career programmer who's obsessed with "best practices" and thinks that anyone who doesn't stuff a program full of design patterns is being incompetent and irresponsible. It's a kind of "sanctimony porn". The attitude is that "if programming is hard for me, it'd better be hard for you too". It's this kind of programmer that shames other programmers for having opinions that result in the creation of simpler code.
I hope you grow out of this phase.
Its not a permission to do work, its permission to merge the results of the work to master. Those are different.
>Google's design document culture is bad.
Then why do we have JEPs, PEPs, and rfcs in perhaps every other major project out there?
Also, to be clear, Linux is relatively small compared to Google or Microsoft, or indeed many corporations codebases. Someone could conceivably read the entire Linux kernel codebase. That's not true for BigCorp.
Compare that to google, where[1] there are almost as many unique source files as the kernel has lines (though to be fair, many are autogenerated).
[1]: http://cacm.acm.org/magazines/2016/7/204032-why-google-store...
Of course a full design doc is not needed for every change. Many changes are small and straightforward. Only big things, where the team needs to discuss and understand options. Or bigger things, where directors etc need to sign off (not really design docs anymore, but same basic thing).
The fact is that, on average, design doc+code takes less time than code without design (that is, without communication). Again, only for certain kinds of changes. Things go faster because problems are found, approaches are adjusted, or unmotivated features are axed.
I have no problem with individual developers choosing to circulate documents in order to solicit feedback. My objection is to rigid processes that force engineers to write documents.
Mandatory design documents for "communication" invariably morph into checklists of required signoffs from people who have little incentive to say "yes". In this way, a culture of design documents breeds a culture of extreme risk avoidance.
It's a tragedy, really. Through numerous small steps, each apparently reasonable, a nimble organization becomes an ossified nightmare in which it takes six months to add a checkbox.
> The fact is that, on average, design doc+code takes less time than code without design
Prove it. Provide evidence. In my experience, your claim is not the case for most changes in most projects. For the changes where design documents facilitate development, my experience is that developers will choose to circulate documents even when not required to do so.
> Unmotivated features are axed.
One developer's "unmotivated feature" is another developer's essential use case.
Another trigger could be a reviewer commenting, "Hey, this diff stack is getting pretty tangled. Can you write something that describes how it all fits together?"
You should trust your developers' judgement.
I don't think you have the causation quite right --- it's not necessity that forces the adoption of process exactly. Process is what you get by default if you don't consciously counteract natural human tendencies in management. A lot of large companies stop consciously protecting their culture, so they get the default big company culture instead. The default big company culture is ever-increasing process.
Everything I'm familiar with has approval and review on everything in government, bureaucracy is synonymous with government to many people for exactly this reason.
Well let's put it this way, let's say that I'm a googler, I work on the android team and I want to rewrite the launcher from the ground up because I have some whiz-bang idea. This falls under my general purview of stuff that my team works on, so I grab my coworker, and for a month we hack away at it and have a good MVP. It has great improvements over the existing homescreen.
We've been working and reviewing each other's code, and so after a month we submit the big change that touches bunches of files out for review by all of the various owners of various affected components. One of them shoots me an email a few minutes later and says that this work is all wasted because another team has been secretly working on an updated launcher for the release of Google's new flagship phone, the Pixel, it already has some of our features, but also has many others, and has been in development for 6 months already.
So now my coworker and I have wasted 2 man-months of employee time on something. That's tens of thousands of dollars that goes poof when someone closes the pull request. And those 10s of thousands of dollars could have been saved with a 2 page document and a 1 hour meeting.
The goal of these things is to not waste developer time, because you or I can't be aware of everything going on in a company, so writing a design doc allows other people to see what you're doing and
1. provide insight and feedback from their experience with similar problems
2. provide prior art from within the company
3. remind you of things you might have forgotten about
4. give you insight into how these changes will affect others
5. most importantly, give you information about other activities the company is doing in this direction that are related to your proposal, so that if you can avoid destructive interference, and potentially have constructive interference.
Otherwise you end up with repeated work, wasted effort, and fragmentation.
I'm not suggesting that communication is bad. I'm objecting strenuously (perhaps stridently) to the design document as a step in a rigid process and a requirement imposed from above.
Personally, I'd rather see two teams work toward our improved hypothetical launcher than to see zero teams do that work because process imposed too high an "activation energy" on the launcher experiment.
Also, why are both teams working in isolation for a month? If both teams check in code, the checkins themselves can provide an indication that other people are working in the same area. Incidentally, it's this effect that makes me strongly dislike feature branches. It's better to develop unstable code behind a feature flag, where the code is visible even if not active, than in a feature branch, where nobody can see it.
(It's true that not all changes can be gated behind a flag, but most can be.)
I chose my example carefully, the 'new' launcher being developed officially was part of a new, unannounced product, and therefore likely secret prior to release. On the other hand, the new version developed by the two people was a top to bottom revamp, which might mean an entirely new application, or something that can't easily be done behind a feature branch.
>Personally, I'd rather see two teams work toward our improved hypothetical launcher than to see zero teams do that work because process imposed too high an "activation energy" on the launcher experiment.
And again, I chose my example carefully: one team would always have created this new launcher, it was creating something as part of a much larger initiative.
>I'm not suggesting that communication is bad. I'm objecting strenuously (perhaps stridently) to the design document as a step in a rigid process and a requirement imposed from above.
Indeed, but if nothing else, a design document is a formalized process for this type of communication. To be clear, my understanding is that Google also has best practices in place for objections and blockages to things proposed through design documents.
The result is that when you are planning to create a feature that is large enough that it is, shall we say, statistically likely to break things or cause someone else serious aggravation, there is, if you are following policy, a formalized process for informing the people who your change might impact, a way for them to provide feedback, and a formalized way to resolve conflicting opinions when each engineer thinks that their workflow or feature is the most important.
This goes back to the other example I gave, which was a real one, where my design documented project could have, if implemented, created a situation where, until rectified, breaking changes would have gone undected by automated tests, and instead shown themselves in the tests for unrelated changes, which would have been an annoying bug to track down, and would have caused some other team a lot of stress and time. They would never have been notified until final review, and conceivably could have been months of wasted effort on my part.
If your culture is one of secrecy (Google's generally is not), then you need process to effect coordination, since nothing else can. Fortunately, good (IMHO) software companies are internally open, so there's no secrecy forcing them into heavyweight process. Process is one of the many costs of secrecy.
> formalized process
The problem with formalized anything is that it involves writing down rules. When you write down rules, you have to distill complex and subtle human interactions into essentially an algorithm for people to follow. The loss of nuance, while creating clarity, introduces inefficiency, since it forces everyone to follow the same steps even when these steps are inappropriate.
You can't write down exceptions for all the inappropriate cases (but you can for some of them). If you tried to allow for a large number of exceptions, the resulting algorithm would be too complex to follow or it would be vague enough to allow anyone to skirt the rules.
I prefer to avoid formal rulesets for human processes. In my experience, the efficiency loss arising from formal rulemaking has outweighed the gain in clarity. Maybe others have different experiences, but in mine, the higher the quality of developer you have in an organization, the fewer formal rules you need.
Applying lots of formal rules to good developers just makes them unhappy.
Internally open does not mean that everyone can always see everything though.
>Applying lots of formal rules to good developers just makes them unhappy.
I'm not sure I agree with this. I'm a developer, and I'm happier when there are formal rules about human code review, code style, automated testing, and minimal test coverage.
I'm happy about those things because they are a small inconvenience (if one at all) that save me time and effort down the line, and insulate me from lazyness, and mistakes on the part of developers (myself included!).
I see design documents as an extension of this. My pay in is that I occasionally need to write a document outlining my thought process w.r.t. a new (large-ish) feature I'm implementing. In exchange for this, I know that I will not waste my time reinventing things and, more importantly, when other people intend to make changes that will potentially negatively impact me, I will be able to provide feedback before those changes go live.
If secrecy is rare in an organization, secrecy-induced problems in that organization are also rare.
> I'm not sure I agree with this. I'm a developer, and I'm happier when there are formal rules about human code review, code style, automated testing, and minimal test coverage.
It takes all kinds, I guess. Personally, nothing galls me more than having to follow a rigid set of rules when I know the original motivation for these rules doesn't apply to my situation. I'd make a terrible soldier.
I understand your social contract argument, but IME, the benefit to me of other people writing design documents isn't big enough to justify my having to do it. I find that subscribing to actual code reviews targeting directories that interest me gives me enough ability to see and affect (and unfortunately, sometimes block or delay) changes before they go live.
In the case that you're asked to implement a project from 'above', they're still useful, for practically the same set of reasons (maybe minus #5, but plus 'other people who you impact can provide feedback to reduce future friction')
That's amusing, that you want solid evidence, yet you're willing to use your own anecdotes.
> Mandatory design documents for "communication" invariably morph into checklists of required signoffs from people who have little incentive to say "yes".
Or they make you think about things that are not obvious on first glance, especially at Google scale. For any customer facing feature, you have to make sure that PII is taken care of, that security is implemented properly (SQL injection and XSS vulnerabilities for example), that internationalization is taken care of (especially right to left languages), and that UI fit and finish plays well with design guidelines, in both web and mobile, and that cross browser compatibility is at least thought about, as well as other issues.
> One developer's "unmotivated feature" is another developer's essential use case.
One developer's essential use case is another three dozen developers' backwards compatibility breaking change.
> Through numerous small steps, each apparently reasonable, a nimble organization becomes an ossified nightmare in which it takes six months to add a checkbox.
When you're serving up traffic at those volumes, with datacenters all over the world, accumulating revenue that quickly, yes, it's worth taking six months adding a checkbox to take every possible step possible to ensure that doesn't leak a security vulnerability somewhere.
Just because your individual progress is slow, doesn't mean that the progress of the team is slow. One breaking change in say Google adwords can undo literally man years of work.
I'm not the one presenting my anecdotes as fact: "The fact is that, on average, design doc+code takes less time than code without design".
Anyway, you've very clearly articulated the conventional wisdom of big companies originating in a certain era of computing. Conventional wisdom isn't necessarily wrong, but it's not necessarily right for all time either. There are companies with data needs, user counts, and codebase sizes on par with Google that don't practice Google-style process, yet succeed anyway. That these companies have succeeded without Google's process is evidence that Google's process is unnecessary, at least in today's environment.
> SQL injection and XSS vulnerabilities for example
Code-level concerns. You're not going to stop SQL injection by looking at some high-level design document. The same goes for r2l text layout bugs.
> it's worth taking six months adding a checkbox to take every possible step possible to ensure that doesn't leak a security vulnerability somewhere.
Keep that in mind when smaller competitors surpass you. It's easy to say that Google's codebase represents 18 years of work. I strongly suspect that it wouldn't take so long to do starting today.
Look at self-driving cars: how long has Google been working on them? How long has Uber? Whose cars are serving real-world passengers today?
These days, we have 1) very good continuous integration systems, 2) good code review tools, 3) fast shipping vehicles, and 4) continually improving static analysis. These things were unavailable (at least at adequate quality levels) when Google started its design culture.
Maybe the conventional wisdom you articulate might have been an optimum some time ago. These days, I think it's far too process-heavy and that Google and similar companies haven't kept up with the times.
> Just because your individual progress is slow, doesn't mean that the progress of the team is slow
It means that the team is inefficient. Communication overhead goes as N^2, after all. Google's teams are notoriously huge. When you see that a startup (or even another > $1 billion company) can do the same damn thing Google does and put a quarter of the people on the task, maybe it's time to wonder whether Google is doing something wrong.
If developers feel like process is slowing them down, maybe you should listen to them.
> One breaking change in say Google adwords can undo literally man years of work.
It's easy to look at a few failures and conclude that you need to add process to fix whatever went wrong. It takes much more foresight and wisdom to see that this process probably costs more man years in overhead and inflexibility than you spend fixing the occasional mistake.
Which? The ones I can think of are Apple and Microsoft, and I'm pretty sure they practice Google-style process. Amazon has its own flavour of process which is heavyweight in its own way. What are you thinking of?
> Code-level concerns. You're not going to stop SQL injection by looking at some high-level design document. The same goes for r2l text layout bugs.
I notice you left out PII and other security implication, as well as design fit and finish. Those can easily be caught at design time, especially I18n bugs. For instance, the average east asian phrase is shorter (graphically) than the same phrase in a western language, and that all has to be translated and dealt with.
> It's easy to say that Google's codebase represents 18 years of work. I strongly suspect that it wouldn't take so long to do starting today
That's a strawman argument. Google has written reams of code for distributed computing (Borg), continuous integration (Tap + Blaze + Forge), code review tools (Mondrian, Critique).
That's like saying that although it took a decade to design the Boeing 737 (just a pretend example), it would take less time now. That is correct, but based on advancements on technology and materials, what's your point?
> Look at self-driving cars: how long has Google been working on them? How long has Uber? Whose cars are serving real-world passengers today?
That's a false comparison. Right now, Uber still has to have drivers behind the wheel, whereas Google self-driving cars strive for a higher level of autonomy. Also, Google has not wanted to get into a directly customer facing role, instead looking for partners to manufacture the cars.
> These things were unavailable (at least at adequate quality levels) when Google started its design culture.
Design culture evolves. Google wrote all of its own integration systems, code review tools, and many static analysis tools. Even though they have top class systems, they still stick to the same way of doing things. That's evidence that it works, and the process is roughly where it needs to be.
> When you see that a startup (or even another > $1 billion company) can do the same damn thing Google does
What's an example? Most startups/competitors to Google seems to do about 90% of the things that Google does for one business division, leaving aside the last 10%, which is naturally the hardest 10% to do.
> If developers feel like process is slowing them down, maybe you should listen to them.
Which developers? People looking in from the outside or actual Google engineers?
> It's easy to look at a few failures and conclude that you need to add process to fix whatever went wrong.
Interesting study on checklists. http://www.nature.com/news/hospital-checklists-are-meant-to-...
If you have institutional resistance towards checklists in hospitals (or process), introducing them doesn't help. But if you actually implement them correctly, they do eliminate many common mistakes.
A lot of companies use cargo-cult like process, thinking if they follow a magical recipe, they automatically get good results. I doubt Google is one of them
Microsoft has not collapsed. In fact, it's doing better than ever. If Microsoft of all companies can reform itself, so can Google.
> PII and other security implication
When your developers are both smart and invested in the product's success, they learn about these things on their own. Sure, they can make mistakes, but so can some damned review committee.
It's interesting to see how people rise to challenges. If there's a security committee tasked with reviewing the security implications of various changes, developers won't take security as seriously. "That's the security committee's job", they might think. But if you entrust developers with their own security, and they're high quality developers, they'll take the responsibility seriously and do a better job.
(I know I'm making a "no true Scotsman" argument, but I think there's a real qualitative difference between developers you can trust with this sort of responsibility and developers who aren't as invested.)
I don't think you can look at security/PII/whatever problems that a committee catches and conclude that those problems would have made it to production absent the committee.
> That is correct, but based on advancements on technology and materials, what's your point?
The Brooklyn Bridge was designed to be six times stronger than it needed to be for its design load. Modern bridges are only about two times stronger than they need to be. The Brooklyn Bridge needed its large safety factor because suspension bridges were not well understood at the time. With modern technology and design tools, we don't need to pay for a safety factor of six.
Imposing Google-style process in 2016 is like building every modern suspension bridges like the Brooklyn Bridge because the Brooklyn Bridge is still standing. "That's evidence that it works, and the process is roughly where it needs to be."
> For instance, the average east asian phrase is shorter (graphically) than the same phrase in a western language, and that all has to be translated and dealt with.
Pseudolocalization and dogfooding help. I'd argue that rapid iteration helps most on UIs. A/B testing and metrics beat heavyweight up-front design any day of the week.
> Design culture evolves. Google wrote all of its own integration systems, code review tools, and many static analysis tools. Even though they have top class systems, they still stick to the same way of doing things. That's evidence that it works, and the process is roughly where it needs to be.
There's an ever-increasing morale cost. How do you expect developers who have experience in process-light environments to come to Google and be happy? "Yes", nobody says, "I want to go from experimenting rapidly on my ideas to writing internal documents to convince people to maybe let me try something."
> If you actually implement [checklists] correctly, they do eliminate many common mistakes.
It's a lot easier to back out a problem diff than to remove the staph you accidentally introduced into a patient's bloodstream.
Process should be proportional to the difficulty of undoing a mistake. If a mistake is easy to undo, it should be easy to do. If a mistake is very costly to undo, it's worth investing in not making the mistake in the first place.
The vast majority of programming errors are of the "easy to undo" variety.
And by the way:
> [Uber's self-driving cars is] a false comparison. Right now, Uber still has to have drivers behind the wheel, whereas Google self-driving cars strive for a higher level of autonomy. Also, Google has not wanted to get into a directly customer facing role, instead looking for partners to manufacture the cars.
Uber recently delivered beer fully autonomously. In a truck. They're definitely planning for L5 autonomy.
It's definitely possible for a team to come along and catch up/overtake the Google (now Waymo) project, but I agree with jimmywanger it's not a valid comparison for the sake of this discussion. Uber is following a different path than Google focused on, and is hugely benefitting (as is the whole industry) from the work done at Google.
At Google's scale even small issues will have widespread impact on real people. Let's say you break the ability to reply to email in GMail for ten minutes - cumulatively that could result in hundreds of hours of lost work across all their users.
Some people like to cower behind "Google scale" as a reason never to change anything. Not me.
I don't know what you think of Google's process. A design doc is needed for any large user facing change or large infrastructure change, and it goes through sections and you skip the ones that are not applicable.
For instance, if you're not storing user information, you skip the PII section and so forth. If you're just making a change to adwords billing, you skip the entire I18N section.
Also, the areas in which Microsoft is revitalizing itself are green field projects like the cloud and some other interesting hardware/software integrations. You can play fast and loose with those, as opposed to Google, which doesn't really have any legacy code and has to support all existing users.
> When your developers are both smart and invested in the product's success, they learn about these things on their own. Sure, they can make mistakes, but so can some damned review committee.
The review committee does this for hours a day, and they see far more cases. That's like saying that it's better for you to assess the condition of the transmission of your car, because you care more and are more invested. I'd rather have the guy who rebuilds transmissions for a living, who has seen dozens of transmissions, and is familiar with common failure modes and pitfalls.
Specialized labor does help.
> The Brooklyn Bridge was designed to be six times stronger than it needed to be for its design load.
Well, heavier than air flight was impossible before the 1890's without investment in materials, engines, and construction techniques. Is it easier to build an airplane now? I still fail to see your point. We're talking about something where you have to invent the tools to make the tools to make what you want to make, vs. already having the tools available.
> Pseudolocalization and dogfooding help. I'd argue that rapid iteration helps most on UIs.
AB testing on wireframes helps and gets most of the edge cases. After you roll out to production things get hairy.
> It's a lot easier to back out a problem diff than to remove the staph you accidentally introduced into a patient's bloodstream.... The vast majority of programming errors are of the "easy to undo" variety.
Not really, when you're working on fundamental libraries that many products depend on. That can cause issues all up and down the product stack, and you're going to cause issues for developers who rely on your code who now have to throw away months of work.
> Uber recently delivered beer fully autonomously. In a truck. They're definitely planning for L5 autonomy.
That's a publicity stunt, and it's unknown whether or not they got paid or not. Also, Otto was based from Google expats, one from Google maps and one from the Google self driving car company. Your point is?
The example I have in mind is in a big legacy product. I can't get more specific without outing myself, but it's very far from greenfield.
> Specialized labor does help.
Specialization of labor can also hurt. I've found myself frustrated with security people in the past because they spend so much time thinking about security threats that they start to veto massively useful functionality on very flimsy security grounds. Broad exposure helps too.
> AB testing on wireframes helps and gets most of the edge cases. After you roll out to production things get hairy.
Why? There's no rule that says that everyone needs to see the same UI in production.
> Otto was based from Google expats, one from Google maps and one from the Google self driving car company. Your point is?
It's telling that Google autonomous drivers experts had to leave the company in order to get their work into a real live product.
It seems as though you haven't really encountered Google process in person, you've just heard stories. Security people do have a day job, they just do security on the side because they've expressed interest/aptitude.
My point remains. I'd rather have a plumber fix my plumbing or check over plumbing designs rather than an enthusiastic amateur.
> Why? There's no rule that says that everyone needs to see the same UI in production.
They don't. Google constantly runs A/B testing. Once you get it to production you've already invested the time in productionizing it.
> It's telling that Google autonomous drivers experts had to leave the company in order to get their work into a real live product.
Or that it's be far more lucrative to be acquired than continue working on Google X. You can't ascribe motives to their actions.
That's a really bad analogy, borderline dishonest.
In this case it's a mechanic taking his vehicle to another mechanic for service, you better believe that first mechanic is both more invested and more familiar with the transmission.
Let's start with "how big do you think Google's codebase is"?
Because the last time we went looking, the number of companies even close was <10, and they all pretty much have the same as Google's process.
If you really have examples, i know the engineering productivity guys would love to hear about them and talk to these folks.
I think I read somewhere that, on an online shop, each step you add to the checkout process halves the number of people who complete it. If true, that is a powerful argument for what you're saying.
Also reminds me of Bastiat's "What is Not Seen". We don't see the developers who chose not to go through the whole process because it was too tedious.
I wonder how much of that is just "more steps bad", and how much is that any step you add to a checkout process beyond what is necessary for the actual sale is (a) forcing you to do something you don't care about, and (b) going to make it more intrusive? Eg. forcing you to create an account, forcing you to validate an email address, forcing you to fill out a "quick survey", forcing you to unsubscribe from their marketing spam.
There have been plenty of times that I've wanted to buy an item from an online store but have decided against it because the store wants to "create a relationship" with me instead of just selling me the damn item.
We just use Google Docs, and it's easy to share the doc and have people add comments or propose changes. There's a template somewhere, but I don't like it much.
But I think you're missing some very practical things here:
> Why are you making this change? What impact will it have? --- can be asked during review of actual code.
Yes, and they're going to be asked every single time. And if the reviewers disagree with the impact, you'll have to either drop the change or rewrite it. So why not ask first?
> There is no need to build a speedbump, not if you trust your people.
"your people" works on a small scale with people you work with continuously. It doesn't work on the internet when someone called "fdsfsaa" submits the code to your project and you hear of them for the first time.
> Most changes don't need this process.
Depends on the stage of development goals, etc. Some projects will get bugfixes and maintenance. Some will be constantly changing. You can't generalise.
While writing a CEP, my gut reaction was like yours, "this is a waste of time, code is art, and I am programmer Picasso," etc. But since my manager and my technical leads at the time were very good, I bit my tongue and did the work of writing up a CEP anyway. In that environment, it didn't take that long to write a CEP. It was about 2 pages. And since I was proposing to change how a critical piece of infrastructure worked, it was really important for the oncall people to be on board, and it had to make sense, and it was important to identify all the failure points in advance, etc. From the egocentric perspective, I'm certain that it was much less work for me to write those 2 pages than it would have been to explain to 20 different people what I'm doing and why, either at the water cooler or during code review.
Trust is often earned and not given. Your coworkers may not know you or the quality of your work. Under the right conditions, I don't think being asked to write a CEP is being asked to dilute your vision, it's merely being asked to define it, and describe how it fits in with how things already work. If that antithetical to you, then you need to work at or found a small company, where everyone is most concerned (hopefully) with making something work, rather than trying to make something that is already working and already making money better.
I have some points of agreement with you, although I wholeheartedly disagree with the conclusions you make.
"Everyone knows that in practice, nobody updates design documents after the fact." Even if this is completely true (and it isn't -- I have written and read many documents that closely follow current practice) does that make it better to not try?
"Sometimes the worth of code isn't apparent until it's made." Does this mean that it is too hard to explain why it's worth doing?
"These questions you highlight --- Why are you making this change? What impact will it have? --- can be asked during review of actual code." You're right, these questions will certainly be asked -- in which case, you can link them to the CEP which will probably answer their questions plus other ones that they didn't even think of to ask.
I can't disagree more that code review can replace a CEP.
I don't know if design documents need to be a huge thing with stakeholder sign off and approval processes necessarily - maybe when you're google size. At my (small) company, we use something similar (we simply call them spec's) and it's more of a way to express your concept than it is to get sign off. Sometimes just spelling out an idea helps you validate it.
Really, you should be able to defend a certain amount of criticism if you're writing code that someone (and probably not you, eventually) will have to maintain. Bad changes can be backed out if caught early, but if left they find ways to creep into other areas and cement themselves when they are built upon.
I don't disagree with your main point though, but I do believe there is a balance.
I'm not familiar with Google system, but in other big corporations the problem is that you end up with system which does not allow exceptions - so engineers end up requiring writing 10 different documents to make a small change which can be explain on one page and everybody will get it: support, testing, doc writers, etc.
And these documents ended up begin written so that management is happy and not actual intended audience: tester, support, maintenance, documentation writers, etc. So these documents ended up being useless so ....
The worst result is that engineers which do not write good code but write good essays end up with promotion and eventually destroying the entire product (not their fault - they just not talented programers).
In that kind of situation, you betcha we need to have design docs! And in terms of making it easy for other developers to shoot down ideas, very often they may know about some dependency or key assumption in some other piece of code that you didn't know about it. And it's better to find out about it during the design phase, than to have to rework 50% of your work when you find out about it at code review time, or worse, if it gets deployed and you get angry notes from SRE's that were woken up at 3am and you need to send them a bottle of whiskey to apologize for your f*ck up....
I don't see such a culture. I often create one or more prototypes as part of my design. One of those might become the final result, or I might throw away everything. As long as that's understood, all is well.
The way I think about it is this: don't do work you aren't willing to throw away until earlier steps have been reviewed. How much work you're willing to throw away is a personal preference. Your design reviewer(s) might say "did you consider this other way that has these advantages?" You shouldn't reply with "No, and I've invested too much time to consider other approaches now. Stop holding me up and lgtm already." No one wants to work with someone like that. Your argument should be based on what's best. What's already done should only be considered if neither approach is (believed to be) significantly better.
My experience has been the reverse, precisely because as you said.... no one updates old design documents.
A piece of code with a design document at least has a historical record of what the original aims of the project were, and a written rationale for why they took certain approaches.
Often you'll find a piece of code with a seemingly inane architecture and wonder "why is this so inane? were the the developers on drugs?". By reading the design document, you find out that sadly no the water fountains were not spiked with LSD in the 70s, but rather they were working around the performance characteristics of hardware that no longer exists and thus these baked in assumptions had reason and merit.
Understanding the original why often illuminates the entire architecture, even if undergone a lot of changes because the original skeleton still remains.
I sort of look at it as Code Archeology, or maybe literary deconstruction as applies to code.
That information is very useful, particularly as a comment or a commit message. Storing this information in a separate unversioned document off on some enterprise management system makes it harder to find. If you put the rationale for a change in the commit message, the rationale is right there when you run blame!
Commit messages are really good for e.g. "patch that due to this bug" or "refactor this so it's decoupled from that feature", but not quite so good for describing arches in development. On the other hand, you can reference an issue in every commit message, and that issue can lead to the design document.
Rather than regurgitate what most people here already said, let me list a few programming projects, domains, and tasks where at least thinking about design if not writing design documents or spending days, weeks, or months figuring it all out is worthwhile.
* Programming Languages
* Databases
* Operating Systems
* Medical Devices
* Safety Equipment
* Streaming containers/formats
* Encryption
* Security
* Manufacturing/Robotics
* Aerospace / Space
* App Dev Frameworks
* Game Engines
I could go on.
The point here is that there are plenty of things where thinking about it up front is beneficial, if not required, especially if some combination (but not limited to) the following are true:
* Lives are at stake
* Changing it later would be hard (programming languages are an egregious offender, I won't name names)
* Customer adoption will completely derail or forbid architectural changes
* Fixing it will require essentially doing it again from scratch
* Changes will force the creation of patches that will incrementally kill the project or slow future development
Frankly, I think we have too many things that are poorly designed. Most projects I see in nearly any domain are mostly set in stone once time and money is added to the mix. Everyone talks about redoing or fixing things, but it rarely happens except for minor changes. As projects scale up, few people can afford to constantly back out lots of changes and rearchitect everything. Those that do usually fail or don't get a good ROI, and those that don't change fail anyway.
I've worked with all kinds of people and though there are people I have great admiration for, I can safely say that 99% of them are idiots and have no business being programmers. I know it sounds harsh, but I've been doing this a long time and have worked with all kinds of people. Too often I see the programmer's equivalent of an illiterate child that gets pushed through high school. So no, I don't trust people to do the right thing, I merely trust most people I work with to not act maliciously. Most of all, I don't trust myself. As the progression goes as a programmer - your code sucks -> my code sucks -> all code sucks -> my code sucks but I'll live with it, hope it is better than most, and ask people smarter than me for help.
Most better developers I know do in fact right some form of design documents, even if it's just notes and justification why X or Y won't work, but Z "might" work. Many also take a lot of time to think about something before writing any code, but once they do, they actually finish much quicker with less bugs than the young programmers who want to "move fast." Of course none of this is universal, and as I said, it all just "depends." What do I know?
I feel the same way. If you have a group of idiots, you need process as a harm reduction measure. A very high contributor bar is a prerequisite for a process-light environment, because if you can't trust people to do the right thing, you need a system to force them into a conservative approximation of the right thing, and this system is called process.
That's what they all think they're doing. Math doesn't work that way.
In other words, you take measures to catch errors, mitigate failures, and protect yourself rather than say "go code and push to production everyone, we can just roll back!" That might work as stated in some domains, but not others.
Not if they have already been shipped.
I've also seen the exact opposite being the case. A culture where everyone does whatever they're in the mood for without running ideas by other teammates who will have different experience in different areas of the codebase, can leave products stagnant for years. Tech debt builds up and the size of change that anyone is comfortable doing gets smaller and smaller until large new features become unfeasible.
If the culture on the team is to mostly work on your own without spending time designing & brainstorming up front except in rare circumstances, then nobody wants to be the only person enforcing process on themselves. You might worry that you come across as a weak engineer if you ask for a lot of feedback and nobody else does, and it does slow you down some so you'll get less work done on individual projects than your teammates (and your solutions might be of a higher quality, but that tends to show itself as the absence of problems or only be visible in the future to the next person who works on your code, and those things are easy to overlook).
So IMO it's not a matter of a presumption of incompetence vs competence, the goal is aligning incentives so that collaboration and taking the time to come up with the best solutions become the most natural way to work.
(FWIW I've never worked at Google and can't speak at all to their implementation of a design document culture).
The kind of design document that I frequently find superfluous isn't a specification of some protocol, but a prose description of the code one intends to write.
In concrete terms, an RFC might describe TCP header flags and the TCP state machine, but it'd be silent on the Linux kernel's sk_buff structure. A design document would describe sk_buff in detail.
It creates a presumption of incompetence
No, it acknowledges the reality, that people, including you and I, don't know what is good for them.Being forced to come up with design documents is very similar to forcing doctors to use checklists. They complained, and still complain, that they are professionals and don't need the bureaucracy and 'assumption of incompetence'. But the numbers speak for themselves: there are much fewer medical errors when they are forced to use simple checklists.
My colleagues are all competent, yet they produce much better work if they are forced to first come up with a design document.
"this will never work", without an actual argument, isn't Googley :)
A big design doc is a liability for sure, in the same way as code is.
Process is the scar tissue of the enterprise. Those scars are there because the enterprise was wounded. Lots of process, many scars.
But try tackling banking, or insurance, or transportation logistics, or any other big-boy problems with that attitude. You won't last long.
This!
Bad changes can be backed out, but in an infrastructure with any size, this isn't a trivial task. I've worked in a million-line codebase with multiple separate deployments, and in that case, design documents were much cheaper to create and maintain than a rollback or even writing working code and then discarding it. In an infrastructure of Google's age and size, the trade off is even more clearly in favor of up-front design.
You probably think the "move fast and break things" ideology you're espousing is agile, but it's not. It's just another plan, and agile is about responding to change over following a plan. I hope that if you ever work on a project of Google's size that you can adapt.
On occasion I've saved myself a whole lot of time when one of the maintainers agrees with the proposal and reveals they're working on something very similar already.
I speak from what I've seen in how they handle their development & releases for Angular 2, which is a fairly large project with an equally large community.
Take a look at this release candidate, which is in between other release candidate RC.5. It has over 100+ breaking changes with new features
https://github.com/angular/angular/blob/master/CHANGELOG.md#...
This is practically against against all definitions of what a release candidate should be: https://en.wikipedia.org/wiki/Software_release_life_cycle#Re... .
So this neatly painted blanket statement of 'at google' we do x,y,z for development, new features & reviewers isn't true. There are very messy projects & development practices, practically like at any other large company.