What Larry Page really needs to do to return Google to its startup roots
slacy.com
slacy.com
When you get to the point where the arguments against something don't make sense outside the context of the internal path you've chosen, you have to ask if maybe you've crossed into an area where you're very much in danger of group think.
I respect what they have been able to achieve, and I certainly respect the vision of data center sized computers. I also experienced first hand the kinds of weird rationalization that accompanies trying to mentally (or culturally) accommodate an externally forced invariant for which the original principles have ceased to be relevant.
The 'Google Scale' problem is one such invariant.
For a large part of its existence there was never enough infrastructure to support things, and if you looked at Google's financial disclosures they were spending a billion dollars a quarter on new infrastructure. An x86 class server costs about $5k, do the math.
When the Great Recession hit that expansion stopped but what was interesting was this: When Google had been growing, an App that was popular would develop a material footprint in the infrastructure, that imprint would cause additional pain if it wasn't designed to fit in with everything else. However, just buy growing so much infrastructure, Google had reached a point where some apps would never have a material impact on the total resources consumed, the infrastructure was just that big.
But the rules didn't change, you had to design to run on what constituted 'Google Scale' as defined in the present not by what it meant when that requirement was put in place.
So the difficulty of implementing the requirement scaled with the size of the infrastructure, but the kinds of products that were being conceived and deployed would never need that level of scale. Stalemate, and what was worse there wasn't anywhere inside the company to even have the conversation about whether or not the requirement still made sense. And that was the root of a lot of problems and a lot of people have reflected that up the chain.
When I was there this blog posting might have been a long missive to the internal list for miscellaneous discussions, it would no doubt become a centi-thread (over a 100 responses) and detecting any change that it produced would be difficult at best (and positive change would never be credited to the person who pointed it out, only the people who changed would get any credit). And it's not that it wouldn't cause change, the social network inside worked through these issues with a plodding but deliberate slowness and sometimes those internal groups within groups might reach out to the original instigator, but more often not. But what the change wouldn't feel like is "startup-y."
I suggested to Alan Eustace once that Google I/O was a really cool way of getting people on board with various API changes and understanding the direction folks were taking, how come we didn't have an internal version?
Lots of potential, lots of challenges. Fortunately lots of folks dedicated to working through the challenges to make things better and its always great to have Larry pounding his fist on the table to add some urgency to an already frenetic environment. Sometimes the table pounding though just made the organization look like one of those table top football games where the vibrating games board caused the game pieces to move, somewhat randomly, around the board :-)
The author started at Google post-IPO. That's hardly "startup" time. In a startup it's not like resources are given out like candy and there's massive amounts of free-time for working on 20% projects.
It's true that there are less meetings at startups, but there are other things that typical hackers find unsavory: developers also have to do sales, sysadmin, customer support and generally be much more aware of where the money comes from.
The thing that makes a startup interesting is that the company is the project. Frequent meetings aren't as necessary because the goals are usually easy to understand and fairly one dimensional, e.g. "We have one product and we need to increase the number people using it."
The way you'd make a big company more startup like would be to make teams far more autonomous and drastically increase the risk / reward gradients. Your product ships 6 months late? Your entire team is fired. You open up and sustain a new revenue channel that's paying big dividends? Everybody gets a $2 million bonus. Team leader disagrees with his boss? He can chose to do it his way, knowing he's risking his team's livelihoods.
I don't know if anybody's really crazy enough to try it in a large company.
I assume things are similar in other large tech companies.
I'll posit that the real problem with big companies is the opposite: the price of failure is mild (you get reassigned to another team and get to do more cool stuff) and the benefits of success are also mild (you might get a promotion and an iPad or something). When your actions have little effect on the outcome, your actions tend to regress to the mean. In a startup, where failure means you starve and success means you never have to work again, you have a much bigger incentive to go the extra distance.
Startups given the market for engineers in particular actually only have a temporal starvation penalty. It's not like they can't get another job after they fail.
Doing this in a large company might lead to even more innovative products because startups are so under-resourced and governed by promises to investors that they risk becoming myopic and less agile.
(Serious question - what I have seen is that the team lead gets bumped up one notch and gets to try on a larger scale).
I know of at least one team that came up with something major that asked for their compensation to be tied to their product success having their idea moved to another team because "they should do it for IBM purely". Of course that product failed for lack of vision and focus.
They need to either remove at least a major part of the overhead and constraints of these highly scalable systems, or create a well planned and organized transition path from something with low deployment and development overhead to something "google scale". This could be at least at first a completely non-technical thing, such as requiring teams to at least have a plan of how they will migrate their redis, solr or whatever data stores to whatever google is using at scale. But reading this makes me cringe, that sounds like a very frustrating environment to work at and sounds very ungoogle-like to me. It appears their public image diverges much farther from reality than I would have thought.
It is so much easier at Google to design something for scalability (which Google is mostly about) than at other companies, mostly thanks to the policies and infrastructures you criticized.
It is easy to just criticize without considering the implications of alternative policies.
For example, re: Switch to team-based distributed source control I've worked at a pretty large software company that does this. The problem is lots of teams are working on similar things and results in duplicated effort.
As someone who worked in the SRE and datacenter/cluster management teams during the same period you were there (2005-2010), I can confidently say that I agree with almost everything you’ve mentioned. If you think engineers on small projects have a hard time dealing with acquiring and managing cluster resources, try being on the team that has to resolve all of those requests. Because many of Google’s core infrastructure pieces are so inflexible and frankly not designed to be used as they are, they end up dying a death of a thousand cuts. Systemic design flaws lead to telling most teams “no” when they asked for even 5 machines worth of resources.
At the end of the day, Google has maybe 5 products that generate 99% of the revenue and operate at huge scale. Should they devote most of their attention and money to these products? Absolutely. Should they do this at the expense of all the small projects? Not if Larry wants the company to act like a startup.
Ultimately the limiting factor to Google’s agility will be its technology infrastructure, not its engineers.
Isn't that a contradiction? It's the engineers who build the infrastructure. As such I would expect them to be the bottleneck, just like in every other company that has an effectively unlimited hardware budget.
Yishan Wong (previously FaceBook's Director of Engineering )'s suggestions there are interesting
e.g
" 1. Fire a broad swath of people in the executive and management staff
I've talked to quite a few extremely talented and didn't-leave-because-they-were-incompetent Xooglers over the past year at Sunfire (and via other avenues), many of them key early employees. A recurring refrain that I hear is that Google has been taken over by overly-political managers who have laid waste to a formerly meritocratic organization where good ideas get turned into good products, and this has been deadly in two ways: (1) truly good talented people who keep their heads down and get things done are motivated to the leave the company and (2) the organization that remains becomes, by necessity, one that revolves around this internal politicking rather than productive endeavor and shipping products.
Identifying these people from above will be hard, because part of being politically skilled includes looking good to the people above you (and the more politically skilled they are, the better they look), so Larry should directly contact a thousand of the best ex-Googlers and ask them to anonymously name 5 people who are still at Google who should be fired, and using a histogram of the results, fire the top 100 names without letting those people "explain" their way out of it. Steve Jobs did something like this when he returned to Apple (except he just walked the halls firing anyone he thought sucked), and this is the data-driven equivalent: a thousand ex-Googlers is enough to even out any personal grudges, and the aggregate information is likely to be highly reliable about who has been climbing without regard for those below them or the good of the organization. "
Isn't that the point? They are there to maintain a larger point of view from individual developers. Giving every person with their own agenda launch authority is disastrous (in a large organization): Of course my project is important. My project doesn't need review, I wrote it.
Agreed that for google.com search (and AdWords & GMail) there should be a few more procedures in place, but applying those procedures to every new small launch is a huge blocker.
Again, Yahoo had this too, and then some. I remember once having a launch shut down three times -- by PR, Paranoids, and Legal.
Actually, what I wish more companies did was have annual "feature killing" and "procedure killing" parties, where you reevaluate everything the company does, and if its cost is greater than its benefit, get rid of it. So "we'll institute a procedure so this never happens again" becomes "we'll try to institute a procedure so this doesn't happen again in the next year, see what its impact is on productivity, and if it saves more than it costs, we keep it."
ANY time there's an accretive process, there needs to be a corresponding ablative process to go with it. Especially around rules, laws, etc.
Otherwise, organizations calcify rapidly.
> Make it very clear that good, small ideas matter.
This is a problem everywhere big, and I agree one million percent.
I also worked at Yahoo, so I can compare what you said to "big internet companies" as well. The dedicated hardware thing can be a nightmare. Using open source stuff internally can be a nightmare too.
Where did you go after Google?
Rather than reducing the load Google needs to better connect the work of doing good recruiting to to individual/team success. The total disconnect between who does the interviewing and what team a hire ends up working on is the big problem, not the total amount of time/effort being devoted to recruiting.
-harryh, googler from 2004-2009
Part of my frustration there was seeing myself doing 2+ interviews per week for most of my career, and my colleagues and coworkers doing 1/month or less. This isn't right. I also saw internal recruiters who had "favorites" who were googlers that were more laxed in their feedback and would more likely lead to a hire. This is wrong.
Interviewing and hiring is really, really important! But, I think that Google's distributed system actually bogs down everyone instead of doing the fast & easy thing.
How would a 3 person company hire their 4th and 5th people? Google needs to do that, at scale.
Yes, exactly. If you're doing 2+ interviews/week this should directly lead to you getting more high quality co workers on your team. While the folks who are only doing 1/month end up with withering teams. Properly aligning this incentive is the key to solving the problem.
"One way Page tries to keep his finger on Google’s pulse is his insistence on signing off on every new hire—so far he’s vetted well over 30,000. For every candidate, he is given a compressed version of the lengthy packet created by the company’s hiring council, generated by custom software that allows Page to quickly scan the salient data. He gets a set every week and usually returns them with his approvals—or in some cases bounces—in three or four days. “It helps me to know what’s really going on,” he says"
However, the key is to understand that these are simply superficial manifestations of a bigger issue: as the company grows in size, communication requirements grow exponentially, and even a slight mismatch in the talents of people will lead to serious issues. I don't think any of the recommendations will work at the scale of 20,000 engineers that Google has.
Take open source software for instance. For a 10 person start up, it works beautifully. Now try convincing 10 other startups to adopt exactly the same set of software, and you'll have a never ending religious war on hand. But you cannot also let everyone to pick their own solution, for then the 10 different groups cannot integrate with each other.
Far too many people keep complaining (I complain where I work for as well :), but the solution is not easy. By far the best approach is for people to realize that there is a need for a company to grow so much, and not anymore. However, human nature will not permit that.
The root of all evil. At Google. That's ironic.
Another great quote: "Amazon EC2 is a better ecosystem for fast iteration and innovation than Google’s internal borg system."
http://searchengineland.com/25-things-i-hate-about-google-re...
The reason he puts it as "LOVE" is directly because products aren't allowed to launch unless they can prove they have sufficient capacity. And AFAIK the policy was put in place because of a few high profile failures where a popular service launched and then keeled over from high demand.
I don't necessarily think it's the right trade-off, but it's important to note that it is a trade-off, and other people absolutely love the effects of something that I find rather annoying as an engineer.
I'm not saying you're wrong, just pointing out the flip-side.
Maybe Google should have a separate network with separate virtualized machines and no "Google stuff" for new medium-sized projects. Test whether they get traction there and then port them into "the Google-way" later.
Or maybe "the Google-way" just needs to be fixed up to make it easier. For example you could use AppEngine.
Making a "wild-west" of infrastructure (a-la an in-house EC2) would really be the right way to go. And, as you said, put small & medium sized stuff there, and make porting over a moderate but not impossible task.
With respect to AppEngine, it's often stated as some kind of panacea solution for scalability for small applications. But, as a web application developer, I see AppEngine as "Googleisms on the outside". And by this, I mean that AppEngine is also a walled garden. Can you access a MongoDB instance from AppEngine? What about memcached for caching? Solr for document indexing and search? These are all things that are easy to do in a dedicated machine environment and greatly benefit small projects, but are impossible via anything Google builds.
https://code.google.com/appengine/docs/python/memcache/using...
Only until the content farms build bots capable of liking each other. It's an eternal arms race. Google figures out a property that quality content has (links from other sites, URL matching keywords, "like button" clicks), and it's useful for a while until the content farms figure it out and morph their content to match those rules.
Also, "behaviors that lead to success" are not necessarily the same as "behaviors of the successful". Lots of folks have effectively won the lottery.
It sounds like what they SHOULD do is streamline and document the system for their engineers better. There should be an internal project started to make their developers HAPPIER and more productive. For example, you have an idea for a new project? Here's the actionable checklist. Need to launch? Please make sure all of these are checked, then you can launch. Treat your developers like you treat your users.
As for capturing people into an incubator before they leave the company? I like that idea. Except of course, one has to wonder how much this will incentivize people to quit google, just to get more autonomy and a better deal :P Not to mention, that once acquired by google, the startups' technologies are just rewritten to live in the Google ecosystem, so this seems like a waste of money... except for possibly the IP licensing costs.
What about ~2 months per year? Like a mini-sabbatical.
http://webcache.googleusercontent.com/search?q=cache:t0H7qG2...
BTW, can someone explain the first point, "Compiling & fixing other people’s code" ? The "world" here refers to other teams inside Google or external libraries?
"The world" refers to all the dependent google source code. This is primarily google-written code, but includes some open source packages that are used as dependencies.
Google's internal code management system won't let you check in code that fails pre-checkin tests. This is generally a 'best practice' sort of approach. The challenge is however if you modify a library, and in the process of modifying it you change a side effect, and other code's unit tests fail because they depended on the side effect, you have some choices:
1) You can go fix their code so that they don't have the dependency and then check-in (this can be laborious because they have to approve your changes to their code)
2) You can re-introduce the side effect so that their code continues to work.
3) You can try to get them to change their code (very hard since they probably have bunch of other things going on that don't depend on you)
4) You can create a new library (or routine in the library) that has the semantics you want and deprecate the older version.
While this was extremely painful for folks who were working lower down in the system it was not as big an issue for folks on the upper levels. And in a perverse way it motivated good interface design.
I'm not sure it really matters if a site is "developed by Googlers". Why do the users care? If Google is playing farovitism with results, then it's probably important to disclose that, but otherwise, just let startups be startups.
The big paradox that brings up is: If all internally developed projects must be developed Google-scale, on Google products, why are so many of Google's most successful products acquisitions rooted in non-Google products?
Does that make sense?
The example that comes to mind (and I may have some of the details wrong) is HP finally breaking through with DeskJet after stagnating in innovation by making the group _completely_ separate. They moved the group to a new location geographically and gave them huge amounts of autonomy to make something outside the bubble / group-think of their LaserJet juggernaut.
I have a personal example of this as well. I worked as a PM for a real estate software firm that owned 70%+ of the market for desktop software for real estate appraisers. 300+ employees with a lot of engineers. They decided to move into the real estate agent segment with a new product, and to avoid the group think and "our company's bread winner and primary focus is on X", they started a satellite office in another state. It was like a startup with occasional oversight from some investors and board. The corporate values were the same, but our autonomy allowed us to truly innovate.
brand matters a lot, especially when trusting a web app with private data is concerned.
more than that there's brand loyalty. i realize this is an extreme example, but i don't use dropbox because it's the kind of internet-feature i want under my google account, not somewhere else. i literally don't want any company other than google to succeed in that space.