The other half of "Artists Ship"
paulgraham.com
paulgraham.com
In a startup there's noone to blame but you, and thus no need for politics. As anyone who has worked in middle management or above in a large company can tell you half of the job is making sure you can always find someone else to blame, while taking credit for the successes of others.
I wonder if the problem at the recently troubled companies (AIG, Merrill, Lehman etc) was not enough checks in place or not enough power allotted to the employees who knew something was wrong.
This mentality, reenforced by the company's bureaucratic change management system, really did not sit well with me.
Fortunately for my summer experience, my "buddy"/summer mentor and I found a loophole which we used to "hack" the change management system so that once we got our initial approval, we could propagate changes to prod without running through the whole process again for each change (which otherwise would have been required).
Look at some of the major Web2.0 kerfuffles in recent years. Off the top of my head, I can think of:
- The HD/DVD mutiny on Digg
- Public suicides on both JoelOnSoftware and Justin.TV
- Reddit storing passwords in cleartext and them getting stolen off a laptop
- Ariel Waldman and the Twitter harassment fiasco
- The Flickr censorship debate here: http://www.flickr.com/help/forum/40074/page3/
Yeah, you could argue that those services are all still around, so obviously they haven't been hurt too much. How much management time was wasted in dealing with them? How many non-users decided not to become users because of something they heard 3rd-hand about how it's a terrible company?
(And I don't think the solution is to never innovate. I can understand why a middle manager at a big company would think so, though - a PR disaster is by definition public and disastrous. In the short run, it always makes more sense to not mess with success, it's just that this doesn't lead to success in the long run.)
I've dealt with the security teams at large companies. They end up implementing sometimes draconian policies, but it seems that programmers refuse to learn to write secure code any other way. I have to admit that, if you consider the large company as being a legitimate entity, they're doing legitimate work.
Most of these incidents caused by employees taking out the data, but there are some cases caused by sloppy web app.
(I could find only Japanese version of the report: http://www.jnsa.org/result/2007/pol/incident/index.html )
A single individual - one who is not even in a high position. But you bet your buttons that it'll be all over the internet within ours.
Sure, it probably wouldn't even cause a blip to the stock price, but it's a huge amount of harm for a single person to cause.
I once worked for a company where anybody can push changes to prod without much feedback whatsoever. Certain more senior guys watched the commit logs for any obvious shenanigans or idiocy, but generally changes were within customers' hands within a day or two (and this was a desktop app!). They suffered some horrifying stability problem for a while that necessitated some new policies. It was certainly liberating to work in that environment, but I can see why many places have safeties in place.
And this is an error that one unchecked programmer would be perfectly able to make if there were no checks in place.
Case in point: I work for a $3Bn company and a year ago I developed a feature that saved the business ~$2M that year. We know the amount because when the problem raised its head we called together a number of people to calculate what the cost would be and estimates were between $1.5M - $3.5M if we couldn't find a solution. Not loss of revenue, but we would have actually had to pay out that much in support costs. The problem was highly visible across the organization and I was able to find a way to solve the it in a way that was completely invisible to the customer, implemented it and released to the field. My manager looked good, his manager looked good. I got an "average" performance review. I was pissed and to this day I haven't regained my motivation to do that caliber of work again. I wonder if I ever will without leaving!
But here's another twist: when you start working for a giant company that provides myriad services under the same login, suddenly everyone has to move as slowly as the slowest part of the business. It doesn't matter if you are just doing Yippee! Backgammon, because who knows -- maybe you could somehow accidentally expose the data at Yippee! Payment Solutions.
This is a completely rational consequence of being a giant company and delivering multiple web services under the same login. It is possible to have slightly different authentication policies for each service. But I wonder if OpenID providers have thought about this enough.
He has made that point in other essays.
I don't think the claim that they would not have broken anything can be backed by facts.
If you look at Microsoft, for instance, their programmers used to write very buggy code, but with the introduction of better processes, like their Secure Development Lifecycle, they saw a substantial improvement in code quality (as measured through security and reliability metrics).
The damage done by not following a quality assurance process and writing buggy code was so big that Microsoft's image will be affected for a long time, even though they have now improved the code quality.
Even these particular guys certainly would have written buggy code. A QA process could correct it.
So I have little respect for formal QA.
Also, bugs do appear in production if you can release code the same day you wrote it, but you can fix it as soon as you find out about it. In big companies that work with milestones (sprints in agile speak), only really critical bugs can be fixed between releases. Not to mention that most bugs are fixed after code-freeze, when all the deliverables of that milestone are ready, long after you find out about it.
So I guess it's best to have QA, but also the freedom to deploy code on production whenever you want.
A two-week release-process lag is so much worse than it might sound. If you can ship instantly, you can get feedback and make improvements rapidly, over and over. Sure, without a careful release process you might occasionally break things, but when you do you can fix them in minutes rather than weeks. When you insert release-process lag into that cycle, improvements that might have been made in a rapid series of releases over a day or two can take months. Sometimes those improvements just never get made, because no one is motivated to keep working on them for so long.
In software, "stable" often means "unchanging", but only rarely means "high-quality".
As companies grow big, they like to emulate larger companies, and hence the people they like to hire are from larger companies since they understand scale -- this comes with a lot of baggage ofcourse. This in turn, most times, brings with it a culture of people who are worried more about not screwing up, covering their own ass, and thinking about getting promoted rather than doing what's right for the company or the business and screwing up as a byproduct of making quick decisions.
Re Joel Spolsky: The cost of the sale is not only the cost of the product, but what an organization would pay for a software/service, and what they would value it at. This is why most companies have a sales force, and don't advertise their prices on their website. As Joel mentions, for a lot of companies, charging less than $50k is a rounding error and not worth their time.
The flip side of people at big companies buying is based not entirely on the performance of the service or the product. As long as the service/product is reliable & ok, and it doesn't screw up in any major ways, they won't get fired for making that buying decision, rather than looking at ways of maximizing/optimizing the performance of the product -- yet another reason committees make buying decisions.
Re SOX: I like what founder's fund + facebook is doing -- letting early employees cash some of their equity out, thereby increasing the time early employees will stay with the company.
I also like Fred Wilson's thoughts on a secondary market for startup stock (similar to what goog does), this would in some sense, I guess let you be a private company, and not deal with the challenges associated with SOX compliance and at the same time not raise a highly dilutive Series D, in addition there is some sort of liquidation event for the investors.
PS: (std. Buchheit comments about Limited Life Experiences + Overgeneralization = Advice apply)
PPS: Say hi to the reddit guys for me ;-)
In terms of time, they have a policy of permission vs. forgiveness. You need to be prepared to fail early – but that’s okay! If a prototype takes less than two days, don’t worry: just go ahead and do it. It if takes more... you should probably have permission.
This actually made things worse. Aggressive releases were often the result of a decision further up the chain that something was "critical" and needed to go out right then and there. We'd got into the pattern of expecting to be able to do this at the drop of a hat.
So now, everyone was rushing to get the "urgent fix" into the daily release. This is an effective way to ship buggy software. Extra emergency releases were frowned upon, and so it introduced a level of panic that "Oh shit, this better not screw everything up". Ironically, with anyone releasing whenever they wanted, there was less panic because you could always make a quick fix and put it out without too many people noticing or caring.
We felt as a team that we'd prefer a more relaxed schedule that also allowed testing and accountability.
Frequent releases were a result of bad management, and unrealistic stakeholder expectations, we as developers didn't necessarily care that our code wasn't being put out immediately. Obviously, it needs to go out at some point, and we've been equally demoralised by mammoth projects that have gone on for months without a release, but that's a different end of the spectrum. The important thing is regular releases, not frequent ones.
We've now decided to move to releases every 2 weeks with a 1 week QA period. We made this decision as a team of developers, it wasn't handed down from above, it just became a necessary check for the scale of our company.
I personally feel a lot more productive as a result, less panicked by the quick fix you need to drop everything to fit into the daily release, and less worried that a release will take out the site.
But sure, it comes at a cost. There's now a lot more going out with each release and more variables that could combine to cause issues, but these issues highlight deficiencies in our QA process that can be fixed. Also, getting our team to work according to this 2 week schedule has meant some significant costs to adopt better development processes, namely SCRUM, but I feel our team has benefited enormously from it. It all depends how well your team can adapt really.
So yes, there are costs associated, but it doesn't necessarily lead to demotivated developers. You need to be sure that the checks are introduced properly with their own test suite. This itself has a cost but can be extremely worthwhile. A half arsed and arbitrary change to your release schedule with no process changes can destroy you, but with a little training, the right team and the right managers, you can make things better for everyone.
Imagine you have to fix a simple bug, like the description of an option in an HTML page.
Imagine that: the codebase is a bajillion lines of Java or C++ that no one person truly understands any more.
Imagine that: fucking up is considered bad. Maybe you're protecting important user data, or maybe it's just code for the online store, and even an hour of downtime can wipe out the day's profits.
The team is so big that your feature is scheduled for deployment along with four or five others this week.
As usual, management has skimped on QA so there are not enough of them to go around. Also, QA was added late, so there's no fast automated test suite. It takes them a full week of manual and semi-automatic testing to go through all regressions.
After that week of testing in development is over, then it goes to staging, where it's tested with real live data and staging subsystems.
At least somewhere in this, something breaks (something is always breaking) and a fix is produced. Then the new build is tested all over again, with the basic smoke tests as well as any test suites that are applicable to that unit of code which broke.
This is causing another trainwreck because we promised feature X to client Y by end of the month. What do we do? Emergency team meeting! Let's rebalance the work schedule.
Meanwhile, your change, which just a simple HTML fix, languishes as this build goes into its second round of testing.
Oh, and did we mention that you have to have submitted your code in time for translation into six languages? I'm afraid that the French translator did not produce a satisfactory translation in time, so we're going to have to hold your change back. Can you revert this and put it in next week's build? Thanks.
I guess it would be possible to note in a bug tracking app "we would have deployed here" and then see how many, and what kind of bugs were found between then and actual deployment, which checks found them, and whether any of those were rollback-causers.
That would give an estimate of how much good the checks did in each case, and a way to choose which checks are ineffective for their cost. (Ish - some checks may only trigger once, ever, to be priceless).
I wonder if there's a way to limit the total cost without hard-and-fast silly rules.
e.g. you can only add a new step to a process if a step is removed from another.
Yes. ;)
When do you institute the rule -- on day one, in which case it is functionally equivalent to "do not have any checks"? On day 13? Day 7386?
What is the quantum of a "check"? Is Sarbanes-Oxley one "check" or a collection of hundreds of "checks"?
Does the limit apply department-by-department or company-wide? Your sysadmins will naturally tend to employ a lot more formal checks (most of which are hopefully enforced by tiny Perl scripts) than your R&D prototyping team. Do you force the sysadmins to compete with the accounting department for a quota of checks? Do you incentivize sysadmins to evade checks by getting tasks reassigned to the R&D prototyping team, which (alas) can't afford to have any checks?
The last thing you want to do is encourage teams of lawyerly meta-checkers to run around enforcing rules about rules. That's costly, squared.
When you realise you have a (potential) problem. It stops it getting worse.
"What is the quantum of a 'check'?"
More tricky, certainly. That's a potential big problem.
"Does the limit apply department-by-department or company-wide?"
Whatever makes sense.
The idea of the exercise would be basically to:
- build an appreciation throught the organisation that checks have a cost
- provide a (albeit imperfect) mechanism for exercising some control over it
Something doesn't have to be perfect to be of use.
"The last thing you want to do is encourage teams of lawyerly meta-checkers to run around enforcing rules about rules. That's costly, squared."
That's just ISO9001, isn't it? :-)
Measuring this, though, is the hard part. It's like trying to do the math behind "no silver bullet" and estimate communication costs (too many unknowns and conditional dependencies).
Big companies on the other hand, start with a company worth $100,000,000, and want to turn it into a company worth $150,000,000. They value steady process improvement that doesn't risk their existing revenue -- that's why they want good managers.
It's important to know this if you're trying to sell to a big company as well as if you're trying to compete with a big company.
That's where you lost me. Sure, there are companies who add checks without thinking of the consequences, but as a company goes from a startup that is not making money to a more mature company that is making money, the cost of breaking things becomes very, very high. That is why most post-startup companies add more checks, to protect their revenue stream. And any company making decent money has a revenue stream that far outweighs the amount of money that the programmers would be willing to pay to have a faster release cycle. And the risk of the company losing their revenue stream for any significant amount of time far outweighs the risk in losing decent programmers because they aren't happy with the length of the release cycle.
You seem to say that these particular programmers would never have broken anything, but I find that impossible to believe unless they are working on something that is not at all complex or is inconsequential to the revenue stream. Every programmer, even the best programmer in the world writes code with bugs in it. The more complex a system is, the more likely there are to be bugs in it and as a company grows and makes money, the more complex their systems will become.
As an example, let's say that you were responsible for a web application that brought in a million dollars of revenue a day and you had these supposedly perfect programmers that never made mistakes who were responsible for the code. Would you let them just throw out whatever code they wanted onto the servers because you trusted them to not make mistakes? Or would you be more reasonable and put some checks in place first to make sure that the application worked correctly?
Obviously there needs to be a sane balance between the risk of breaking the application and the opportunity cost of not updating the application and the frustration of slowing down development is one of the costs that should be weighed, but to say that there shouldn't be any checks added as a company matures is almost laughable.
To me that's the most important paragraph here. That takes you from "let's make sure this never happens again" to a cost-benefit analysis. The issue is not whether or not to have checks, as much of the discussion here assumes. It's about realizing that not all checks are equivalent, and using that knowledge to get the greatest safety at the lowest cost.
Steve Yegge has had the unenviable luck of working on 3 products all of which have not shipped (cited as business reasons) yet he still works for Google ~ http://blog.stackoverflow.com/2008/10/podcast-25/ So it seems there is something else going on that keeps programmers on board. Is it money, having the freedom to talking about it the process? I don't know.
I worked on it because I love programming and trying new things. Everything I was doing at the time was new to me so it was fun, engaging, and a learning experience. Plus, it beat the hell out of anything else I could have been doing in that wasteland ;)
Good explanation.
I put it as a question as I'm wasn't conclusively sure why people would sign on to assignments that have a probability of failure. But I've since found that smart organisations allow pursuit of training and learning as a reason to keep staff and motivation. I figure the USF know a thing or 2 about motivation. Hope your putting the lessons you learned into some hack your working on :)
Surely if the costs of checks could be "discontinuous," "non-linear," "step-change," or however else you would like to term it, then perhaps the potential cost to the company of NOT having the check in place could also be the same. Larger companies potentially have more to lose by having recurring lapses in supply, or in the ongoing case of your essay, software service. Of course, software/coding may also be a special case. Additionally, it's true that your "committee" example represents a discontinuous change in bureaucracy for that firm, which would make it much more reasonable to expect an equivalent affect on cost.
In the end I think the solution to both the article's and my concerns are the same: It is always better to ask the members of the team that experienced the problem firsthand what they think the solution would be or what they would do differently next time to avoid the problem. Especially if you phrase it as "what could you or your team have done differently to improve the outcome" and urge them to avoid proposing major process initiatives if at all possible, the improvement will be more practical and lower cost in all senses.
Programmers, though, like it better when they write more code. Or more precisely, when they release more code. Programmers like to make a difference. Good ones, anyway.
Salesman like it best when they can sell more. When I used to sell used cars, I loved it best when I could keep pushing cars out of the lot.
Writers like it best when they can write more. They like it best when their imagination can produce thousands of stories.
The difference is that coders usually know what they will tackle and have a strategy so it is easier for them to get to the end point. Also they are almost sure they will hit the end point of a certain problem.
With other professions you cannot predict much. Your next customer is not guaranteed, your next story just does not want to show up in your brain etc....
That being said, Yes, a lot of programmers I know work hard.
Also, since most workers don't love their job, it is probably also true that good programmers are unlike many workers. To say that programmers are unlike many TYPES OF workers we'd have to believe that a higher proportion of programmers love their job than do workers in other types of work. From what I have seen on HN (albeit sample biased) I'd say that is true.
People like to make a difference. Good ones, anyway.
They like to get paid for it, they can even be greedy. They like to get recognition for it, that can go to far too.
But they also enjoy making the sale.
to me, hacking > sales (no, charity/non-profit/giving away money doesn't count -> there's moral incentive there)
Programmers are unlike many types of workers in that the best ones actually prefer to work hard. This doesn't seem to be the case in most types of work. When I worked in fast food, we didn't prefer the busy times.
Who are the people who are "best at" fast food? I think perhaps you mean that people who build things get enjoyment out of building -- perhaps more precisely, of seeing their built item in action. I think programmers in particular are used to having short feedback loops between building and trying out, and that pattern of building up and testing as you go until a larger problem is solved is a strong motivation.
I guess I'm claiming that harder-working isn't really the issue. The issue seems to be who is having more fun. And in that sense, your claims still make sense. Part of the fun of building is a short feedback cycle. More checks in the process make the feedback cycle longer, therefore less fun. That explains why teachers' like their jobs. You can see it in someone's face when they learn something new.
I have two pieces of evidence.
First, it was well-known and easy to see that Sarbanes-Oxley would be extremely expensive. It was well-known when passed that its costs fall disproportionately on small public companies. Because IPOs by definition are small public companies, the costs fall disproportionately on them. Thus, it must have been known when enacted that the bill would harm IPOs.
Second, suppose in fact it is the case, as Graham seems to advocate, that Congress "inadvertently" harmed IPOs. His argument is that Congress just accidently happened to overlook the harm of the bill to IPOs. In that case, when the harm to IPOs became factually clear, the bill would have been changed or amended. The fact that the law was not amended even after the harm to IPOs became clear proves that it was Congress' intention all along to harm small public companies (the companies most dangerous to the large corporations who have the most lobbying pull).
Working at Motorola on cell phone code could be a great way to give everyone in the modern world a cool new feature - or accidentally break 20+% of all phones out there. A terrible Mickey Mouse movie could destroy a brand. Take it a step further to NASA, vehicle engine design (and safety), oil pipelining and refining, etc. and the cost could be human lives.
That said, I'm very interested in the cost of these "checks" on those industries. Not monetary costs (although important), but rather the cost in stifled future innovation.
Thoughts?
Once you get large, and your company is providing sustenance for all your employees, people rely on your product, and you've got too many users for effective feedback-checking, you have to close up, take fewer risks. Because suddenly, people want you to move slowly. They don't want you constantly skyrocketing ahead with their playing backup. Look at any big company - even Google, which was once famous for moving quickly - and you'll see that part of what gives a big company a good reputation is its being "solid." They have to give things up for an advantage.
It's why newspapers are so relied-upon. Of course, now it's what is hurting newspapers the most. They're being beaten by the flexible Internet. But even there, we're seeing a trade-off. Look at the quality of stories by the top writers online and by the top NY Times writers, and the online writers are much more amateur. They're faster, occasionally they're more interesting, but the Internet is thus far not retaining a high level of professionalism among reporting. Similarly, start-ups are much less reliable on the whole than large companies - look at Twitter and its problems, for instance.
So I think that PG's article is right. You can't restrict people and expect them to do as well. However, too much freedom leads to less stability, so it becomes a trade-off. Everything in moderation.
If you have the agility to make rapid production changes, you also have the ability to rapidly rollback. So the argument that larger companies require more checks and testing than startups isn't really valid, especially when you consider the costs.
This is just not true. Rollbacks are always more expensive than changes, because you can't rewind time to undo the consequences of having your software be broken for minutes, hours, or days. Worse, in the absence of "checks", the cost of making a production change tends to be roughly constant as the company grows -- it takes the Amazon sysadmin no more time to type "make deploy" than it does me -- but the cost of a rollback scales directly with the size of your company's customer base.
Within a few seconds after Amazon.com breaks S3, thousands of companies begin to lose money, and they lose money second by second until the rollback happens. Even if Amazon is only down for a minute, that's one minute of downtime multiplied by its number of customers. The larger the customer base, the larger the stakes.
And, unfortunately, the cost of downtime is nonlinear. If Amazon goes down for a mere two minutes, hundreds of peacefully sleeping system administrators will get emergency pages from their uptime-monitoring systems. They will get out of bed. They will check their logs and their failover mechanisms. They will lose a lot of sleep, and soak up a bunch of overtime pay, and a lot of their good will towards Amazon will dissipate like the morning dew. Once you lose your reputation for quality it takes a lot of work to get it back.
This is why larger companies have more controls. The controls are in place to try and pass the ever-increasing cost of a rollback back to the team that causes the rollbacks. The reason it seems so gosh-darned expensive to add a trivial feature to your flagship app is that it is expensive: If the average rollback costs $1m in revenue and every new feature is only 95% reliable, every new feature costs the company $50k to deploy.
The secret here is: If you want to deploy changes rapidly, don't work on a product that has a lot of uptime-sensitive customers! Start a different product line, or start a beta program, or found a smaller company.
Let's say I own a video site and I want to add threaded comments. If I have 5 users and the site goes down for 5 minutes, those 5 users will get 5 minutes each of annoyance. If I have a million users, each of those users will get 5 minutes of annoyance each also. There is no difference to the user there. So, by adding more checks to make sure the site doesn't go down for 5 minutes when you have more users, you're saying the more users you have, more the important each user becomes. I think that's a strange way of thinking.
(The same is true here of an infrastructure service-- if S3 had 5 users and were more cavalier about their release schedule and broke something, those 5 users would exact the same net effects of downtime as if S3 had 5 million users.)
The awesome benefit of getting threaded comments developed, tested briefly, and pushed in one evening is worth the risk of 5 minutes of downtime compared to the 2 weeks of rigorous testing and approval-by-committee. No matter how many users you have.
If I have 5 users and the site goes down for 5 minutes, those 5 users will get 5 minutes each of annoyance. If I have a million users, each of those users will get 5 minutes of annoyance each also. There is no difference to the user there.
No, but there is a big difference for you! If a user is worth a dollar per year, the five-user site is worth five bucks per year, but the million-user site is worth a million bucks. If each patch to your code causes 0.1% of users to abandon your product (a number which depends on the odds that a patch will cause a rollback, and on the odds that a rollback will annoy a user enough to make them leave), patching a 5-user site costs you half a cent per year on average (most likely it has no perceptable cost, since odds are no users will leave) but each patch to a million-user site costs you $1000 per year in revenue. And that's just the linear cost. There are nonlinear consequences: one or zero annoyed users is nothing to worry about -- unless that user is Michael Arrington -- but a clique of 1000 annoyed users is potentially a movement: a critical mass of people who will all start complaining about your company on Twitter on the same day, potentially costing you your next 10,000 or 100,000 or 1 million users while simultaneously empowering your competitors, who may begin building the site that will take you down by poaching those dissatisfied users.
This is just the flip side of scalability. As a programmer you enjoy mighty economies of scale: Running a site with a million users is more expensive than running a single-user site, but it is much less than a million times as expensive. But this leverage also applies to your mistakes: a mistake that costs you a dollar when your site is small might cost you $1,000,000 when your site is big. And it's the same mistake! Typos are just as easy to make on big sites as on small ones.
Obviously, this doesn't mean that you shouldn't ever change the site. Presumably each and every one of your patches is valuable, and will bring in revenue to pay for its own insurance premiums. Right? :) But you do need to think about that calculation, because you do occasionally make mistakes. As your userbase grows, you may wish to test each patch on a subset of users to be sure they will really like it, and that the additional revenue is really going to be there. You may wish to institute tests and internal audits that lower the risk of rollbacks, or failover mechanisms to lower the cost of rollbacks. And before long, lo, you will be that which you deplore: A company with a bunch of annoying internal controls! But at least you'll have revenue to console yourself with.
I agree with both of you that it varies considerably based on what the site does (infrastructure, videos, games, etc).
I'm not going to argue with that. Just because a certain increase of caution is rational doesn't mean that caution isn't being overapplied in many cases, just as PG suggests in his original post.
You can change and roll back features quickly, but that gives an impression of instability among users. If things are constantly changing, they'll seek out something more stable. It's "lowest common denominator" thinking. It results in something that nobody dislikes, and that's the goal of larger groups. Niche companies are able to work much better, but even then there's some slowdown.
The big issue is that the level of fear increases as you move up the command structure.
Let's assume one of your startups has crashed and burned. Something really spectacular must have happened if this makes potential customers of your next startup spurn you. Taking risks may have consequences, but they are temporary. But if you are a high-level manager for a large company, spectacular failure will make you unemployable, or in the best case push you several steps down the ladder - steps which took years of boredom and politics to climb, and which you will never get back. Failure is punished disproportionately, success is rarely compensated. Conservatism is the correct choice for each individual actor.
We all know bureaucracy, culture and politics - existing companies won't change in this area.
It doesn't matter. If you make a mistake with other people's money (e.g. calculating a payment wrong, crediting/debiting the wrong person, etc), even if you put it right quickly, they'll start losing trust and looking at your competitors.
For example, competent people who knows their own weaknesses well and often check their decisions with others should be immune to checks
Good programmers often let other good programmers check their codes. Good writers often let other good writers proof read. Good managers often discuss his decisions with his employees before deploy them. They should be check immune.
You can even pass out check-immune badges to encourage non-check-immune people to be more critical of themselves :)
Of course, this system itself is a check. Who decides who will get check-immunity? What's the successful-decision-ratio threshold? I think it's possible to implement a light-weight system little cost.
In a startup, it's acceptable for service to be accidentally down for short periods while the geniuses running the show fix their latest cockup. In large companies, it isn't. How do you get the best of both worlds?
It is not by removing all safety interlocks. Sysadmins, especially good ones, can tell you in excruciating detail why not. Most programmers are not very good sysadmins, or at least not as good as they might like to believe. Coincidence?
I would call it part of institutionalisation. You have a need to be effective via policies, checks, procedures & such. This replaces the shoot from the hip of small groups.
I think you can apply this to schools that need to teach approved courses with approved grading. Then it gets worse when they need examinations that are to be applied across an entire country.
I'm sure there are principals that can be extracted & applied to lots of places. The costs can be very serious.
Interestingly the biggest company I worked for actually ran the best because perhaps its roots being 80+ years before any of the others would allow the significant difference to be tracked.
Well, only if they still buy every product, which would defeat the purpose of having a gatekeeper committee. If, instead, the committee approves only 1 in 25 products that would otherwise have been bought, the committee saves them money (leaving aside the cost of the time of the committee-members).
My experience as a manager of a team of programmers for many years is that for sure they would love to be able to write and release w/o a QA process, they do in fact take the site down when this is allowed!
Mental Poisons are ideas that harm people, life and the universe. They are remarkably common.
Paul expresses this one very well. There's a lot more out there too.
Fossilised heirarchies and thought-free-rule-driven organisations are great for the student of poisons.
http://computinglife.wordpress.com/2008/11/14/hands-free-or-...
I'm surprised you didn't tie this in to due diligence checks done by VCs, and its impact on startups, etc. The parallels are strong, and you've commented on that impact before (at FOWA in 2007, I think).
Regards, Terry Jones
I think it is very cool to have the comments here instead of his website. It is nice to have them consolidated and not some comments here, some there.