So, they don't write tests at justin.tv, and they don't do automatic deployment? Sounds like a great place to work at...
So, they don't write tests at justin.tv, and they don't do automatic deployment? Sounds like a great place to work at...
- The code assumes that an external service is highly available
- The page looks glitchy in IE
- If you use that flash api, computers with an nvidia graphics card will show a green screen
- If you open and close a socket quickly, flash will hard-crash the entire browser...sometimes
- Under peak load, when a vacuum is running, this new query will hose the DB
That said, it sounds like this bug COULD have been caught by testing (though I don't know the details and can't be sure). And tests can be darn useful, when used in the right circumstance. But in the hierarchy of important things for running a production website, here's what I would list in priority order:
- Extensive monitoring and alerting
- Detailed, easily accessible logging
- Code review required and taken seriously
- Hiring good developers
- Automated deployment and rollback
- Integration testing
- Unit testing
We've been working our way down the hierarchy :-)
2) We have 100% automated deployment and rollback, and a neat internal tool called Brigade for managing it that I really should get around to open sourcing some day...
My shock and amazement comes from the 24 hours to restart your payments system? Whaaaaat? If that is correct you guys have far, far more issues than not having tests.
If you want to be kind to your users, you stop accepting connections on a new server, wait a period of time (the longer the better), and then restart. It's a matter of how many users you're willing to disrupt for how quickly you can restart the system.
The canonical example at Justin.tv is the video system: people broadcast for days to weeks at a time. If you restart the server they're connected to, their stream will be disrupted (even if the auto-reconnect works).
We have a separate system that handles most of the complexity which is stateless, but sometimes you need to restart the actual connection-holding-daemon itself. How would you suggest doing that without disrupting service?
When someone wanted to do a deployment, we had a completely automated system that we called the cluster immune system. This would deploy the change incrementally, one machine at a time. That process would continually monitor the health of those machines, as well as the cluster as a whole, to see if the change was causing problems. If it didn't like what was going on, it would reject the change, do a fast revert, and lock deployments until someone investigated what went wrong.
When you're two guys in a garage, unit tests probably slow you down. But at some point, they begin to dramatically speed things up -- working on someone other person's code becomes orders of magnitude easier when there's a comprehensive test suite to validate your changes.
The problem is that at some point between first employee and a dozen (or more) engineers, you start to wish you'd spent a bit more time on testing code. Writing those tests from scratch becomes "a project", and it only gets done with herculean effort.
I think a lot of people also see tests as slowing things down but, they also provide a lot of value in documenting intentions, make upgrades easier, etc.
There's a lot of counter-intuitive stuff that happens in production code for a big website, and it's not totally fair to throw inexperienced people at it without at least some safety belts in place.
I think, as much as I dislike pair-programming overall, that's something I would prefer to use as a safety belt here (because it's only expensive for the period during which a new employee is still learning). Pair programming hasn't ever been tried at justin.tv either though, as far as I recall. I'm probably going to give it a try at ZeroCater for new engineers actually (we're hiring!).
Being able to change core fundamental components of a system and trust that everything will Just Work™ when you're done is an awesome feeling.
I think most small teams start with minimal qa and testing (not the recommended approach) and thus things like this happen often. Rapid development becomes blitzkrieg.
I'm trying to bootstrap a startup now, and if it fails to gain traction then justin.tv would be one of the first places I'd interview. (Hey, I like watching competitive TF2, and they provide a hell of a platform for casts, can I say.) Now I know this blog post was sort of written for the purposes of recruitment, but it's sort of making me think twice about whether I'd want to interview there. Bleh.
Here are some things we now do at Justin.tv because of experiences like the one you mentioned:
- Everything gets code reviewed
- We have started to introduce tests (sorry Bill, they're really quite helpful in the right circumstance)
- We have extensive monitoring for all systems in case something DOES slip through
- We have fully automated deployment (and rollback)
There's a cost to all of these things, either in setup time or in constant maintenance. When you have few users, maybe they're not worth it - as you grow they become essential.
I don't think I've ever claimed they're never useful ;) But I do still believe that, out of the things you listed, they have the lowest utility for the amount of effort involved.
Personally, I would do fully-automated deploy and rollback first, closely followed by monitoring, then code reviews, and then if problems were still slipping through the net I'd tell people to start writing unit tests.