Continuous Deployment at IMVU: Doing the impossible fifty times a day.
timothyfitz.wordpress.com
timothyfitz.wordpress.com
Also, it's clear to me why your daily routine might sound like science fiction to the median HN reader: A lot of programmers have never seen a system like this. As those of us who were online during a specific half-hour period a couple weeks ago can attest, even Google doesn't have a system that's remotely as reliable as this: It appears to be possible to break all of Google search, worldwide, in ten minutes by misplacing a single character in a text file.
Meanwhile, I'm sure that the original submitter would agree that tests ain't perfect. If you read the link at the top of this blog post:
http://timothyfitz.wordpress.com/2009/02/08/continuous-deplo...
...you'll find that this isn't merely an article about automated testing. Automated testing is just a part of the mighty continuous-deployment ecosystem being described here. It isn't even the real heart of that system: The heart is a planned, well-designed, semi-automated routine for rolling back changes in production. They roll out a change to a subset of their servers, monitor for statistical anomalies in the usage patterns of real, live users, and only continue the rollout if there are no anomalies. If they run into trouble, back they go.
And I agree about resiliency of the deploy -- it's what I meant by sophisticated testing of these momentary guinea pig users. Google's presentations on this stuff are about analysis and data gathering of changes both for immediate functional snafus and user preference for changes. i.e. probably state of the art in this regard.
I have often posed questions here about things I've been doing for years just to see what others are doing or if there is a better way. Invariably, someone tells me it won't work when I already know better.
This tells me 2 things. I've encountered someone who speaks when they should be listening (what else is new), and, more importantly, that I'm pushing the envelope enough to make otherwise knowledgeable people uncomfortable. Good.
I also didn't mean to rag on news.yc, on average the responses here are better than the original posts which is an unparalleled level of quality.
I left the technical issues unspecified in my first post on Continuous Deployment, and the comments on that post already had started discussing the path of solutions we ended up building ourselves!
I still have another post to write about this, because we also ship a native windows client. We ship daily prereleases of it, and roughly biweekly full releases (offered to all users). That's close to, but not quite as impressive, as the update system that google chrome uses. Google definitely still has us beat in certain categories.
How can we conclude that the same isn't true of IMVU? The fact that such a rare event hasn't happened to them yet tells us very little.
It may be hard to imagine writing rock solid one-in-a-million-or-better tests that drive Internet Explorer to click ajax frontend buttons executing backend apache, php, memcache, mysql, java and solr. I am writing this blog post to tell you that not only is it possible, it’s just one part of my day job.
For test strategy in general, you need a whole seminar. There's a good blog on testing over at http://googletesting.blogspot.com/ <- go there!
As for code:test ratio, I'm not sure what metric you want, but using just php files, there are ~1k active test files and ~7k php files total. So if you take into account that 1k of that 7k is the test, and some is test-setup and obsolete files we no longer use, it's maybe about 5:1 code:test. However, a lot of the code is third party OS software packages. In-house code is about 1.5:1 test:code, in my experience. These, by the way, are very rough estimates.
I find this unparseable. (English is not my native language). As far as I know "one in a million" means something like "very rare". Help?
He means that test failures are not acceptable in their culture, and they have a LOT of tests.
Our Internet-Explorer-based tests are very reliable; they fail less than once per million executions.
Let's say your team makes 25 commits per day.
25 * ~300 working days = 7,500 commits per year
That would take 133+ years to reach 1 in a million.
The more interesting metric to me is how often the build gets broken.
Edit: of course that assumes a peak commit rate matching or exceeding the commit-test cycle period. The point being that even a considerably low rate of failure in the testing mechanism could manifest itself as a blocked commit-test-deploy cycle at least once a day, hence the importance placed on rock-solid testing systems that should only ever fail when the tested code itself fails.
It just seemed to me like you were bragging that tests get run over and over again. They only need to get run if any new code is committed, of course.
And what kind of commit is being checked in every 9 minutes? How big is the dev team? Seems like an awful lot of commits. Is each one a full-fledged feature / bug fix for the site, or are many 1-line changes to the code?
It's true that any particular test that spuriously fails one in a million times may never fail. But if you have tens of thousands of tests, and you do tens of test runs per day, you'll have a test spuriously fail once a day or so.
A necessary, but not sufficient, requirement for a nimble start-up.
I think there's an idea that if something goes wrong because you let an automated system do it, it's somehow much worse than if something goes wrong because there was human error. I don't really understand the reasoning.
The ultimate solution is to have business metrics drive your UI changes, usually in the form of an A/B test. Then you have a clear winner. This A/B would be run separate from the roll out structure (and indeed, we do LOTS of A/B tests).
Sometimes that's not possible, for a new feature or for content without a clear business metric to evaluate for. Either way we often have someone manually test new UI, so that we're not exposing users to something fundamentally broken. We usually do this by using the existing deploy system, but turning the frontend on only for QA users.
In the end, you do what works and is cheap, and that's usually something slightly different for every project.
A) You can quickly fix it reducing the number of people that see that bug.
B) You can probably release the fix without breaking other things.
Is it better for you to release 5 bugs that 5% of your users sees for a few hours or release 1 bug that 100% of your users see for 2 weeks?
The major problem is rolling back client side changes (that are located in scripts or CSS). This is pretty costly to rollback, because of browser cache - we solve this by having real versioning of the static files so we can force a refresh of browser cache (real versioning = script_{timestamp}.js and not script.js?v={timestamp}).
The major advantage to using "script.js?v={timestamp}" is that it maintains a consistent URI for the resource. Whereas with "script_{timestamp}.js", everything that points to it needs to be updated every time it changes.
You could create a symbolic link or rewrite rule that directs requests for "script.js" to the latest "script_{timestamp}.js" but it's more convenient to use a URI parameter.
Also, if you ever move to a CDN, then you are forced to use real versioning (at least with Amazon Cloudfront).
The versioning scheme we use is `md5 hash of name + file contents + file extension` (and not timestamp).
Moving to a CDN does not force you to put versioning in the path or filename. The URI parameter merely tricks the browser into thinking there is a new file. The parameter itself is otherwise ignored.
Unless you specify "Cache-control: no-cache" header you aren't really sure how the browser caches your static files (especially if the user is behind a proxy - and even "Cache-control: no-cache" can easily be ignored).
The urls would be something like:
/static/r12345/foo.js and /static/r12345/foo.css
This describes pretty well what we do at Justin.TV too. These days I push new code about 5 times a day.
Were you saying don't write automated tests that test your code, instead focus on monitoring the actual production invironment?
Or were you saying that specifically the "unit test" class of automated tests are not worth their time?
I can imagine a system that monitors the business metrics well enough to prevent defects from slipping into production (it's a stretch, metrics are soft and squishy moving targets), but I can't imagine using only those metrics to find every bug you ever slip into production. Metrics are so distant from the bug that caused their downturn; you'd waste so many cycles debugging. The gap between writing the code and finding the problem would be much larger than if unit tests found them; that has to slow things down as well.
- Monitoring the production environment, tons of effort. We record and analyze an incredible amount of data about everything that happens on the site, and have more and more automated processes looking for anomalies (though still nowhere near as many as I would like).
- Automated testing not including unit tests, some effort. I wouldn't be opposed to us doing more of this, but it's not incredibly high-priority and there always seems to be something else that's more important.
- Unit testing, yeah, not worth our time as far as I'm concerned.
I guess you could make the frontend code aware of the code version, include it as a param with each XHR request, have the server check versions and return a "version mismatch", and then produce some alert on the browser asking to refresh the page. But this would tradeoff far too much usability.
Then when the AJAX stuff saw a version mismatch, it would wait until the user completed any operation that -wasn't- stored in the fragment and put up an "updating, gimme a sec" box, and refresh itself.
It was a hell of a lot of work but -extremely- slick (which I'm allowed to say because it wasn't me who wrote that part ;)
What's not horrible is having thousands of tests, on dozen of machines, 9 minutes to-live, with selective updating of users, and rollbacks, as this article has explained.
The original post was too light on details, I guess. Its intention was not to be comprehensive anyway, the focus was why recently changed code should be put in production ASAP. But it looked like the author was simply FTPing after commit. And the whole "SOMEONE IS WRONG ON THE INTERNET" thing kicked in.
I'm also one of the developers on a hobby project called http://TIGdb.com (Jeff Lindsay is the other, and has written the majority of the website) We don't have a big Continuous Deploy infrastructure, but we also don't have the users and business requirements of IMVU.
We started with the usual, completely manual deploys and hard-to-setup sandboxes, and have been iterating towards a fully automated setup ever since. The entire time we've been doing this, we've been committing and deploying often. Our users are patient, because we're giving them something they can't get elsewhere and we're giving it to them for free. As we do introduce regressions, we'll post-mortem them (probably using the 5 why's technique) and we'll slowly evolve a system to prevent regressions. If the site is a success, we'll have evolved a world class deploy system. If the site never makes it that big then we won't have wasted time on infrastructure. It's classic lean startup thinking (even though TIGdb is really just a hobby project).
I've never worked in a team big enough that it could devote resources to maintaining all of the following kinds of tests: * unit * functional * AND acceptance * plus writing the actual code
IMHO, a neutral third-party group like QA should be responsible for writing & maintaining acceptance tests.
Looks like meeting that goal would constrain you to write code to be used by a robot and not by a human. There may be many cases where this is both doable and acceptable to the end user. So no problem with that.
I am greatly challenged to see how this could be done for a highly interactive, visually oriented, subtle pattern generating response to user input, type application. Computers are still not as bright as earth worms when it comes to generalized pattern recognition. Which means we programmers are about as bright as earth worms when it comes to writing such code.
How then could computers automatically test all the software reactions to the wonderful and totally unpredictable behavior of mere humans as they interact with your software? The test cases would expand to consume all the resources available for development. All you would get done is writing all but impossible test cases. At least you wouldn't ship bugs.
This does not consider the explosion of combination and permutations of inputs that prohibits exhaustive testing that no matter how many systems you run tests on.
It would be much easier and cheaper to go out of business. Your certainty of being free of shipped bugs would be much better than one in a million.
Why not have your local tests, automated or not, cover the common cases and error conditions to catch programmer stupidities? Then let the actual humans do the strange corner cases.
If your design is even close to correct, testing repeatedly tested code is pointless. If your design is corrupt and your implementation is sloppy, no amount of testing is going to save your ass.
I do very rapid turns and I am a one man team. I can turn my system in less that 30 minutes and have the user testing it in a live situation on the other coast. If I want 10 turns a day, I can easily do it. Low coupling, high cohesion, clean correct design, and disciplined implementation makes it possible.
I agree that doing things in small chunks is a great way to do it but doing the equivalent of a weeks worth of global automated testing for each small change seems like a silly exercise. That is except for the server hardware salesmen and system admin people.
The sales commissions and payroll look rather good. The production of real value is questionable. Bang for the buck is as important for testing as it is in any other part of product development.
I have found from working in large teams, there is a core four who get things done. The rest are simply dead weight dedicated to shuffling paper and attending meetings. At best, they do nothing. At worst they create more work than they do.
Use the right four and dump the other sixteen. You will get at least ten times more productivity and ten times higher quality without even breaking a sweat. If you don't have the right four, you are hosed from the start.
When you begin to take that into account you realize you have to find ways for the larger team to work together and still produce a quality product. Hence the techniques being used by the author and other companies out there trying to address similar problems.
I am not sure its possible. The communication overhead of so many linkages forces incoherence. The resultant incoherence forces still more additions to process and body count. That adds still more communication overhead. The result is still more incoherence - not less. If something is "finished", its simply because time, money, resources, and toleration ran out. The end result was simply called "done".
Maybe that is the best we can do but I am hard pressed to call products produced that way quality products. See Vista et.al. for instructive detail.
A well written concise introduction to continuous integration / constant testing would be a boon to this community.
There is a certain non-zero probability for errors to occur during deployment. Binaries have to be reloaded, database connections have to be reconnected, sessions have to be restored, etc, so the more you deploy, the larger the coefficient before this probability in the "will something go wrong" equation.
So, what we do is break up our system into deployment groups where some handful of users gets updated a few times an hour sometimes. We test the deployment on this small set of users, usually they know the change is coming and are ready to test the change in real time.
Sometimes we repeat this process using different deployment groups. Test in this one, then test in that one, until we get a final small errorless deployment and then we roll out to the masses.
If it is successful, we roll it out to the masses.
Your site doesn't have to be /all/ beta or /all/ production. You can have batches of users in different groups.
No, seriously, I'd be much obliged if you could tell what tools go into your setup, how much of it is created in house - and thus unavailable - and how much of it is off the shelf, preferably open source. I'd very much like to spend time on recreating what you've done there.
http://startuplessonslearned.blogspot.com/2009/02/continuous...
http://startuplessonslearned.blogspot.com/2008/11/five-whys....
http://startuplessonslearned.blogspot.com/2008/09/new-versio...
http://startuplessonslearned.blogspot.com/2008/12/continuous...