Deployinator: Etsy's deploy tool open sourced
slideshare.net
slideshare.net
http://codeascraft.etsy.com/2010/05/20/quantum-of-deployment
Edit— Found this from NYLUG in April: http://www.archive.org/details/NYLUG-2011-04-20-Etsy-Deployi... (Talk starts at around 4:20, runs to 1:21:06.)
Haven't watched it through yet, but it looks like mostly the same slides.
The only solution I can see is one where databases are handled differently (they generally are anyway at the deployment stage), and IMO, this is the challenging issue. Continuously deploying app servers is not trivial, but I would consider it a mostly solved problem.
Another option is to use a VM for your database and do a snapshot before deploying.
Also, for a DB, LVM's probably just as good as a full VM.
The context of the discussion was performance, scalability, and engineering resources, so we didn't get into how the logs are used, but it would be my guess that if the "primary" database had some sort of failure that lead to a rollback, affected accounts would be locked out until the transactions can be properly recovered from other storage/logging mechanisms.
A simple approach is not to delete data, but rather simply set a flag which marks the data as deleted. The application storage layer acts as if this data were deleted. A separate process (changed separately) is responsible for actually purging data which has remained deleted for a certain period of time.
In any case, do you really wish to delete the data? Keep it around -- it might be useful. The only reason to delete data, outside legal concerns, is to control the cost to retain it. Given that you mostly never really want to delete data, you probably want to orient your schema toward hiding it instead. In that case, if you need to roll back application changes, you can do so by backfilling the data to restore it to a visibile status.
Keeping detailed programmatic logs is also helpful for making the backfill easy to execute. For example, in addition to your main textual log file, you keep something like a log consisting of JSON objects, one per line, each which represent the essential details of some application action.
For schema changes, there are a bunch of approaches. Just a few:
* Treat (non-purely additive) schema changes as a special case and take extra care deploying them. (This sounds like a cop-out, but it might be the right effort tradeoff for many startups.)
* Write forwards and reverse migrations for each change, and avoid data-losing schema changes.
* Implement all (or all non-purely additive) schema changes by shadowing data operations onto a new version of the schema (asynchronously migrate all previous data onto the new schema in some back-end process.) Write tests to verify the migration/shadowed data is correct. Switch data reads to be on the new schema. Eventually delete the old schema in a subsequent deploy after an appropriate stabilization/verification period.
I find #3 to be the most general, but there's plenty of overhead (computational, storage and man-hour) that may not be appropriate for your case.
1. Use some sort of schema migration tool to keep track of and version all changes you've made to the database since your very first commit.
2. Never do a backwards-incompatible schema change. Never. Time and time again these are the ones that will completely bite you in the ass when the day comes that you broke production with a deployment and desperately need the old version back up and running. You're going to end up losing data and shit will hit the fan. It doesn't take much code-wise to support the old scheme, and it will ensure that you have zero downtime while rolling out your updates.
3. Never delete data. The only exception being if some user-submitted content violates a law, like child pornography. In which case, it shouldn't be part of your migration/update process to begin with. Just use deleted flags / boolean fields. Note in your code when certain fields are no longer used anymore because they're only there to support legacy versions.
Concerning 2: how do handle schema fixes (example: at my previous job, the tables were very badly designed, and the application+schema combination prone to frequent race conditions). What to do in this case ?
As for 3, I already came to this conclusion on my own, so I was not completely on the wrong track :)
We avoid data loss by doing soft deletes as much as we can. Sometimes we do want to do real deletes, especially for data that is heavily denormalized, but in those cases we keep an audit trail so that the data can be reconstituted in the event of a mistake.
Databases are handled differently from this tool. Iirc he said Thursdays are database day and all changes are done/initiated then.
Even in svn, and much better in git, if you can keep a sane release and merging strategy then it's not going to introduce much overhead to getting new features or bugfixes out. Our 20+ person team deploys a few times a day using a branch system of (sprint + trunk/QA + release), but always looking for different ways to do things.
Server affinity + staged rollouts are another approach, but for long timeframes (e.g. A/B UX tests, feature prototypes) I'd think you'd want to go with flagging anyway from a codebase management point of view.
Some things do take more effort to make dynamically flaggable than it would take cost to delay integration, to be fair. Judgement call.
When a feature is in development, its config flag is off or enabled for admins. We deploy it as we develop it, in pieces that are generally not more than a few dozen lines of code long. This is another critical idea: we don't push out all of the code for an entire sizable feature all at once, ever. Doing that is a recipe for having a problem and having to review thousands of lines of code to try to figure out what it is.
For new features, we generally do not remove the config flag once a feature is live because we want to be able to disable it if anything goes wrong.
If we are replacing a feature with another one, we generally want to keep the old one around for a little while (we generally do A/B testing or ramp up new versions whenever we do this). After that, we do delete the old code and reduce the config flag to an on/off switch which was probably there in the first place.
[1] http://www.amazon.com/gp/product/0321601912
[2] http://timothyfitz.wordpress.com/2009/02/10/continuous-deplo...