How Balanced Does Database Migrations With Zero Downtime
blog.balancedpayments.com
blog.balancedpayments.com
I like the simplicity of your approach, it's literally 2 lines of code. Bravo guys.
[1] https://www.braintreepayments.com/blog/how-we-built-the-soft...
(Disclosure - I work at Braintree)
(Disclosure - I also work at Braintree)
"How we handle something <TECHNICAL> at balanced"
or
"How we handled <SITUATION> using <TECHNOLOGY> at balanced"
is of interest to this community.
No need to ask, get writing, hop to it! :)
(We're planning on doing something similar in the future.)
Basically, before any deploy (not just drastic ones like this), not only do all existing unit tests for the new code have to pass (with a certain minimum code coverage threshold), but also a full acceptance suite, which tests the new code and how it connects to all of our other bits of infrastructure. We simulate an extensive set of operations that a client might perform against a test instance of the server, and also run a number of tests with all of our services loaded in-memory, which allows us to mock/patch arbitrary points in the code, to assert that what we expect to happen is actually happening. We also run each of our various clients' test suites against the new code, to make sure that each client sees the behavior it expects to see. This testing suite has dramatically increased our confidence any time we have to do a deploy, and best of all, it's all done automatically.
Since we use services internally, being able to confidently test interactions between all our services (6+ at this point), it is a HUGE win for us.
Open up an issue here: https://github.com/balanced/balanced.github.com if you're interested.
Judging by what just got posted, please assume you can post anything about your tech stack and we will love it.
I also wrote a more detailed post about it for SysAdvent 2012 - http://sysadvent.blogspot.com/2012/12/day-3-zero-downtime-my...
Using Nginx also lets us do fun stuff with routing using Nginx's Lua integration, which we may end up writing about in the future as well.
That sounds a lot like transactional schema updates in Firebird to me. Plus being careful about how the app handles the data. Schema updates in Firebird are essentially instantaneous, with the row updates performed lazily (although if you need it, a SELECT COUNT(*) will force an update of all rows immediately).
You obviously thought about this a lot but I don't see why you'd want two humans doing it instead of a script.
The whole point of this was so API requests wouldn't fail, they would just take slightly longer. API requests don't get the option to see a maintenance page -- they would just return errors to the client, which potentially means lost business for them.
If I'm reading this correctly, your suggestion would alter a single table online, and at the end, I would end up with a table with a new schema (assuming I had no foreign keys referencing the table being modified, which seems to introduce additional complications). Presumably, this change happens while my application was running, which means that during the migration, I would have to use the old table format, and then cut over to the new one instantly once the migration has completed.
Our migration at the time involved multiple table changes, many of which had foreign keys referencing each other. It doesn't sound like this tool would atomically switch all tables to the new schemas, which would have led to broken data for us. Does that make sense?
EDIT: grammar
While you can't use pt-osc to do multi-table updates directly, you can use the same strategy. All it does is creates a new table with the new schema, adds triggers to the old table to duplicate row modifications to the new table, and then copies old rows across. Then, when all the copies are done, atomically renames the new table into place, then deletes the old table.
There is nothing to stop you from delaying the rename until all new tables are ready, except that it is more hassle than just using pt-osc as it comes.
But, point taken: your case is more complicated than the one I was thinking of. And thanks for your thoughtful response to my somewhat dismissive comment :)
As it stands, during your upgrade you've lost all of your fault tolerance and can't meet your performance requirements - you've gone from 5 nodes able to process the traffic to 1!
This would provide a fail back mechanism assuming you could resolve data continuity.
I'm sure there is a reason this wasn't viable (possibly the data issue), but I was curious.
Being a payments-processing company, we have a variety of users (marketplaces), each of whom has users who are widely geographically distributed, so our usage doesn't drop off as much as some other types of sites might on weekends. That being said, on weekends we see slightly lower traffic than during the weekday, so we performed this migration on a Saturday evening.
We ran our migrations multiple times on test instances of our database, because we were worried about this exact issue. We optimized the migration to remove extraneous changes a few times in order to cut down the time taken. Also, 13 seconds was actually the upper bound of what we saw -- many times we ran it, it took closer to 9-10 seconds.