How Lanyrd moved from AWS to SoftLayer and MySQL to PostgreSQL with no downtime
lanyrd.com
lanyrd.com
The moment we've gone in to read-only mode we can create a brand new instance of the site from a copy of the (now frozen) database, make any necessary changes to that, then switch traffic over to the new instance once we've tested that everything is working properly. If the new version of the site has a problem we can turn read-only mode off and continue to run on our original database.
If you're running a more content-oriented site it's well worth taking the time to set this kind of thing up - it's not too hard to do, and it gives you an enormous amount of flexibility for maintenance further down the line.
It will basically instruct it to continue to serve content that it considers "stale" (up to a point) until it is able to update it.
This is where you are super glad you did the right thing (did you?) and your DB layer is abstracted and your SQL is standards compliant, and this saved you hours of headaches. Anyway, you maintain reads from the old db.
Then, as the site is running, you migrate all data pre-fork, to the new database. Finally, after validating that you got it right, you flip the switch again to have reads come from the new DB.
But you're not done yet. Keep the reads forking for a while, till you're sure everything went okay. If not, you can flip back to pre-migration instantly with zero data loss.
Presumably, all this flipping between dbs is done through some kind of flag that can be modified at runtime, so you don't have to do more code deployments, because you're dancing with the devil here already.
In any case, with proper capacity planning and good code, this is also doable, but not super required unless writes are mission critical for your customers 24/7.
But the last thing you want to do is to flip the switch, pray, lose customer data.
A better approach would architect to write all the updates to a persistent message query. Then a updater can read the update messages from the query and apply them to one or both database. The updates are welled ordered with respect to time and to both databases. The potential failure scope is limited to one place and it's easier to go through the recovery cases.
Persistent queue has other benefits in scalability and fast apparent saving to the users.
Second, I guess it's true that there are some consistency issues since you have to commit one transaction only if the other succeeds. But this is way smaller a risk, I think, than yours.
With your method you're writing way more code just for the update and adding a whole new layer (both failure modes in their own right), and then just moving the transaction inconsistency problem to the updater instead (when you choose to write to both databases). The benefit is that you can "replay" the updates from the updater, but the truth is not all changes are idempotent so replaying some queries may fuck things up without storing DB state (e.g. inserts, increments) which is a mess in itself.
So, I guess, choose your poison :)
I was wondering about disabling writes to databases that are critical to application to support read-only mode with application logic to gracefully handle write failures.
Out of interest, how do you catch and block all POST requests when the site's in read-only mode without duplicating code? Not sure if you use CBVs at Lanyrd - if so, do you use a common mixin? If not, how?
This could work: https://gist.github.com/4066816
I used to agree with you on the price until I realized that with Heroku you're essentially paying for a dedicated server with that much RAM for each plan. Reason being, every single database is "hot" and the entire thing lives in memory to avoid I/O.
On second glance ... you're right. It is kind of pricey.
Speaking of moving to dedicated.. I'd love to hear some of your general feedback on how AWS has been at some point. Essentially where you'd use it again or where you'd go directly to dedicated (or a diet option like a VPS) if you were starting on a new project.
I guess it depends on the type of workload (reads vs writes).
We might be able to schedule downtime for a period when no active events are going on (tricky considering the number of conferences happening around the world) but it's much easier for us to be able to make these kinds of changes without worrying about events that are using us to serve critical information. As it is, we still make sure to communicate planned read-only mode periods in advance so conference organisers have a chance to plan around them.
It would be great to know more about your new setup, i.e. do you use streaming replication and some resource manager like Pacemaker?
Even Cari.net, which I think prices a bit higher than others is offering more for the money. I've had several dozen machines with them since 2006 without issue, top notch support. I also use the really cheap and no-frills folks like Ubservers for when I want a bunch of disposable cheap dedicated machines. Their support is absolutely shit, but by god are they cheap and the bandwidth real. I've been through probably dozen and some change of these guys and that's really all that matters is cheap solid bandwidth and a vague sense of support.
The only criticism I have of Softlayer is their RAM pricing is sometimes extortionate, but is negotiable.
I've named the ones I feel comfortable mentioning, but if you really want to know more about what's out there, there's a number of forums like Web Hosting Talk where the actual companies maintain presences to run promotional deals -- and there's a lot of public opinions aired about how said companies are doing. There's going to be some rather uninformed opinions aired, but most people can tell you when there's a real problem.
[1] http://www.aeracode.org/2012/11/13/one-change-not-enough/
"If it fits on an iPod, it's not big data."
That said, out of any company I've ever been with their support was really good. Just understand you're paying like crazy for it.
Every server you launch is connected to your own private LAN.
Doesn't AWS have physical disks for RDS? What am I missing?
EBS has historically also had somewhat variable performance (on top of the extra network latency), but the new provisioned IOPS feature should help with that[2].
[1]: http://aws.amazon.com/message/65648/
[2]: http://aws.amazon.com/about-aws/whats-new/2012/07/31/announc...