When is the right time to introduce high availability for web site?
stackoverflow.com
stackoverflow.com
To the first question - daily backups are too rare, consider incremental SQL log backups throughout the day, like every 15 minutes.
To the second question - 5 minute downtime is likely a lot less important for your business than you think. But your hardware will fail at some point and cause several days of downtime it takes to order a new machine. It's probably not acceptable so you will need a standby machine.
Personally I run all my SQL Servers on AWS EC2, so if one of my servers fails I can simple reattach the EC2 EBS volume to another machine and I am back in business a few minutes later. I also do EBS disk snapshots every 15 minutes. Even if entire availability zone fails, I can restore the EBS disk from a snapshot and get up and running in about an hour.
The setup you are describing will work as well it just seems very capital intensive with all the hardware and software licenses. And you're still prone to Geo failure (power failure, flood, tornado, fire, theft).
I split this in two parts - "provisioning", that is installing generic stuff (IIS, SQL Server, python for scripts etc), and "setup", that is installing my own databases, scripts, and ASP.NET apps.
So, I have created an instance, did "provisioning" on it, and then created a "pre-canned provisioned image". Now I can create as many EC2 instances from it as I need, and only thing that is left is deploying my own stuff - that takes one-click and however many minutes it takes to copy the data. Actually, the data does not always have to be copied - if you are tearing down old machine you simply reattach the EBS volume and it takes zero seconds. If you are creating a copy of older production machine, you can create an EBS volume from a production snapshot, but that volume is available immediately, populated from the snapshot in the background, prioritizing parts of the volume you're trying to access. Hence the volume is immediately online, but is very slow at first.
Re cost, please have a look at this blog post of mine, in particular the TCO columns of the spreadsheet:
http://blog.altudov.com/2010/11/03/amazon-ec2-reserved-insta...
IMHO, a standby server is worth it.
2) Another alternative could be RAID 1 or RAID 10. Catastrophic failure of two drives at once is extremely unlikely. Still there would be occasional maintenance downtime, but it looks like maintaining RAID solution would be easier than Stand By server.
Moreover, humans fail. RAID cannot save you from human error that corrupts the filesystem, for example. RAID is just a little too efficient at replicating changes from one drive to the other. There's a tradeoff between how up-to-date your mirror is and how much time you have to detect a problem on the primary and stop that problem from replicating to the slave; RAID is firmly parked on one edge of the space of solutions.
And, as you have noticed, RAID also cannot help you deploy your new code on one machine at a time, secure in the knowledge that you can fail over to the other machine in case of big trouble. (That doesn't work for every deployment, but it sure is a handy option when you need it.)
But as another comment suggests, unless you are grumpy about Azure ToS or US laws (maybe Azure services exist outside the US) or whatever, you should take a look at Azure services.
In terms of pure revenue break-even, he's looking at 6 days downtime to make back the money. But he's talking about 5 minute scheduled downtimes to apply patches.
My company runs on .Net and SQL Azure with multiple web and worker roles. Its worked great for us so far.
How would performance be affected? For example, I use SQL Server Full Text search -- how would it work with Windows Azure?
How much would it cost? At this moment I have ~50 GB worth of data.
My configuration worked great for me so far too. We are talking about potential risk of unexpected downtime. Did you go through unexpected downtime in your Windows Azure account?
Install your web app on a Windows 2008 VM on serverA ; move the VM to ServerB, reboot serverA, then move the VM back. For applying patches to the VM image itself, I figure it should reboot in what, about 60 seconds or less?
You can get fancier, have separate dev, staging, production VMs etc., adding web balancer and what not later. Then when you want, you can go all-VM should you ever tire of the hardware maintenance.