EBS volumes are great and all but not for database where the dataset is many multiples of the working set.
EBS volumes are great and all but not for database where the dataset is many multiples of the working set.
It’s totally safe to use local storage if you build it right. But those raided EBSs caused a lot of problems. In short, when one gets slow the whole volume gets slow because software raid isn’t hardware raid.
The main advantage of RDS is that they take care of the mundane redundancy for you.
As a database operator I treat safety on i3 similarly where I have multiple hot replicas of my data so that if any fails I’m good to go. Additionally, there isn’t any reason you couldn’t have a EBS replica of an ephemeral node.
What we typically do with i3 is mirror the data locally, replicate it, have an EBS replica, and take backups. This is probably overkill but the data needs to be both accessed quickly and secure so that’s where we are at.
Is it automatic or manual?
On infrastructure I handled from top to bottom, I used VIPs with keepalived (only the vrrp part, with a weight linked to success/failure of a check script).
But in AWS, I'm wondering how to do it properly, maybe DNS records with low TTL (like 1 second).
As you mentioned, it is limited by instance size, but for a DB that fits it works great and has fewer moving parts. Knowing that your entire database is essentially ephemeral raises the stakes too and forces you to take replication, backups and restore testing seriously.
I don't think you should trust your data to a single disk, whether or not it's a physical device in your own datacenter or an EBS in AWS. Everything fails eventually.