Wait a second. You run your production database on ephemeral storage? Wow.
I see the replication setup and the S3 WAL archiving and whatnot but still... that's brave.
Wait a second. You run your production database on ephemeral storage? Wow.
I see the replication setup and the S3 WAL archiving and whatnot but still... that's brave.
May not be as durable as EBS, but it's enough for me to sleep soundly at night. And with a highly concurrent WAL-G download, it takes like an hour to catch up a new replica from scratch.
Netflix went full ephemeral storage for their Cassandra clusters since the beginning, at the time when they were just spinning disks. Years later, they still insist on doing this, and had to come up with creative solution to fix the uptime issue: https://netflixtechblog.medium.com/datastore-flash-upgrades-...
Anyway, from the other comment here, I think tommyzli might not have realized that a reboot is still possible, which would partially explain the 3 years uptime.
> it takes like an hour to catch up a new replica from scratch
That means it should be pretty easy to replace an instance - just create a new replica from scratch, then fail over (if you're replacing the current master/primary instance), and remove the old one.
I wonder how big the delta was for CMB between EBS and ephemeral?