About synchronous disk replication
cloud.google.com
cloud.google.com
Although it is called synchronous replication, it appears to be asynchronous. That is the primary disk will acknowledge writes while the secondary may not have acknowledged them. To make something synchronous usually requires 2 secondaries where writes occur on 2/3 disks total: this allows one to fall behind and the primary acknowledges writes after replicating to 1/2 secondaries.
this GCP offering allows for monitoring when a secondary falls behind. With just one secondary that means when you get your alert, there’s a potential for data loss. If there was a write to 1/2 replicas then when you get the alert, you know you can still fail over to the secondary that is caught up and don’t have to panic while trying to deal with the replica that is falling behind.
> If the disk replication status is catching up or degraded, then one of the zonal replicas is not updated with all the data. Any outage during this time in the zone of the healthy replica results in an unavailability of the disk until the healthy replica zone is restored.
There isn't a binary log that the replicas can catch up to, if the healthy disk goes down you are out of luck.
MySQL's semi-synchronous replication is "more synchronous" than this. If it's enabled it won't acknoledge a transaction until a replica has the transaction saved in it's binary log. Then the replica could be out of sync but if the master exploded, the slave would eventually catch up to the master using its own binary log.
I'm either misunderstanding Google's service here, because the name doesn't seem right.
As block devices aren't ACID, their challenges are greater.
But note the following from the MySQL docs, similar problems exist when you have to fail over with uncommitted transactions.
> With semisynchronous replication, if the source crashes and a failover to a replica is carried out, the failed source should not be reused as the replication source, and should be discarded. It could have transactions that were not acknowledged by any replica, which were therefore not committed before the failover.
For durability, you'll indeed want something like 3-way replication. But that's a distinct problem. If durability is your concern but you're fine with the availability SLOs of a single disaster domain, then you don't need regional replication.
Users don't continuously check replication status. They rely on it being synchronous almost all the time.
3 way quorum replication is great, but you then need to send to more data centers, potentially affecting performance. There's a tradeoff.
(I work on GCP storage)
The availability story isn’t incredible either. Once the secondary is behind an outage of the primary means the system is no longer operational until that primary can be restored.
Three-way replication with raft and 2/3 or 3/5 acks is what modern distributed databases use. It’s for both availability and durability.
The same underlying system is also available for blob storage.
It's strange that GCP and AWS are so far behind on this feature.
(I work on GCP storage)
I did notice in an article about using blob witness for SQL clusters that they’re not all interchangeable.
Then users can decide the cost vs reliability vs durability of data written milliseconds before an outage.
Perhaps give users a web-based calculator where you can put the numbers, and see how much it would cost in $ per gb per day, the mean time to committed data loss based on historic data in years/centuries, and the typical increase in write latency (loss of performance) compared to a single replica.
Then the user can decide.
CREATE PERSISTENT DISK
----------------------
Replicas: W
Zones to split replication across: X
How many replicas must complete a write to allow the VM to continue: Y
How many zones must complete a write to allow the VM to continue: Z
With the above settings, you can expect approximately:
Write latency 5-95%: 0.5-2.5 milliseconds
Mean time to committed data loss: 37 years
Mean time to failure to write: 3 years
Mean time to data loss of data over 1 minute old: 18327 years.
Cost: $0.18/GB/month
The user would set those 4 parameters, with help text for guidance and a big red warning if you set the parameters to something insanely slow/unreliable.Probably because the mental model for the underlying implementation that you are basing this idea on does not accurately reflect the actual reality of the implementation.
As a very broad rule of thumb, whenever you find yourself saying, "Given <thing A>, I don't understand <thing B>," you probably want to re-examine Thing A.
As a follow-up of the availability topic: what's the best to simulate a zone outage on GCP?
Typicality, for this kind of disks (or previously "regional persistent disks"), you want to test that whatever automation you have for failing over to the secondary zone works correctly. It would be great also to be able to test other kind of failures at the zone level.
It feels better to me to have the replication at the data store level (e.g. database) rather than try and hide it under APIs that really aren't meant for it