MySQL Cluster Auto-Recovery
magicofsecurity.com
magicofsecurity.com
[1]: https://dev.mysql.com/doc/refman/8.0/en/mysql-cluster-overvi...
So we know there is very likely an infrastructure problem. The question is, can we mitigate it on the application layer?"
That seems unwise. It's on DigitalOcean, so they could have migrated to different VPS instances and checked to see if the problem followed.
Might reduce bandwidth or increase latency, but a total outage is pretty painful when you're taking about distributed/replicated databases.
I guess my point here was that Our database should be resilient to this kind of infra issues and ideally self-heal if these are transient events.
For what is described in this article specifically, TiDB would make you change the storage engine (from Innodb to Rocksdb), and if they have all their data in a single Group Replication cluster they probably don't need sharding either.