Note: source for this information is from discussions with DO via support and the kubernetes slack. I am not affiliated with DO in any way.
The downtime issues with DOKS are directly related to the resources allocated to the control plane, which is not set up in a HA capacity. The resources assigned are directly related to the size/number of nodes you use. API-heavy applications can very easy knock out the control plane, and in that situation is takes it a (relatively) long time to recover on it's own (I would typically see ~4hr before it became responsive again).
Their support team is able to modify the master resources for a given cluster (to assist with recovery), but the turn-around time on that shouldn't be considered "production ready".
At this point my advice for DOKS would be:
- Are you using very basic, out-of-the-box Kubernetes to host "apps"? You will probably be fine, but be sure to have a back-up.
- Are you planning to use Operators, or anything that heavily interacts with the kube-api? I would recommend not using it, or over-provisioning your cluster (which would very quickly offset the advantage of the "free" managed masters).
I know that they are working on fixing some of these reliability issues and I have hopes that it will be more stable shortly.
At this time I have a "stable" cluster which (unless their support lied to me) had master resources manually increased by their support team after the 3rd incident of it dying. I haven't had an issue since then.