Following this course will give you a solid foundation for debugging issues, and as has been mentioned, the logs are surprisingly useful for that task. Once we got the packages installed on the cluster, we tested various ways we could imagine it failing (application failing, docker container failing, image restarting, image dying, etc.) and found it pretty intuitive to fix most issues. Sometimes, the fix is necessarily outside of dcos - you'll need to set up an autoscaling group to ensure you always have the proper number of nodes running. You'll need to set up VPCs to ensure your public/private slaves are actually public and private. You'll need to tag those instances as they're launching so Marathon will know they're public/private respectively.
Assuming your team is pretty ops-savvy, I'd say dcos is surprisingly simple to manage and debug, and this advanced course does a great job of walking you through the entire stack of technologies used.
Now that we are post-v1 I'm certainly hoping to replace kube-up with something more readable and maintainable (or at least start replacing it). Hopefully we'll find something which better promotes reuse between the different clouds, and hopefully we can document it a bit better as well. The split between kube-up and Salt is also not exactly elegant.
If there are any particular questions you have about how the magic happens, I'd be happy to try to answer them (or feel free to file issues and tag me on github); this will help me make the docs better.
The logs are _very_ good, and most unintuitive behaviour will often trigger a decent explanation in the logs.
That said, kube is a lot more integrated; you prob will spend more time setting up stuff like Mesos DNS or Bamboo (that's my plug).
So if you invest in Kubernetes and later want to take advantage of Mesos you can lift up k8s and put Mesos underneath.
Mesos is definitely more mature, and you can use Marathon for container orchestration, which is a little more mature than Kubernetes, but it's also a bit simpler. Marathon doesn't have the service, secrets, or pod abstractions, for example.
DCOS adds some additional benefits on top of Mesos, like the dcos-ui and dcos-cli. One of the more compelling features is one-step cluster package installation (ex: dcos package install kubernetes). Currently, the community edition is available on AWS (https://mesosphere.com/amazon/).
I do agree that the kube-up scripts in k8s are a bit of a mess. There's really hard to read and reverse engineer, and each provider has their own divergent deployment methods. That's one of the things DCOS is trying to standardize, to have a consistent deployment pattern for Mesos frameworks that deploys their core components inside Marathon, giving them a more battle-tested platform to live within, and granting automatic resurrection in case of failure.
Disclaimer: I work for the company which donates the majority of engineering effort to Cloud Foundry.