Observability is becoming mission critical, but who watches the watchmen?
simme.dev
simme.dev
We're running our OpenTel stack on a Hashicorp Nomad cluster, and so far the best solution to this problem that we've settled on is to have a totally independent stack in a different DC, with both doing full monitoring of each other -- with the secondary only monitoring the primary OpenTel stack, and probably running on a single VM to simplify things.
This lets us leverage some of the infrastructure wrappers, but minimise common dependencies, and gives us some convenient operational views of 'the other side' during patch cycles and the like.
There will always be edge cases that can drop all our components off the network before anything can bleat out an SOS, so it's more a matter of reducing those scenarios rather than trying to cover the last 1%.
In practice, anything that takes out two of our DC's / cloud presences is not going to go unnoticed anyway, and 'there'll be bigger problems then' etc. LDAP, DNS, and so on definitely fit into the category of services we just have to assume will always be available. In a full on DR situation, our big questions during recovery will be what bits of our stack are functional and which are struggling, and the design tries to preempt that.
A "dumb" heartbeat message, and a check for its absence, is reliable precisely because of its simplicity.
So if the question is "How do I monitor my complicated observability stack?" then the answer, to me, starts with "What is the simplest method of accomplishing the minimal version of my goals?"