> I never saw anything like this done with metrics or logs done by anybody.
Well, I assure you it can be done and it is not nearly as scary as you presume. I migrated an old Nagios+Cacti system to Prometheus (well, VictoriaMetrics) for a DC with thousands of servers.
Nagios using 4-valued datatypes for virtually EVERYTHING is frankly just poor design. As soon as you move away from binary state, the semantics becomes ambiguous.
Consider these two monitoring systems:
- Nagios-style: 1+ state metric for your HTTP service. "Warning" indicates something wrong but not critical. What exactly? Usually you will see the warning and then look at metrics in sysstat or Cacti as a follow up.
- Prometheus-style: 1+ state metric for HTTP probe, and a few relevant performance metrics. RPS, latency, etc.
The merit of the latter is you use the exact same alerting + visualization infrastructure for both types of data. You create a CRITICAL alert if the service is down, and WARNING alerts if RPS/latency crosses thresholds.
Importantly, this data is not siloed into Nagios and can be used for other purposes such as usage analytics and cross-domain incident analysis. The alerting service is just one possible consumer.