o Logs are there for ignoring. You need them precisely twice: once when you are developing, and once when the thing's gone to shit, but they are never verbose enough when you need them.
o Using logs to derive metrics is an expensive fools errand pushed by splunk and the cloud equivalents(ie cloudwatch and the like). Its slow, inaccurate and horrendously expensive.
o using logs for monitoring is a fools errand. Its always too slow, and really really fucking brittle.
o metrics are king.
o pull model metrics is an antipattern
o Graphite + grafana is still actually quite good, although time resolution isn't there.
o You need to raid your metrics stores
o We had a bunch of metrics servers in a raid 1, which were then in a raid 0 for performance, all behind loadbalancers and DNS Cnames with a really low TTL.
o Cloudwatch metrics are utterly shite
o Cloudwatch is actually entirely shit.
o tracing is great, and brilliant for performance monitoring.
o Xray from AWS is good, but only really for lambdas.
o tracing is fragile and doesn't really plug and play end to end, unless you have the engineering discipline to enforce "the one true" tracing system everywhere
but what do you monitor?
http://widgetsandshit.com/teddziuba/2011/03/monitoring-theor... this still is canonical.
In short, everything should have a minimum set of graphs, CPU, Memory, connections, upstream service response times hits per second and query time, at a minimum.
You can then aggregate those metrics into a "service health" gauge, where you set a minimum level of service (ie no response time greater than 600ms, and no 5xx/4xx errors or similar) red == the service isn't performing within spec, yellow == its close to being outside spec, green == its inside spec.
if you are running a monolith, then each subsection needs to have a "gauge". for microservice people, every microservice. You can aggregate all those gauges into "business services" to make a dashboard that even CEOs can understand.