Observability is about more than crashes or slowdowns, serious investment in observability is a must-have for any SaaS/cloud product to have reliability, auto scaling and velocity. It’s more than just crashes/slowdowns.
My team use Grafana’s open source LGTM stack. We use Prometheus metrics to track anything from JVM/Go runtime stats, K8S metrics, saturation of CPU/memory, scalability issues, crashes/OOMs, custom metrics for business insights, debugging. We use USE/RED metrics (see: Google’s SRE handbook) to track our production services performance in an objective way. We track SLAs and SLOs so we know when it’s time to focus on features and business impact, and when it’s time to put that aside to focus on stability and maintenance before our customers notice reduced reliability.
As a developer it’s really helpful for testing changes. For example, I added a new database index in dev, then run some load tests and check our dashboards before and after. I look at Q95 latency of APIs and database load to see if it has the desired effect, then when I roll out to production I can monitor those same dashboards and make sure the same desired improvement can be seen for real-word usage.
I used traces recently to discover that something that should have been happening in parallel was instead happening sequentially leading to very long/timing out requests. Adding visualisations via traces helps get your head around how something is working.
I added annotations to our dashboards that shows when our K8S pods restart alongside the metrics. This made me realise that some requests were failing exactly around deployments because we were not cleanly handling SIGTERM in some services.
We have started adding horizontal auto scaling based on metrics for the number of queued messages on a specific Kafka queue. If a large number of messages are waiting we spin up more K8S replicas, and then once this reduces, we reduce the replicas to keep costs down.
I optimise the resource allocations on our services by looking at historical CPU/memory usage so we make the best use of our K8S cluster and avoid OOMs as we scale.
We use Loki for log querying and parsing, you can create really advanced/domain-specialised log querying dashboards and provide that to your support team, and integrate those logs with traces to debug different stages of a request as it traverses your microservices or different processing stages.
You can even build dashboards from logs, which is helpful when debugging a particular type of error over time that you were not specifically monitoring with metrics, or determine which customer(s) are affected by this error. Alternatively if you have a legacy system that does not have effective metrics, you can build metrics from its logs.
We use our metrics for alerting and paging in a way that provides a better signal-to-noise ratio than old-school alerts like “high memory usage” so people don’t get woken up as much (we’ve had zero pages since my product launched 6 months ago!). It’s better to alert only when we have a measurable impact on customer experience, like when a smoke test has failed more than 80% of the time, or HTTP requests 5xx rate is elevated to abnormal levels.
It’s also really reassuring when you do a prod rollout to easily see that stuff is still working without digging into logs, so you can spend more time coding and less time babying prod.
Overall I think having good observability is definitely a worthwhile investment. There are cheaper ways to do it than datadog. I expect much of the trouble is that switching providers is a huge job, we have invested so much time building our observability stack, the challenge of moving seems massive. Thankfully we picked Grafana’s open source LGTM stack and self-hosted it. Even if you picked their SaaS offering, switching to open-source self-hosted is an option so you are less tied in.