I'd love your feedback on how this process could be easier for me, some resources on learning the Grafana query languages, and general comments.
Thanks for taking the time to read + engage!
I'd love your feedback on how this process could be easier for me, some resources on learning the Grafana query languages, and general comments.
Thanks for taking the time to read + engage!
* ZFS pool errors. Motivator: one of my HDDs failed and it took me a few days to notice. The pool (raidz1) kept chugging along of course.
* HDD and SSD SMART errors
* High HDD and SSD temperatures
* ZFS pool utilization
* High CPU temperature. Motivator: one of my case fans failed and it took a while for me to notice.
* High GPU temperatures. Motivator: I have two GPUs in my tower, one of which I don't really monitor (used for transcoding).
* High (sustained) CPU usage. I track this at the server level, rather than for individual VMs.
I wanted the ability to quickly see the current & historical state of these and other metrics, not just configure alerts.
I’m also omitting the fact that I have collectors running inside different VMs on the same host. For example, I have Telegraf running on Windows to collect GPU stats.
You can run numbers manually but I think designing for it up front is really important to keep performance targets on lock. That's where Prometheus and Grafana come in. And I think looking at performance numbers is a really good way to help understand systems dynamics and helps you ask why something is hitting some threshold. On the other hand, there are so many tools and they're often fun to play with, it's easy to get carried away. There's also a pretty reasonable amount of complexity involved in setting it up, so it's also easy to just say fuck it a lot of times and respond to issues on demand instead.
[1] http://k6.io/, it's also a Grafana project.
[2] It can test both normal REST endpoints but also browsers thanks to the use of headless chrome/chromium! So you can actually look at first paint latency and things like that too.