Explaining modern server monitoring stacks for self-hosting
dataswamp.org
dataswamp.org
Other stacks can do the same, but I haven’t found anything as simple to set up; netdata is just a simple service.
If you wanted to be practical about things you very likely would just give up self hosting entirely.
Much of the crowd around here and certain subreddits IS probably practicing for more, but I have a counterexample. You won't find him around the internet. Not interested.
I have a friend who's recently learned just enough Linux and Docker to get through a Debian install with some prompting, and run a Matrix home server, Nextcloud and Jellyfin because he thought they were cool. Its an old (for him) repurposed desktop.
He's a mechanic. He doesn't care about anything enterprise. He doesn't want monitoring because he'll know if it's down. He doesn't care if backups run because he's sure he's got the important stuff on an old disk somewhere, probably.
Alerts work well and it is not too complicated for simple things.
If you're going to be working with it on the job, Prometheus is relatively simple to provision, though, and can provide some experience if you use it at home.
The UI sometimes feels a bit dated and not everything is as straightforward as one might expect, but for my use case (monitoring a bunch of GNU/Linux hosts) it's sufficient.
Some of the things that are good about it:
- can be run in containers if you want to, use a familiar DB like MySQL/MariaDB/PostgreSQL
- the agent installation on the hosts that you want to manage is also pretty simple
- supports both active (monitored host sends data to Zabbix) and passive (Zabbix asks the host for data) configurations
- depending on the template that you use, has lots of built in metrics out of the box
- easy to integrate with something like e-mail or SMS messaging for alerting, other plugins exist AFAIK
- also has built in alerts, such as when disk space is low, CPU load is high, memory usage is high, swap space is low, host is unreachable etc.
- has the ability to build dashboards with graphs, network maps etc.
Some of the less nice points about it: - last I checked (older version, now using Uptime Kuma for this use case) the web monitoring didn't send alerts by default
- last I checked (since haven't bothered to set up again) maintenance windows straight up didn't work and still sent notifications about the agent going down
- in general the UI can be a bit cumbersome and there is some legacy to be found, at least a while back there were legacy graphs and the "modern" ones, each of which had different sets of functionality available
- sometimes certain parts of the OS template also decided not to work, e.g. currently have disk usage not showing up in like 1/6 graphs, even though the configuration is pretty much the same for all of the nodes in question
- in general, it's just not as popular as other solutions and might not have as much tutorials around setting things up in it
I'd probably compare Zabbix against something like Nagios or LibreNMS, rather than netdata or Prometheus/Grafana, but perhaps that's just me.Edit: That said, the docs of Zabbix are pretty decent.
Here's how configuring something in it typically looks like in the UI: https://www.zabbix.com/documentation/current/en/manual/web_m...
Here's an example of some of the dashboard functionality: https://www.zabbix.com/documentation/6.2/en/manual/web_inter...
And here's maps that you can embed in the dashboard: https://www.zabbix.com/documentation/6.2/en/manual/web_inter...
But this requires different templates for each mode.
Why they did this way is beyond me... though I have a suspicion what they aren't using it themselves on anything than a tiny lab with a couple dozens of hosts.
Dealing with 3 things... feels about as much work as zabbix. That one requires server + db + agents. Basically the same components I'm running.
Promql/flux - while I know it, I almost never use it - instead in grafana click the database name, metric name, aggregation and I'm done.
While you need people to use it effectively, it's the same for everything, including mysql for zabbix. You can dump prometheus/influx somewhere with no configuration and survive for quite a long time.
I looked around but didn't really find anything that fit for me. There are a lot of complicated (albeit powerful) options, but I want simple, easy, lightweight, quick. These days I'm juggling so much, I want to be as efficient as possible with my time.
It's still early days but hoping to be able to onboard people towards the end of the year, for anyone who is interested feel free to join the waitlist: https://serverduty.co
Interestingly, I'm using ServerDuty to monitor ServerDuty as I build ServerDuty. I mean, if that isn't dogfooding, I don't know what is.
It uses a HTTP client to poll all instances every minute which return /proc/stat so that I can see how much CPU they use:
This is my main gripe with Prometheus (didn't use it yet).
For backups, you ideally want the backup server to do a pull from the backed-up servers, to avoid that a security incident in one of them could damage your backups.
But for monitoring, I feel that push is the way to go. I want only the minimal, indispensable connections into my production servers.
I like this approach because, for security data, you want it off of the box as soon as possible. In an ideal world a log would go straight from the kernel to another box, or as close to that as possible (to avoid tampering/ DOS'ing). So for your security data you're already going to want to push that latency metric down, and now the question is what to do with the rest of them - obviously your service logs are less sensitive, but at the same time getting things shipped off of a box can save you a lot of headache.
This is hand wavy though, I'm honestly very curious to hear what others think as I haven't built the system I'm describing.
In response to your vector solution:. What I found highly adaptable is having all servers forward to a central vector/fluent-bit agent, have that agent forward to kinesis firehose, then attach a transform lambda to create the final output. You end up with 2 configuration points from which you can control and transform all your logs. Instead of having to push this down to the servers through config management.