Show HN: Self-hosted, open-source infrastructure monitoring and alerting
cabotapp.com
cabotapp.com
The DSL isn't that confusing, and there are many many options there. I've been able to set up a very comprehensive monitoring service using it, and while I too once cursed its complexity, it was worth it in the long run. Certainly didn't take as long as it would take to write a clone.
Also it was surprisingly easy and fun to write. Probably has taken longer than learning Nagios but I don't think I would have spent Christmas hacking away at that task...
Now, the problem is that to get to the links I cited, I had to click through 5 links, and I knew exactly what I was searching for. Nagios's biggest problem is that their documentation looks archaic. Updating that would give the project so much more appeal.
[1] http://nagios.sourceforge.net/docs/nagioscore/4/en/hostcheck...
[2] http://nagios.sourceforge.net/docs/nagioscore/4/en/servicech...
Monitoring is looking at the services and making sure they're running within tolerances. Graphing is not useful for monitoring, its useful for postmortems and planning. Inventory management is not useful for monitoring, it's useful for building out servers. Service discovery has nothing to do with monitoring, though it's certainly be useful for automatically populating your monitoring configuration.
However, trend detection, I agree; nagios needs more of that. I should definitely get alerts when my disk space usage suddenly starts growing 4x its normal amount, or when my CPU usage spikes to 3x its normal usage.
I'm sorry, I have to completely disagree with this point. Graphing is a fantastic way of visualising information to show trends, such as disk space growth, or memory usage, which can and have, in our company's case, led to diagnosis of problems that hadn't yet manifested in a complete crash. I realise you add in a comment saying that Nagios should be doing these things internally, which it also should, but what about situations you haven't planned for? Or intermittent events that don't look like anything individually, but when graphed over 6 months turn out to reveal the subtle failure of an aircon unit in a server room?
My intention was to say, with regards to graphing:
Graphing should be part of your overall tooling for monitoring boxes, but it should not be part of process monitoring. It should be a separate component (which should in turn be monitored by Nagios), for use in post-mortems and future planning.
It should not be built into Nagios, since the graph itself is completely irrelevent to whether Nagios should alert about a problem.
I also think that graphing is different from heueristic rate monitoring; graphs are for humans to spot trends/correlations. Alerting on rate changes is a binary alert/don't alert.
I.E. Cacti/Graphite should be creating graphs about services for human consumption, Nagios should be watching for problems with services.
I also tend to think service discovery or at least a flexible API that allows other systems to create new checks is a huge part of monitoring. Nagios' configuration files are a nightmare to deal with especially in a configuration managed environment.
Ideally, this also resolves your configuration management problem as well. If it does not, there are tools which can probe a running Nagios state file, and provide you feedback on what's being monitored, and its state. The python library nagparser is one that I've used in the past to probe Nagios status for a status aggregation tool.
Also nagios 4 has changed the way forking checks work - http://labs.nagios.com/2013/09/20/nagios-core-4-now-availabl... - I appreciate this was too late for you.
Anything that has to ssh, ftp and http to a few thousand servers is going to be slow.
* All the (virtual machine) host boxes.
* All the routers, switchers, and firewalls.
* Status-checks, on hosted sites, etc.
In the end I designed and implemented a system which was capable of running all the tests in around 90 seconds, by virtue of being distributed. One host does all the parsing and such like, and N-other hosts could pull out tests to execute. (As it happened we ran everything on a single box, but it was designed to be distributed, it just transpired that having 6-10 worker process pulling jobs from the queue to execute was good enough.)
Introduction:
http://blog.bytemark.co.uk/2012/12/19/custodian-a-network-mo...
Code:
It's both a web-configurator and a replacement for the default interface - no need to learn much about Nagios (unless you want to) and should be satisfactory for a simple use-case. Certainly not a replacement for the more complicated and scalable configuration/frontend solutions out there, but it does serve its purpose.
I have services running remotely, that are accessible through a variety of communication methods (VPNs, SSH tunnels, 3G modems, and occasionally just straight up IP!) and it's been awhile, but when I set up nagios, it seemed there was no way to tell it: "Alert if this host is not reachable in any of these 50 ways" - it monitors all 50, and alerts me whenever each one of them changes state (which is quite often).
What I want to say is: Here are all the ways I can reach server "pinky", I want a warning if there are less than 3 ways for more than an hour, and an alert if non of the way works for more than 3 minutes.
Can cabot do that?
(for that matter, can any other monitor do that?)
https://www.zabbix.com/documentation/1.8/manual/config/host_...
Though I'm not sure about the "warning if there are less than 3 ways for more than an hour" requirement.
In our case, it was more that there were multiple things we could monitor for that potentially indicated a problem with a service but it was hard, using Nagios or whatever, to tie those individual indicators back to a single service going haywire at 3am, especially if all 50 of them blow up at once.
We don't have the precise problem that you describe though - for us the ability to monitor the number of data series in a graphite collection (so that if a server disappears we can notify, and then if another drops off we can alert) is sufficient. Cabot can use this and the ability to tolerate a number of failures to give behaviour very close to what you talk about. However I don't know if Cabot would support the kinds of checks that you're carrying out out-of-the-box, you'd probably have to extend.
I'd suggest you spin up a clean Ubuntu instance in Virtualbox and just run the fab deploy script against that locally... It won't be fast, it won't be pretty, but I'm pretty sure it will work.