Why I'll be letting Nagios live on
laur.ie
laur.ie
In response (obviously) to http://www.slideshare.net/superdupersheep/stop-using-nagios-... (also available through HN).
Author uses Nagios at Etsy, with 10,000 checks (mostly in the 2 minute range), and it seems to work well for them with some minor tweaking. They have a plugin that provides a REST API. And he thinks the rest of the complaints are about backend stuff that he doesn't deal with much or at all (configs, wire formats, et al.)
He considers Nagios to be simpler than the proposed Sensu, and prefers "Unix style" applications rather than monolithic ones, where Nagios is certainly on the simpler side.
So he will continue using Nagios because it works well for his use, and good luck making something better.
Ops guy here. If something better than Nagios/Zabbix came along, ops people would use it. Nagios gets used because 1) people know it and 2) its pretty easy to setup.
Is it going to scale to Netflix or AWS scale organizations? Probably not. But at that scale, I presume you have buckets of money to throw at problems (i.e. build a custom monitoring platform).
A simple 'strings' on the proprietary app's agents revealed a whole lot of 'zbx'.
Link: https://www.zabbix.com/forum/showthread.php?t=10155&page=5
https://pbs.twimg.com/media/BNELF1GCUAExynU.png:large
for reference the architectural diagram of sensu
My colo server died earlier today (completely unrelated to this) and didn't come back up because GRUB hadn't reinstalled properly on a replaced software RAID disk.
But yes, Nagios did tell me it was down :)
This strikes me as a pretty lazy argument. Admittedly, that diagram is not the best. But that you can separate the components of Sensu isn't a bug, it's a feature. No one says you have to do it this way. In fact, you can exactly replicate the architecture of Nagios by having a single server. The point is you have choices, choices which are dependent not on the monitoring software (which is very lightweight by design in the case of Sensu), but on the other open source software it relies on (Redis, RMQ).
So I would argue that the "server instance" scope is way too broad a category to measure complexity. If you attempted to diagram the workflow Nagios uses (ignoring for a moment what server instance each component is on), you would come up with something equally bad (if not worse). That's if you even understood anything at all about how Nagios works to know what to diagram.
So let's replace one crude measure with another. The Sensu core repo is ~3MB total. Nagios core is about 10x that (30MB). NRPE is about 1MB all by itself. Mod_gearman (to pick out an add-on) comes in at a whopping 6MB. Suffice it to say, but for something that's basically a glorified exit code validator, this seems like a lot of complexity. Sure, Nagios has a lot of features that Sensu doesn't have, and that accounts for some of this. But there's a lot to be said for modular systems vs. monolithic ones.
Not sure why people are making Nagios configuration more complicated than it needs to be.