Nagios vs Icinga: the story of one of the most heated forks in free software
freesoftwaremagazine.com
freesoftwaremagazine.com
However, Icinga2 and icingaweb2 are both coming along very nicely. In a position I just recently left, I was using Icinga2 in production, and I really loved it for its performance and clustering capabilities. It really blows Nagios out of the water in those aspects. When you configure the active/passive "master" nodes in a datacenter, the checks get split between the two nodes with only one updating the DB (per the docs), and in practice, this seemed to keep the load very low for us, even with over 20,000 checks polling at least once per minute.
I realize the story is more about history than a direct comparison between Icinga 1.x and Nagios, but I just wanted to share our experience with anyone who may be considering the two options - don't forget to consider Icinga2!
It should also be noted that I still really love Nagios, and the later versions have resolved at least some of the issues we've had with it in the past.
It's only free for up to 25 hosts. Then it costs at least $65/mo (for up to 250 hosts) plus the fees for extra features.
One of the biggest wins of Icinga2 for Nagios shops is the fact that it can be sort of retrofitted to work with your existing Nagios hosts using NRPE.
Of course, NRPE is highly insecure and this is not recommended, blah blah, but it works, and I'm sure there are more than just a few people still using and older version of NRPE in production.
Yeah, they could make that more appealing :)
On the other hand, we're getting really sick of every upgrade breaking our environment (if it's not the post-install scripts choking on our use of a non-default DB name, it's the configuration parser choking on previously-valid configuration files, or...), and we wish that the new UI could catch up with the capabilities of the classic UI.
We also wish we could pin a package version, but the Icinga devs remove the previous version from their repo the moment the new package goes up. Our solution to this will be to host the truncated packages on our own repo; we just wish we didn't need to do that.
There are no good APIs (add/remove/modify hosts or services sanely).
The configuration files are not easy to templatize, and hard to understand even if you are able to programatically generate them.
They are even more clunky in a dynamic environment like AWS.
Check_mk is the sanest Nagios improvement add-on I've seen: http://mathias-kettner.com/check_mk.html And that has its own issues.
But now I wonder - is there any technical reason to choose one over the other?
I set up Nagios to monitor about 50 servers last year and now I'm almost done switching to Icinga2. I think it's worth it if you have spare time, but the major improvements I've actually noticed have been UX, clustering, and graphite plugins. other than that it's Nagios with a nice GUI.
Not worth the switch if you have pressing issues at the moment.
http://webcache.googleusercontent.com/search?q=cache:www.fre...
My feeling always says the first, but all my successes (which I define as "someone was getting his job done using my open source tool") came from the latter. So I wonder if my sample size is simply too small, or if I was too naive in the first place, thinking clean architecture is good and important.
This kills the company when their competition outdoes them.
> or someone who will integrate nearly everything even if it means the code base, build process etc are a mess?
This kills the company slightly slower, when technical debt means slower development and less features in the long run, allowing their competition to outdo them.
> My feeling always says the first, but all my successes (which I define as "someone was getting his job done using my open source tool") came from the latter. So I wonder if my sample size is simply too small, or if I was too naive in the first place, thinking clean architecture is good and important.
It's also possible your bar for "clean architecture" is high, or different from what others would consider "clean architecture".
I have written some terrible, terrible code, that hasn't caused major problems. Either because it was replaced with a proper solution before growing too unwieldy, or because it was contained well enough that it never became unwieldy in the first place. You could say that even though the micro-architecture was terrible, the macro-architecture was at least acceptable - which IMO is the vastly more important of the two to keep "proper" for long term maintenance. Although one has to be wary of a system growing until the micro-architecture becomes macro-architecture, if the former is of poor quality.
The fundamental argument that generated the fork was the community permanently diverged between simple and complicated and the primary maintainer decided to ally with the simple team, probably from a strong historical "eat your own dogfood" culture. Maybe rephrased some people have to ping 10K boxes per minute and other people have to check 1000 services or files or aspects on 10 boxes and its not REALLY possible for the same codebase to serve both groups equally well in actual real world practice.
Yo, heres some patches to make your screwdriver(tm) more hammer-ish to help professional nail installers work somewhat more effectively. What are you crazy the wood screw installation professionals such as myself will not tolerate that, patch denied. OK we'll fork and BTW we're taking the name with us so on the other side of the planet, people will be installing nails using our hammer shaped fork of Screwdriver(tm) what could possibly go wrong? Oh no you won't, at least with my trademarked name. (consults with lawyers inserted here) Um, yeah I guess you're right we're gonna rename our hammer shaped screwdriver the Hammer(tm). And years later the worlds carpenters happily use Screwdriver(tm) and Hammer(tm) and sometimes both at the same jobsite without knowing the somewhat contentious history of those different tools.
The two user communities have little if any overlap. This near total lack of overlap lets the two communities live together in peace and plugins generally interoperate etc etc. More like relations between the Word and Excel communities than like Android vs iPhone communities. Aside from that exciting little trademark spat a long time ago relations have always appeared pretty calm and cool.
I admin'd a large nagios system a bit more than a decade ago that basically pinged the heck out of a large number of customers using configs generated by polling our own network devices automatically (read only access, auto generated configs, etc). It depended on the system being simple stable reliable and fast, it was always speed limited in deployment. I needed Oracle support like I needed a hole in head, so new complicated slow unstable features held negative appeal for that implementation. So I stuck with nagios, although I can totally understand a user with a completely different use case needing icinga. I wonder if that system is still running at ex-employer? Probably. It was pretty bulletproof software.
Really? Why? The first thing seems quite easy all around. Configuration-wise one would expect it to be simple. Computation-wise, pinging 10K boxes per minute means sending <200 packets per second (and receiving at most the same number). Maybe 3X that or so to get multiple samples. That's nothing for a modern machine.
The second one obviously requires some more complexity but I don't know why the extensible configuration and plugin system it'd require would have to interfere with simply pinging a bunch of machines. Computation-wise, again sending <200 requests per second should be trivial unless they're spectacularly heavy-weight.
The Nagios plugin protocol does seem to be somewhat heavy-weight, though. Looks like it still fork()+exec()s on every probe? https://nagios-plugins.org/doc/guidelines.html That seems like a much greater problem in terms of computation than the actual probes themselves, particularly if probes are written in a scripting language with non-trivial startup overhead. Doesn't look like Icinga is any different? What significant architectural differences exist between the two?
[0]: http://nagios.sourceforge.net/docs/nagioscore/3/en/embeddedp...
[1]: http://nagios.sourceforge.net/docs/nagioscore/3/en/epnplugin...
The default out of the box config is something you monitor a handful of machines with, but Nagios is regularly run with 100k+ active checks. That requries some careful architecting.
It's been working like a champ. I will be reinstalling that server soon - does anyone have any insight on Shinken vs. Nagios or Icinga?
We migrated over to Icinga 1.x without much trouble. Our existing config still worked, NConf still worked, plugins still worked.
Shinken sounds interesting, but I'd never heard of it. What I'm trying to understand is why anyone would have so much invested in their Nagios configuration syntax. If you have more than a few dozen hosts, you are probably already using a higher level configuration abstraction. NConf, for example.
But once you stop counting servers by the dozens, you are hopefully using an even higher level configuration abstraction for data center management. In which case you either already committed to a Big Vendor years ago, or you'd have no incentive to switch away from a legacy Nagios deployment.
Anyone who needs to replace a Nagios deployment and prioritizes configuration file level compatibility in their new monitoring platform sounds crazy to me.
(edit: plugin compatibility is all that anyone should care about, and even that's just based on good decisions about uniform output)
I totally agree, and as I mentioned, configuration compatibility was NOT a selling point at all - I just mentioned it as a property of shinken that makes it (supposedly) a drop-in nagios replacement without even the need to reconfigure.
I see that as a selling point for testing (not relevant to me because I didn't have a previous nagios installation) - to test shinken for real, you just install it (a couple of apt-gets) and run it - you don't need to configure it.
Once you've actually decided to commit to a new system, configuration compatibility with your old system is much less important, of course (provided it's not a tangled mess of 1000 hand edited files)
Your point about the ease of compatibility testing is quite salient.
https://wiki.icinga.org/display/Dev/Bug+and+Feature+Comparis...
Ethan's comment about the wiki page above is quite critical: "This comparison on the wiki is both inaccurate and skewed, as it contains incorrect bug data and mixes new features that have not yet been implemented. This has apparently been done with the intention of trying to make Nagios look bad in comparison to Icinga."
BUT
ffmpeg vs libav was rediculus and Ubuntu used EDIT avcon and put a weird warning on ffmpeg on it being deprecated and please install EDIT avcon just got under my skin!!! [1]
[1] http://stackoverflow.com/questions/9477115/what-are-the-diff...
Nagios is a monitoring server. It's the thing that sends emails and other notifications when various checks it runs periodically fail. Stuff like: process dead, CPU 100%, disk IO through the roof, database not responding, etc.
The plugins that were talked about so much are small pieces of software that check if enough free disk space or RAM are available, that processes are running, ports are open etc.
And then the core component aggregates the results from those checks, do alarming if configured, threshold checks (for example only five failed checks in a row result in alarming) and offer a web GUI where one can see all checks for a given host, get an overview of hosts with failing checks etc.
I didn't have any contact with monitoring tools before joining an ISP, which relies heavily on them.
So it's not the coolest. Gotcha!