Stop using Nagios (so it can die peacefully)
slideshare.net
slideshare.net
I mean, "don't use X, use ShinyX instead" is one thing (and most of the time it's a bad thing but it does occasionally turn up good ideas), but this is just So Much Worse...
Have you ever been woken up by a nagios page that automatically cleared after five minutes because the incoming queue was delayed past the alert interval?
Have you ever had your browser crash because you click on the wrong thing in the designed-in-1996-and-never-updated nagios interface and had your browser crash because it dumps 500MB of logs to your screen?
Have you ever had services wake you up with alert then clear then alert then clear again because some new intern configured a new monitor but didn't set up alerting correctly (because lol, they don't get paged, so who gives a flip if they copied and pasted the wrong template config, as is standard practice)?
Have you had to hire "nagios consultants" to figure out how to scale out your busted monitoring infrastructure because nagios was designed to run on a single core Pentium 90?
Being pro-nagios is like being pro-Russia, pro-North Korea, and pro-Rap Genius while arguing "but at least we know how bad they are and can keep them in line."
To give context about my comment about context:
* was nagios setup before you started the job?
* did you setup nagios yourself?
* is your internal process for managing nagios broken?
* culturally do you work at a place where ops is an afterthought?
* if Nagios is your technical debt do you have a way out? are you crushed by other commitments? Maybe it's more of a management/culture issue.
... hmm actually I should stop. From re-reading your comment, I can't tell how much of it is trolling (in a entertaining Skip Bayless, right wing radio, Jim Cramer kind of way).
Yup.
did you setup nagios yourself?
god, no.
is your internal process for managing nagios broken?
It was the best they knew after a dozen years of experience at other companies.
culturally do you work at a place where ops is an afterthought?
Nope. We had three people for 500 machines and two data centers.
if Nagios is your technical debt do you have a way out? are you crushed by other commitments? Maybe it's more of a management/culture issue.
It was "good enough" and nobody wanted to build out an alternative (or knew how to—the others were pure "sysadmin" people without programming backgrounds).
The problem: apathy. The solution: leave...after three and a half years.
I can't tell how much of it is trolling
Never trolling from this account, just unfocused anger with no other outlet. :)
No, because I know how to configure escalations properly.
>Have you ever had your browser crash because you click on the wrong thing in the designed-in-1996-and-never-updated nagios interface and had your browser crash because it dumps 500MB of logs to your screen?
Actually, no. I've had my browser crash due to AJAX crap all the time though. Nagios' (and Icinga classic's) interface is clear, simple and logical; it's just not 2MB of worthless javascript that wastes half my CPU time, so I can see why unpopular with some user types.
>Have you ever had services wake you up with alert then clear then alert then clear again because some new intern configured a new monitor but didn't set up alerting correctly (because lol, they don't get paged, so who gives a flip if they copied and pasted the wrong template config, as is standard practice)?
No, because I know how to use time periods, and escalations again.
>Have you had to hire "nagios consultants" to figure out how to scale out your busted monitoring infrastructure because nagios was designed to run on a single core Pentium 90?
No, because it isn't, because I know the basics of Linux performance tuning, and because I've heard of Icinga and/or distributed Nagios/Icinga systems for very large scale.
Your post reads like "Have you ever crashed your car head on into a concrete wall at 70mph because it didn't brake for me?". No amount of handholding a program can do will protect users who have no clue how to use it.
I do not by any means consider myself an expert in Nagios either - if there was such a market for consultants as you claim, I'd likely be doing it and therefore be rich, but in actual fact, it's a skill just about any mid-level or better admin has.
I've inherited a Nagios config before that was a mess, that I rebuilt from scratch in a maintainable way, as well as extended. If Nagios (or MySQL pre-Oracle, for that matter) has a problem, it's amateurs attempting it, making a mess, and others judging the quality of the tool on their sloppy work. Not unique to Nagios, by any means. If there's a criticism you can level at Nagios for that, it's the lack of documentation and examples in the config files.
I'm also not denying the existence of alternatives - OpenNMS is ok, as is Zabbix, but both are far more limited in terms of available plugins and extensibility, and by nature harder to extend. Munin is good for out of the box graphing, but relatively poor for actual monitoring/alerting and hard to write new plugins for with limited availability of additional plugins. Each one is a standalone tool that's good for a purpose, and not some vaguely defined set of programs, partly nonexistent, that everyone has to hack together for themselves.
So, nagios thinks "Service Moo hasn't contacted us in 60 seconds, ALERT!" when the update is actually in the event log, but nagios hasn't processed it yet.
That, and setting up proper retry intervals for checks that take a long time to execute.
We had a college student during his summer time write up a quick nagios add/del/modify app. Took him a few hours to bring it up, now it is so easy to replace the whole configuration(s) through it.
Same here, has never crashed on us.. on the server or on the client. I don't know what this guys is talking about, maybe he's confused with the old OpenNMS or ZenOS?
Everyone knows that Rap Genius is an oppressive regime in which you're executed for dissent, North Korea are SEO spammers, and in Russia... there is no need for monitoring software because Russian computer does not fail!
I like the idea of Sensu, and I would love to have a system built on it, but I don't have the time or energy to build that system, create the parts of it that need creating, tie them all together, and then (ha!) make sure they're all monitored so that none of my monitoring system fails.
Being pro-Nagios is like being pro-USA: it does a lot of things that annoy people, but in the end it means well, gets the job done, and you know what the caveats are. It's not perfect, but I have other problems to solve and systems to build.
The point is that you cannot know that something is better if it has unknown failure scenarios.
http://www.centreon.com/ ..or install via FAN: http://www.fullyautomatednagios.org/wordpress/
Makes life so much better.
Nagios, and a host of other popular software that's difficult to use, exist because the alternatives are so poor.
We use it for a small ISP since nearly two years and are quite happy with it. All limitations that came up during integration could be fixed by writing some small scripts. Even upgrades worked.. :)
So if you have any Zabbix pain that might be waiting for us, please share.
They do offer an API that is quite usable. I've build a simplified interface for our less trained staff members with it.
You won't find any RRD files which might be a problem for you. Until now that was only a theoretical problem for us.
Shinken was written by Jean Gabès as a proof of concept for a new Nagios architecture. Believing the new implementation was faster and more flexible than the old C code, he proposed it as the new development branch of Nagios 4.[3] This proposal was turned down by the Nagios authors, so Shinken became an independent network monitoring software application compatible with Nagios
Automate that shit. We use Nagios to monitor our infra (5000+ checks, hundreds of hosts), and chef maintains the config. Works without a hitch - and best of all - it's been running for years, and not once has anyone had to poke it with a stick.
Yes, NetSaint is old, yes, the UI is worse to look at than Putin's crotch, yes, the plugin architecture is whimsical as all shit - but... IT WORKS AND YOU CAN RELY ON IT.
This is why we moved to Zabbix:
Number of items monitored: 77342 Number of triggers enabled: 25405
Try and scale Nagios to those values and let me know how it works. We tried and we let it go a few years ago.
And yeah, it's horrible.
rspec-puppet in particular; so much pain.
People use Nagios because it works, and it gets everything right, including config if you have any clue whatsoever how to set up a good object hierarchy. The only real problem with it is maintenance, an issue which Icinga resolved long ago.
The thing is: nagios doesn't actually work for more than monitoring about 100 machines. Everything after that is a hack. It's made where you know your systems by hand, are hand configuring everything, and setting up all alerts with love and care and a kiss as you edit every file manually.
[This is your brain.] [This is your brain after losing sleep every night for a week because of nagios spurious alerting.]
The UI is also horrible, though I see they've remade that in the latest versions. In particular the UI does not update to show changes in state in a timely manner. If you tail the logs, you see that the state has changed as expected when expected, but the UI is still reporting old data for minutes to come.
I haven't tried a lot of systems and am really no expert in monitoring, but saying Nagios gets everything right... just makes me feel oily. It can be best-of-class and still be shitty.
http://rbcarleton.com/send_nsca_service_check.shtml
Use some kind of automation system like CFEngine for the distributed scheduling. Some assembly required ;)
That being said, there _are_ limitations (but scaling is not one of them) to nagios, and the configuration is definitively not something you do in cute widdle config.yml.
Combined with the recent negative developments with the corporation behind the Nagios trademark and the enterprise version - which the author fails to mention, and should be even more alarming - one should at least consider using and contributing to bareos, the (hopefully) true OSS fork of nagios (I will for future deployments).
Oh hey, look at that, it would even pose the possibility to _improve_ the software. (I really don't see why an almost-complete rewrite of nagios should be necessary. Even after reading these slides(2)).
(1) That includes the author.
(2) Or rather especially after reading them.
Bareos [1] is more an fork of Bacula [2], Isn't it? Did you mean Icinga [3]?
The irony here is that with bacula a similar thing happened.
I haven't tried to write plugins for it but it comes with a lot out of the box and it's working well.
If you want to be lazy and not set up a complete graphing infrastructure, just run collectd, have it automatically log all your statistics, and use it with the bundled https://collectd.org/wiki/index.php/Collectd-web package to view your history when needed.
Recently I've been implementing salt stack, and collectd is really easy to automate config of. I've got salt fully configuring collectd, and then using the circonus api to setup monitoring rules automatically per applied states. It's a beautiful thing.
(Full disclosure: I'm working with Scalyr, but you should still try it.)
My small startup company already has more than 10 (virtual) servers which would cost us 500 dollars a month to monitor which is more than the servers itself cost.
1. With our custom parser you can replace or delete confidential information from your logs before they're stored on our servers.
2. We're exploring a different pricing structure right now that would address this scenario. If this is your only hesitation, I hope you check us out anyway and contact us about pricing.
Also -- lots of companies are seriously concerned about pushing their data externally. Making it a hard sell.
However, as a long tail service, this looks great.
This is both error prone and utterly defeats the purpose. Why would I pay a bunch of money for somebody else to manage my logs when I'd just have to keep them all anyway so I can get at the unredacted versions when there's a problem?
We take security very seriously, but let me turn this around into a question: what would it take for you to trust an external service to manage your logs? Some of the things we're doing:
1. SSL everywhere (including internal traffic between our backend servers).
2. We add a tag to the raw representation of every string value, so that we can verify that data never leaks across accounts. (This has never detected a problem -- except in tests, because yes, we do test it.)
3. Implementing in a "safe" language (Java), to rule out low-level buffer management bugs.
4. As Greg noted, we make it trivial for you to redact sensitive data before it leaves your server.
We are sometimes asked for an on-premises installable version of our service. We don't provide that because we're using economies of scale on the backend to completely change the log management experience: when you give us a query, every CPU and spindle in our entire cluster is briefly devoted to that query. This means that you aren't limited to graphing predefined metrics; you can do ad-hoc exploration of your entire log corpus in on the fly. E.g. display a histogram of response latencies for all requests for url XXX on server group YYY in the last 48 hours, and expect a near-instantaneous response.
I think that's a really weird question that completely fails to address the concerns that some people might have. We do logging of sales, profit margins and stuff like that. You can't have access to that because: "You're not us". If you can read our data, then we're not going to use your service and to do anything useful the logs you really do need read access.
Of cause you might have no reason to spy on our data, but the only safety is that you promise not to. We could seperate logs for different things, so webserver logs go to you, but email logs goes to an internal system, but then we would need two systems.
Because either way, that doesn't change at all the concerns he voiced regarding this particular service.
For a publicly traded company, a relaxation of relevant law to allow arbitrary data to leave the company.
An act of God. This implies, of course, certain further prerequisites that are themselves probably quite challenging to meet.
> economies of scale
Under your current plans, the money spent putting my entire infrastructure on Scalyr would pay 2-3+ developers (or some mix of developers and sysadmins) in Taiwan, where my employer is based. We are not a tiny company, but that would be a substantial and welcome manpower increase for the server team, and only a fraction of that manpower would need to be spent meeting even our "would be nices" for logging and monitoring.
This would only get worse as we grew. Realistically, I would expect us to be able to open an office in Silicon Valley and start hiring there at market rates for the amount of money we'd be giving you.
Well the no-brainers are:
1 - Encrypt everything before living my infrastructure (with my key, not SSL), only decrypt it in my infrastructure when generating reports. 2 - Anything that runs on my infrastructure is open source, and widely distributed. Bonus points for a simple protocol that I can write plugins for. 3 - Make it possible for me to back everything up, and restart everything in case you go out of business.
Those are the must have "I won't even let you get through security otherwise" features. Now that I think about it, #3 alone makes anything you can offer worst than doing it in-house.
But none of those are features that'll make your system look any good in my eyes, they are just enough for your system to not look like an enemy.
https://www.scalyr.com/login?prefillEmail=demo-account%40sca...
There's your answer to "Why won't you host your server logs (which are usually key for troubleshooting flaming boxen)?"
Everything in life is a tradeoff. If you entrust your logs to us, you run the risk that we have an outage or failure of some sort. On the other hand, internal systems can fail as well. We hope to serve people who prefer not to carry the responsibility of maintaining their own monitoring infrastructure, and/or are interested in the features and performance we provide.
Well, yeah.
Out of curiosity, what is your target customer? People running their own dedis probably are alright to grudgingly setup a monitoring solution, and people who are just using the ~=cloud=~ probably don't care.
What's your ideal customer?
This looks more like documentation and less like a product landing page.
Here are some nice examples that might help: http://land-book.com/
https://www.gosquared.com/ (pulled from the first page of land-book)
FWIW, the price actually works quite well for a lot of people. On a GB-for-GB basis, we're actually much cheaper than other hosted log management solutions -- we work hard on backend efficiency and we pass that along. But yes, if you're using small virtual servers then the pricing model breaks down. We originally went with this model to provide more predictability; log volume is often more volatile than server count. We've heard enough complaints that we've decided to move to more volume-based pricing model, we're just working out the details.
1. Dynamic registration of ephemeral systems with a monitoring platform.
2. Security monitoring of same
3. Meaningful graphing
When we are optimizing the purchase of hundreds upon hundreds of spot instances daily, where we are looking at grabbin hosts for just a couple cents an hour, the model of per host fees for things like StackDriver, CloudPassage and your service makes per-host pricing completely a no-go.
I don't have a good idea how these should be priced; but I think its important for people to understand all the other costs associated with having a solid management platform for your environment that covers all the bases and doesn't require another round of funding! :)
As for the licensing model: we're going to move to per-GB pricing anyway, so no worries there. If you'd like something more concrete today, e-mail us at contact@scalyr.com and I'm sure we can sort out any pricing concerns.
Yes, Nagios' configuration is ugly and occasionally requires sacrifices to elder gods. At the same time, I've never found any sort of monitoring/alerting I've needed done that it can't handle. As much as your service looks cool for a specific subset of monitoring, it is still missing half the hooks as to why nagios is stubbornly sticking around.
https://laur.ie/blog/2014/02/why-ill-be-letting-nagios-live-...
Not sure whether that's true or not, but they did address it...
It's also the opposite of nagios in that rather than lots of smaller moving parts, it is one big mega (java) process that does everything. Again, not necessarily a problem (though I happen to think so :), but different.
I also found opennms to be VERY complicated. I suppose nagios is though, first time around.
For some reason though, I really want to use opennms and keep going back to try it out, but eventually give up.
Your traditional monitoring systems have hand-selected features to monitor and alert for. OpenNMS will just go out and discover everything you have (and graph everything without any intervention too).
You probably aren't monitoring all the statistics on every interface of your switches (what? people have switches?), but just throw OpenNMS at your networking management subnet and it'll pick up everything for later review.
You can use OpenNMS for alerting and inventory tracking, but I prefer more extensible tools for those. Just use OpenNMS as a largely hands-off sanity check of your existing monitoring and graphing systems.
I'm not a huge fan of having yet another debian package with it's own version of ruby packaged. It does make the plugins easier to write though.
Checks need to be installed on the client (like nagios). It means that some coordination is necessary when you want to add a new check on the server side. This is largely resolved when using a configuration management system but it doesn't seem clean to me. The sensu-community repo has a lots of checks which is great to get started, some of them need some ruby gem dependencies to work though.
I had issues with malformed json config or rabbitmq disconnections which would crash the server. Because the debian packages uses the old sysvinit it wasn't restarting. Moved the init scripts to upstart and added json validation when generating the config and now it's fine.
If the tool can't trace a transaction end-to-end. I.e if a user visits your page or uses your application you need the ability to trace it from Http to EJB across any webservices and queues and ESB's right down to which queries were used in the database, if you can't do that you're using a shit monitoring suite.
Knowing infrastructure metrics is useless without knowing if it's actually affecting end-users and in which use-cases.
I recently had an opportunity to do a clean sheet build out for monitoring, so I evaluated Zabbix, Munin, and combos of statd/Graphite, etc. and none of them were better.
That said, I have a stock Nagios base config that I can install and have monitoring in five minutes. The key to Nagios configs is to define hostgroups in one file, and then create config files for each host, assigning it to a group. Then you put the service definitions in a service file. Easy peasy.
I think nagios is a piece of shit, but it is a working piece of shit.