Why monitoring sucks — for now
gigaom.com
gigaom.com
For example, email. Our email sending library requires various references to internal ips, external ips, and/or domains it can send to/from/through. Then our cfengine files have the same exact references in multiple locations: sendmail settings, dns settings, and network settings. Then our nagios alert system, again, has references to many of them. Oh, then we need to have a wiki for human readable format rather than having to dig through these scripts. And that's just for email. Nevermind databases, caches, app servers, etc.
We have health checks within custom developed applications for failover, and Nagios health checks. Sometimes Nagios will rely on the app health checks, but the app never relies on ops related monitoring. It just seems like there's quite a bit of effort being wasted here as we're doing the same thing twice in two different places.
The frustrating part is when I change one, I have to go through and change them all. It seems like there could be a much better system than this. I would imagine many come up with custom scripts for all of this, but there has to be a better way.
Also, I don't see the issue with nagios asking apps "you still breathing?". And why does a library know about IPs (if it's strictly a library)?
You raised many interesting points.
[edit] I can also attest to the awesomeness of Boundary's platform. These guys have a killer UI, excellent reliability, and collect important, typically invisible data. I started work on Reimann because I needed to handle more than just network traffic--from the rate of feed item fanouts to a breakdown of memory consumption across all hosts. I also have different dashboard requirements. Regardless, I'm excited to see Boundary's take on the problem.
disclosure: I'm just a really really happy customer
New Relic is indeed an excellent monitoring tool but it's not perfect. Alerts do lag behind anywhere from two to 5 minutes, and there are no context-sensitive alerts... at least not that I'm aware of.