We have definitely outgrown Orion, and a lot of stuff in Orion is very rough, sloppy, and not well integrated.
We don't really like the idea of cloud hosted monitoring (which is a lot of what more modern monitoring systems are). Alternatives also seem very expensive.
So if we are going to make a investment (cash or labor) I would rather we get a system that fulfills all of our needs (fit) and share it with everyone.
The logging tool is the only one I'm not sure about, if you're parsing 1200 events a second. I'm not familiar enough with Orion to understand why it was insufficient.
Would you say these in-house tools are primarily about integrating / presenting the data, or are there custom parts doing more heavy lifting too?
A lot of these tools are about integrating / presenting but that isn't entirely the case. Status does its own polling of things like Redis and SQL. Realog parses the data, structures it in Redis etc, and the patching dashboard can kick of updates.
One thing Nagios handles well, although it isn't exactly well polished, is distribution, scheduling, and aggregation of the polling. Also, anecdotally, "nobody" seems to know how easy it is to do simple-intermediate monitoring of MS-SQL in Nagios. I happen to be using this, at the moment:
https://github.com/scot0357/check_mssql_collection
edit: And thank you for the concise explanation of Hiera. I had entirely ignored it as just another add-on as Puppet "goes enterprise"
We use Puppet to manage our Linux infrastructure and are testing Desired State Configuration for our Windows systems. Puppet and DSC solve a different problem than monitoring.
Our tools span the gamut of aggregating and presenting data, to doing "heavy lifting" of things like installing patches, managing our load balancers, and removing bad query plans from our database servers.
Our in-house tooling fills in gaps where existing tooling wasn't responsive enough or made it difficult to deal with certain edge cases. General purpose tools like Orion satisfy 80% of our use cases, but we are fanatical about performance and functionality so we want to fill in that additional 20%. In fact, the majority of our Orion monitors are custom script monitors, which we have to create ourselves as it is.
If what we do can work for others (I've worked in several environments with Orion and had the same issues in each environment), that's a bonus.
We've found we are building a significant number of these projects and that's why we are looking for a dedicated developer for our team.
These are definitely distinct from monitoring, but monitoring configuration should be populated from the same configuration store as Puppet. (Sometimes Puppet is the configuration store.)
I've never evaluated Orion in particular, but it's slightly puzzling if you're creating a lot of custom monitoring scripts. In the low-touch Nagios deployments I've seen, this is often because someone didn't understand good places to use macros, and centralize more of the parameters.
Configuration systems have historically looked at config as bits-on-a-disk.
Supervisor systems look at bits-in-memory.
And they emerged and evolved independently. So there's pain points and impedance mismatches regardless of which one you start with.
What's needed is a system that sees that configuration and supervision are the same problem: you have a directed graph of what a system can look like, plus a compare-and-repair mechanism to drag the system to that state frequently.
I've personally looked at Chef, Puppet and Cfengine, none of which really does all of them the way I would like.
http://chester.id.au/2012/06/27/a-not-sobrief-aside-on-reign...
BTW, we made heavy use of setting and logging http headers too. One trick I liked was capturing performance timing metrics as a request was processed and stuffing it into a response header as the response went out. We then logged the response headers, which gave us the ability to report on the performance metrics. We also had a debug mode in the app on the browser side so we could see the performance metrics from the headers there too.
~70 events per second doesn't sound like much to capture and aggregate. How much of this parsing did you need to perform in real time? Creating a unique token to pair requests/responses shouldn't add much overhead at all.
- No real-time parsing; it's all nightly batch processing after devops rotates the Apache server logs to a storage volume. The logs sit there for a while then get compressed and moved to offline tape archives.
- No DB storage of the logs; space was too expensive and the Oracle database we had couldn't have kept up. It was already heavily burdened with a completely separate usage statistics system that fed into user-facing reporting and billing, which had a much higher event rate, about 100x higher, than the http logs.
- We had unique tokens, but they identified a particular user session that tied together all of the user's http requests from login to logoff/abandonment, and which also tied into the Oracle-based statistics for that user, that user's organization, and the customer responsible for the user (often multi-organization). My reports had breakdowns for individual user experiences, session-level metrics, and user type/organization/customer/region/etc metrics.
- I don't recall how long the analysis took; it was between half an hour to two hours I think. A lot of that time was spent on disk I/O reading the logs. I had optimized the parsing, analysis, and results recording about as much as I could.
- This stuff was written in Perl, and ran on Solaris servers from that time era... probably not a lot more powerful than a handful of smartphones today, though they did have lots of cpus. I don't think traffic has grown much since I left the company (we had pretty full market penetration already) so it's likely those servers haven't been upgraded.
I think I have a good idea of how businesses (at a high level) have failed to understand Moore's law from 2000-present. I'm curious what those failures of understanding were like from 1985-2000.
We all know that technology has been advancing rapidly, but these specific anecdotes of organizations paying a million dollars just for the backing storage of a system that you can essentially get for free from Google now...
Actually, they're probably still paying over $100/GB. The whole datacenter was outsourced to Perot Systems in the mid-2000s, and the storage fees were astronomical. We calculated that Perot must pay a separate tech to stare at each individual hard drive with a replacement in-hand in case any errors were reported. At least, they could afford to do that with what we were paying them for storage.
2. System lock-in, I like to have the data and be able to query it as I see fit
3. If an event happens where the facility is cut off from the Internet you won't be able to tell what happened inside the facility (unless maybe there is a store and forward agent, but even then it is only useful after the event.
4. Latency monitoring (again maybe an agent can help) but if I see changes in response time I can't tell if that is the WAN or not.
Basically it is an additional perspective and a redundant level of monitoring. But I view it mostly as an up/down layer of monitoring. Extremely important, but simple and not really meant to give insight into the complexity of our system.