Riemann – A network monitoring system
riemann.io
riemann.io
- You must pick up Clojure to understand and configure Riemann (we're not a Clojure shop, so this is a non-trivial requirement)
- Config file isn't a config file, it's an executed bit of Clojure code
- Riemann is not a replacement for an alerting mechanism, it's another signal for alerting mechanisms (though since it's Clojure and the configuration file is a Clojure script, you can absolutely hack it into becoming an alerting system)
- Riemann is not a replacement for a trend graphing mechanism.
- There are other solutions which can be piped together to get the 80% of the functionality we wanted from Riemann (Graphite + Skyline) in much less invested time
Skyline link: https://github.com/etsy/skyline
>- You must pick up Clojure to understand and configure Riemann (we're not a Clojure shop, so this is a non-trivial requirement) >- Config file isn't a config file, it's an executed bit of Clojure code
This is actually great -- static files quickly become their own franken-languages, with code generating config files.
>- Riemann is not a replacement for an alerting mechanism, it's another signal for alerting mechanisms (though since it's Clojure and the configuration file is a Clojure script, you can absolutely hack it into becoming an alerting system) >- Riemann is not a replacement for a trend graphing mechanism.
You probably don't want another alerting mechanism; you probably already have pagerduty or something else -- what you want is a rich way to create the alert.
This is the heart of why we use Riemann. When we first started using it 2 years ago, we had thousands of different types of error emails per day (due to monitoring thousands of retail stores, all with their quirks). Because Riemann config is just code, we were able to build systems and abstractions on top of it for describing the various error types and their semantics. E.g If 500s are being returned from service A, only alert us if > 1% of those requests failed in the last 2 minutes. You can get these kinds of rules in something like Nagios, but if you want customization, you have to deal with plugins. Here, it's just code. If we don't like it, we change it. The result is that there's no excuse to setup gmail filters. You can ensure that all errors are actionable.
For stream processing engines, configuration will be code. Unfortunate, but unavoidable.
> Riemann is not a replacement for an alerting mechanism
> Riemann is not a replacement for a trend graphing mechanism.
Indeed it is not. It's misadvertised as a monitoring solution, while it's a stream processing engine.
What I think of it is that you're supposed build a monitoring system on top of stream processing engine. It's a pity Riemann doesn't allow to subscribe to its streams from the outside, so to add any message destination you need to update its config.
I honestly don't think it's unavoidable, so long as you separate the configuration (i.e. hosts, thresholds, outputs, etc) from your processing logic. Of course, this requires additional development work from within the "configuration" file.
There aren't many examples of when code is a configuration parameter for service (generic RPC server for sysadmins being another example I've encountered), but there are some.
How so?
Kafka is a stream processing engine that uses plain old Zookeeper data structures for config.
Edit: Kafka also seems to have the missing features you mentioned if Riemann should be taken seriously as a general-purpose stream processing engine.
I'd also argue that Zookeeper nodes are anything but "plain" :)
With my programmer's hat on, they're also harder to populate programmatically, so I have a hard time justifying their use.
Any complex config file runs that kind of risk though, whether it's in a well-known programming language or an ad-hoc DSL. My preferred approach is to include most of the config in the regular code (subject to the normal review/release process), with the only thing on the server being a one-line "which config to use" setting (e.g. dev/stag/prod). Of course that has its own problems.
> With my programmer's hat on, they're also harder to populate programmatically, so I have a hard time justifying their use.
Not at all true in the case of Clojure - it's just S-expressions, very easy to write, parse or modify programatically. I agree that a config structure should have good programmatic access, but to my mind that's an argument for using a language with a good metamodel rather than anything else.
The major difference is that Clojure (Python, Lua, Perl, et al) gives you all the tools right out of the box, whereas with a DSL you should be severely restricted from doing things like reading/writing to disk, making network calls, or executing other binaries.
Granted, there are possibly ways to break out of the sandbox, but it's the difference between giving the thief a set of master keys and $50 for a U-Haul and making them work to enter every safe you have on the premises.
/me takes off the tinfoil hat
The point GP is making, and with which I agree, is that executable configurations can be dangerous if not sandboxed and even then still carry an elevated risk versus a parser. We are speaking relatively; it is absolutely still a risk to parse user input as a config, but less so than a full programming environment being immediately available to a malicious config writer.
Stepping back and identifying the malicious vector is worth it here, though, as there's a case to be made that configurations are the domain of administrators and should be secured accordingly via external means. Then the problem is recentered.
I actually think Clojure is a huge selling point.. seriously you should see the crap that Rackspace has https://www.rackspace.com/knowledge_center/article/alarm-lan... which I'm ashamed to say we use it (the lua monitors though are cool and its free monitoring infrastructure).. and yes its not the same as Riemann as Riemann is not exactly just an alerting tool.
And that is sort of the problem.. Riemann is a tool that does one thing really well but has not that good of a UI.. sadly we want prettier graphs and less granular tool.. a better nagios.
to be fair though it does say it's no longer maintained.
What exactly here is the problem for you? That config is not a config, or that it's specifically Clojure?
I like it so much that I did an experiment to implement it in C++
https://github.com/juruen/cavalieri
My implementation sucks, but I had a lot of fun working on it and I got to learn how Riemann works better.
Was surprised to see the most comprehensive and will written documentation I've ever seen on Github!
There's a sample chapter available, which covers the initial Riemann implementation and a Clojure "getting started" guide which should help anyone - even if you're not interested in the rest of the book! :)
If you're coming from Nagios (or not), and you'd like something that will schedule Nagios event scripts (and others) and send them to Riemann, I have been using this in production since mid-2013: https://github.com/bmhatfield/riemann-sumd
It allows you to tap into the huge ecosystem that is Nagios monitors, without requiring any other Nagios component at all. It just translates the output into a Riemann event.
It's only part of the stack, but it's great for routing some stats to this TSDB and other stats to that TSDB. It's also great for detecting anomalies and sending updates to wherever you want them to go.
For us, this happens in 150 lines of clojure plus 150 lines of unit tests. I know that's fairly meaningless without knowing more about our system -- but the point is, it's very expressive so you get a lot done in a few lines of code. And therefore, don't worry too much about it being clojure.
While I'm unconvinced their custom JavaScripty DSL (TICKscript) is actually preferable to Clojure or even can be read without careful, quite LISP-like indentation, it is pretty similar in basic functionality to Riemann and is definitely not-Clojure*
see: https://influxdata.com/blog/announcing-kapacitor-an-open-sou...
https://docs.influxdata.com/kapacitor/v0.2/tick/
https://github.com/influxdata/kapacitor
*at worst, it's a mangled subset of clojure with extraneous dots and the parentheses in the wrong places :-)
Maps and lists are gonna be maps and lists...
It's a good thing that it's Clojure and not some freakish turing-complete XML-based configuration file.
Greenspun's 10th rule? http://c2.com/cgi/wiki?GreenspunsTenthRuleOfProgramming
It's been a breeze, rather worry free and its very good collectd support has enabled us to cover very interesting use cases at Exoscale.
I wrote a starter guide that might be helpful https://yogthos.github.io/ClojureDistilled.html
There's also a free book that's very good http://www.braveclojure.com/
Although Brave and True looks like a great book, its not for me. I would love a book which blazes me through things rather than hand holding me through every concept. I have been looking at Living Clojure and will probably get that.
I know that JS isn't the same thing as Clojure but the ideas in this book work with really any language. After reading this book I'm a better Python programmer.
But yeah, there's the Lisp tradition and the typed tradition and they're almost entirely separate, but through accidents of history we call them both "functional". So in the same way that you'd learn an OO language and a functional language, I'd say it's worth learning one of each.
Riemann rocks, just not as monitoring system.
Bosun link: http://bosun.org/
(Ok, you do need some knowledge at the parent if you want to raise alerts but you'd need that anyway.)
If you can't easily do that with your existing infrastructure, you should fix that first. I've written about this at http://www.robustperception.io/you-look-good-have-you-lost-m...
Your article is a good one but in my experience, many companies are still many years from being able to implement that kind of database:machine knowledge consistency.
Alerting based on state changes is fragile, it's better to compare against what you expect to see.
The lack of a plugin system, and the reliance on HTTP to collect, means writing small collectors is a pain. You can't run 23 different daemons, each on their own port, to collect stats from things like PostgreSQL stats or RabbitMQ.
We opted for the "text directory" way, where we populate a directory with .prom files that node_exporter automatically picks up. It's not ideal: It means a whole bunch of collectors run via Cron jobs, which themselves need to be monitored; it means if we remove a collector, we also have to clean up its .prom files; and in the end it meant we had to invent our own little plugin system in order to avoid writing a lot of boilerplate code needed to manage the different collectors we use. We'd love to share our collectors, but due to that last point, our collectors are less reusable than we'd like.
Prometheus itself has been quite stable, but it still has some rough edges:
* If anything goes wrong with its database files, it tends to just crash, and the only way out is to wipe the entire database (e.g., see [1]).
* There's no way to do snapshots of the database.
* The team is rather cavalier about backwards compatibility. We've experienced at least one version upgrade where they changed the database format and didn't provide any upgrade tools, so people were forced to start their metrics history from scratch. I know that it's pre-1.0, but still, they knew perfectly well that people were running it in production. The alert manager was also written from scratch recently, with a whole new config format. With several releases, every tool has had its command line flags changed ever so subtly, too.
* The lack of packages (Debian/Ubuntu in our case) is also problematic. Fortunately I've got a script now that grabs a release and bundles a .deb from it, but I'd vastly preferred real releases.
* No syslog support is not acceptable in this day and age. Our Upstart scripts spawns a "logger" subprocess to catch stderr. Not everyone is running under Docker.
We have many ways to plugin to Prometheus across the ecosystem, the textfile collector you're using is one of them.
> You can't run 23 different daemons, each on their own port, to collect stats from things like PostgreSQL stats or RabbitMQ.
There's no fundamental challenge with this approach. If you've got good basic infrastructure, particularly configuration management, the rollout of each should be a small operational task. If it's a major challenge, then your problem probably isn't with the Prometheus architecture.
> it means if we remove a collector, we also have to clean up its .prom files
There's several problems arising from this approach, this is one of them. You can also expect odd artifacts in graphs.
The textfile collector is only intended for machine-level metrics, by putting service level metrics in there you're missing out on a big win of Prometheus by thinking in terms of machines rather than services.
Fighting against the architecture means you're not getting the maximum benefits from Prometheus, this would be easier with exporters and service discovery.
> which themselves need to be monitored
Are you aware that the node exporter exports the mtime of all the textfile collector files? That's there to make monitoring of them easier.
> If anything goes wrong with its database files, it tends to just crash, and the only way out is to wipe the entire database
As far as we're aware, the only way that happens is if you run out of disk space. If you've evidence otherwise please let us know, so we can prioritize accordingly.
> We've experienced at least one version upgrade where they changed the database format and didn't provide any upgrade tools, so people were forced to start their metrics history from scratch. I know that it's pre-1.0, but still, they knew perfectly well that people were running it in production.
We broke backwards compatibility in the storage format once, and there's no plans to do so again. The core developers who were all running it in production didn't see it as worthwhile to write a converter, and noone else stepped up.
> The alert manager was also written from scratch recently, with a whole new config format.
The old alertmanager has always been flagged as very experimental, as it was a functioning PoC. The rewrite was always been on the cards, and this came up regularly.
This is all part of evolving the system to be better for everyone. If we tried to keep perfect backwards compatibility then we couldn't remove warts, bugs and misfeatures. We aren't afraid to deprecate where it makes sense to do so, and have transition plans where practical.
> The lack of packages (Debian/Ubuntu in our case) is also problematic.
There are packages in Debian proper, and nightlies at http://deb.robustperception.io/
> No syslog support is not acceptable in this day and age.
That's in the latest versions.
The high level problem is that there's so many different ways to do logging that we can't sanely support them all. For every X there is someone who thinks it's essential.
We do have good basic infrastructure, thanks. We use Puppet and have a decent deploy system that performs atomic deploys from Git.
We also do think in terms of services. But the exporter has to run somewhere. About half of our exporters are machine-specific (reads local stats from files or proc or whatever), about half run on the Prometheus node itself and talk to services like ElasticSearch or Postgres.
The problem is operational overhead of maintaining a dozen daemons per box, each with its own allocated port. It's not rocket science, just annoying. What could be a small script becomes something unnecessarily big. That's why we're sticking to "textfile" for now. When I have time, my plan is to write a small HTTP server that spawns plugin as subprocesses that emit their metrics via stdout, which seems like a much more reasonable, low-maintenance solution, and something node_exporter ought to support in the first place.
If there's useful stats from /proc we're missing, we accept PRs.
> about half run on the Prometheus node itself and talk to services like ElasticSearch or Postgres.
That doesn't sound right, those two should run on the ES/Postgres nodes.
Only things like the blackbox exporter and snmp exporter should be considered for running beside Prometheus.
> When I have time, my plan is to write a small HTTP server that spawns plugin as subprocesses that emit their metrics via stdout, which seems like a much more reasonable, low-maintenance solution, and something node_exporter ought to support in the first place.
We went with the textfile collector approach, rather than reinventing cron.
> rather than reinventing cron.
But Prometheus already has a job interval setting. With text files emitted with cron jobs, you get the silly effect where metrics are produced at intervals which don't correspond with the job interval.
Performant - In the realm of 6 oid monitoring of 50,000+ devices in 5 minute intervals
I like the look of the config structure, being Clojure.
Jokes aside.. while the stream processing on events seem powerful there was something similar in graphite but probably not as advanced or easy to use. However the push approach brings its own limitations specially on existing setups.
- nodequery source (and ecosystem) is closed
- "periodically store various system data" "minute" VS "low-latency"
- email notifications only
- no query language, but a JSON API → overhead
Riemann looks more free (libre as well as liberty of usage, you can just monitor anything), more in-depth (does more with lower latency objectives) and better designed (DSL for rules, protobuff over tcp and udp instead of json over http over tcp).