Riemann, a distributed systems monitor (built in Clojure)
aphyr.github.com
aphyr.github.com
Bigger projects i've used either citect or wonderware, both of which get the job done but show their age and are at times painful to use, although not as horrible as many of the other legacy control system HMI software out there.
Mostly data is collected by polling modbus slaves, although OPC would be another important protocol to support.
It seems like wonderare or citect is ready to be replaced by a distributed system that uses the web browser as the display client. systems monitoring software such as nagios, openNMS, cacti overlaps with the control system HMI software arena as both
1. display real time data, preferably with some context (eg gauges to indicate how close to maximum or minimum limit, alarm or shutdown thresholds the variable is) and sometimes overlayed on a diagram to assist in visualizing or understanding the process
2. "trend" (log and plot) historical data for analysis and reporting purposes. Better yet would be interactive plotting (zoom etc).
I've been wondering about graphite (which riemann can use) as part of the solution, and people seem to be producing great plots with d3.js. As an aside I've used kst and veusz for desktop interactive plotting with success.
In summary: if riemann supported the modbus protocol it could be useful for control systems.
2. Yeah, that would be great. The historical event store space is pretty terrible right now, and it's such a big problem that I doubt I could realistically tackle it. Librato Metrics and Hosted Graphite are both approaching this as a service, and there's openTSDB if you have Hadoop people in-staff. Riemann has out-of-the-box integration with librato and graphite, but I haven't set up an openTSDB cluster yet.
Modbus: that'd be cool. Implementing a Riemann server (i.e. a thing that accepts events from the wire) is pretty straightforward, though I'd need to understand the protocol. If you're interested in building it, I'm happy to discuss how.
It might be easier to think of streams as literal streams, rivulets, deltas, and tributaries, which events flow through, rather than a query language with well-defined clauses like sql.
From skimming the docs, it looks like Esper is much bigger, much cooler, and with a more abstract version of events. It implements a lot of the primitives I've been considering but haven't built yet. It looks more difficult to set up, and has a commercial offering for support and HA; neither of which are present in Riemann right now.[1]
I recently finished building a monitoring system using Esper and JRuby since the client asked specifically for that, but I wished I had used Riemann from the beginning.
[1] https://groups.google.com/forum/#!msg/riemann-users/GhVMYJow...
Riemann is also more general than Esper, in that you can define arbitrary operations on events. It sacrifices having an up-front query language in favor of composable functions with stateful side effects. If you want to write a stream that restarts an EC2 instance on failure, it'll be a composable first-class citizen and you can write it right in the config file. Same goes for a stream that pulls in, say, parallel colt to do some heavy statistical lifting. On the other hand, Riemann doesn't include the full range of Esper queries as builtin streams yet, and the ones that are there haven't been optimized to the same degree.
Hope this helps clarify things. :)
That said, I think you'll find many of the ideas in Riemann to be radically simple. The config file is just a Clojure program. Streams are just functions that take events. Events are just maps of keys to values. Everything is an event: there is no concept of a first-class host or service, no need to update the config when you add a host, and no poller loops.
Riemann tries to draw strong boundaries between the different layers of monitoring. It speaks a simple network protocol and interoperates with other systems for event collection, visualization, alerting, and storage, instead of building in those systems. In many ways Riemann is defined not by what it includes, but by what it leaves out.
That said, there's a lot of work required to make simple abstractions behave correctly, especially around IO and error handling. Wherever possible I try to draw clear internal boundaries to isolate this complexity, but it's still there. If you have specific complaints about code or interfaces which seem too complex to you, I'd be happy to try and explain or change them.
Reducing hairy things to simple abstractions can save weeks of work in a matter of minutes. It is the single most powerful programming technique I know of. And I'm going to seriously consider switching a bunch of stuff over to Riemann.
i see that you use Protocol Buffers. from google's page, it seems like they only work with C++, Java or Python. now this could be a problem for us. what if i want to pull events from a bash script, a delphi gui app, SNMP, Dell idrac interface, or any other event? do they have to interface over Protocol Buffers, or am i missing something here? would i have to write a glue layer?
and what about, if you have two seperate networks, and want one server to forward data for its entire lan to the other server, to process and graph them?
For pulling from other tools, I usually write a little daemon to poll and relay the data. See riemann-tools for a collection of existing tools to do just that; and you can require it as a library to write your own in just a few lines of ruby.
Forwarding between servers is built in; it's easy to aggregate events in hierachies for large-scale analysis.