Prometheus: A Next-Generation Monitoring System [video]
usenix.org
usenix.org
They recommend starting with PromDash [1], their Rails-based visualization tool, but I didn't find the console-based templates [2] very difficult to dive into (Although more examples would be handy).
Worth checking out, IMHO.
I still think something like Riemann/Influxdb/Graphana is the best bet for a roll your own type solution. Riemann is a really flexible router/aggregator/alerter with realtime metric display and it can throw everything over to influx/grafana for dashboarding/time series analysis.
- Assuming metrics are stored in memory, they'll be lost if the app or server recycles. Of course, a push model should probably batch so there's probably some kind of overlap here.
- I've heard pull is good for detecting if a server is down. The flip side is no automatic discovery and dynamic scaling where a server going down might not necessarily be a bad thing.
- Security. Not super thrilled with exposing something that says "Hey come get my data".
Similarly if the server dies before sending the metrics, so that's a draw.
> The flip side is no automatic discovery and dynamic scaling where a server going down might not necessarily be a bad thing.
That's more about top-down vs. bottom-up target discovery than the scraping itself, see http://www.boxever.com/push-vs-pull-for-monitoring
As of Prometheus 0.14.0 which was released earlier in the week, Prometheus supports target discovery. http://prometheus.io/blog/2015/06/01/advanced-service-discov...
> Security. Not super thrilled with exposing something that says "Hey come get my data".
That's a fair concern, though I'd expect your metric endpoint not to be internet accessible (and some of the push options will have a HTTP endpoint for debugging).