A Practical guide to StatsD/Graphite monitoring
matt.aimonetti.net
matt.aimonetti.net
I have no alternative to suggest, however. Perhaps Cube [3], but unclear if it has any user community.
[0] http://stackoverflow.com/questions/10820119/graphite-is-not-... [1] http://graphite.readthedocs.org/en/1.0/functions.html#graphi... [2] https://github.com/etsy/statsd/issues/98 [3] https://github.com/square/cube
"xFilesFactor should be a floating point number between 0 and 1, and specifies what fraction of the previous retention level’s slots must have non-null values in order to aggregate to a non-null value. The default is 0.5." - http://graphite.readthedocs.org/en/1.0/config-carbon.html
Re [1]: How would you expect the presentation layer to present >n data points using n pixels?
Graphite doesn't "change your data". Presentation of data != the data itself, just as a map of a city != the city itself.
Sure, and many people do exactly that. The point is that a new user to graphite is likely to be surprised by this behavior. (I would further bet that a reasonable fraction of statsd+graphite users end up viewing incorrect data without realizing it, especially given the statsd focus on count data, for which the default aggregationMethod setting is exactly the wrong choice.)
(And even awareness of this behavior isn't quite enough, since every user needs to also remember their server's exact storage configuration, lest they inadvertently expand their plot across a retention boundary.)
> How would you expect the presentation layer to present n data points using n pixels?
The same way that most plotting tools do so: by overdrawing. Yes, one ends up with a solid block of pixels if the data are noisy and the plot is small, but that outcome is easily understood and has the easily understood solution of explicitly aggregating appropriately. Graphite instead takes the approach of implicitly aggregating based on how wide the plot is rendered in a given interface. That behavior is, at the very least, surprising.
One nitpick- You don't need to use statsd as an intermediary in order for your application to send metrics via UDP; just set ENABLE_UDP_LISTENER to True in carbon.conf and graphite will accept metrics on UDP itself. Other options are TCP(obviously) and AMQP.
I love how simple Graphite's plaintext protocol is; it's nothing more than a line of text with <metric path> <metric value> <metric timestamp>. This has lead lots of software to integrate graphite support and makes it easy to do yourself. In a pinch I've even set up a cronjob reading a value from /proc and sending it to graphite via netcat.
Graphite shines at generating graphs, but it's ability to return JSON is also very useful. For example, I've written a script (https://github.com/sciurus/grallect) that plugs into Nagios and generates alerts based on system metrics sent by Collectd to Graphite.
My two frustrations with graphite-
You have to choose a single aggregation method. I'd like to be able to store the average, minimum, and maximum values.
Sometimes I find it hard to query for the data I want. E.G. To check the percentage of space used on each filesystem I have to fetch example.com.df-.df_complex-used and example.com.df-.df_complex-free separately and calculate the percentages myself because asPercent(example.df-*.df_complex-{used,free}) would combine all the filesystems.
It's written in pure c and behaves like you would expect statsd to, with some additional improvements. I'm definitely more comfortable deploying it as opposed to installing and managing a node.js application.
[1] https://github.com/etsy/statsd/blob/master/lib/set.js [2] https://github.com/armon/statsite/blob/master/src/set.c [3] http://research.google.com/pubs/pub40671.html
Feel free to point out useful things graphite could do better (constructively only) and/or some of your favorite posts or tools used with graphite. We aren't too far off from 2 quite massive releases (0.9.11 / 0.10) and are thinking about departing from some of the legacy bits moving forward. I'm looking at you python 2.4
When I last checked there didn't seem to be solid documentation on how to get it all setup. Searching now, this looks promising: http://amin.bitbucket.org/posts/graphite-mac-homebrew.html
Thank you for graphite and thanks for being so receptive!
I think graphite would greatly improve by having the "web" part split into different apps/packages (API, graphing and frontend/dashboard).
That way, people could install whatever they wanted. Imagine "only" having to improve the backend while other people create amazing dashboards (which, right now, is already happening anyway... there's such a fragmentation in the available frontends...)
I'd love to help with the split if you deem it worthy & if you need a hand :) I will create a GH issue anyway :D
Again, THANKS!
It's very affordable and dramatically simplifies management of your stats. I am so glad to put Graphite behind me.
If you have the same metrics coming in from multiple app servers, how can they be viewed in aggregate and separately?
When you display your metrics you can query for: .accounts.*.http.post.authenticate.response_time
To get a breakdown per machine (and per client). or you can sum the metrics still using the wildcard.
Sometimes the best solution is to ignore the system provided libraries and build your own environment.